datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lar-echr
Dataset Card for LAR-ECHR
Dataset Details
Dataset Description
Curated by: Odysseas S. Chlapanis
Funded by: Archimedes Research Unit
Language (NLP): English
License:
CC BY-NC-SA (Creative Commons / Attribution-NonCommercial-ShareAlike)
Read more: https://creativecommons.org/licenses/by-nc-sa/4.0/
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Uses… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/lar-echr.statcan-dialogue-dataset-retrieval
Statcan Dialogue Dataset (Processed for Retrieval Tasks)
This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately.
Quickstart
from datasets import load_dataset
repo = 'McGill-NLP/statcan-dialogue-dataset-retrieval'
# load english queries, training split
queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.tiny-singleturn-chat-kolc_quad2
Dataset Card for LC-QuAD 2.0 with answers
RoJBMO
RoJBMO: Junior Balkan Mathematical Olympiad Benchmark
RoJBMO is a benchmark of 508 competition mathematics problems drawn from the Junior Balkan Mathematical Olympiad (JBMO), its official shortlists, and the Romanian Team Selection Tests (TST). It is designed to evaluate the mathematical reasoning capabilities of large language models on problems that are underrepresented in existing benchmarks and resistant to training data contamination.
Sources
Source… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/RoJBMO.CHASE-QA
CHASE: Challenging AI with Synthetic Evaluations
The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-QA.ph_en_text_detoxedPhEnText Detoxed is a large-scale and multi-domain lexical data written in Philippine English and Taglish text. The news articles, religious articles and court decisions collated by the original researchers were filtered for toxicity and special characters were further preprocessed. This dataset has been configured to easily fine-tune LLaMA-based models (Alpaca, Guanaco, Vicuna, LLaMA 2, etc.) In total, this dataset contains 6.29 million rows of training data and 2.7 million rows of testing… See the full description on the dataset page: https://huggingface.co/datasets/NLPinas/ph_en_text_detoxed.genetic-counselor-freeform-questionsA collection of open-ended questions about genetic counseling, curated from:
relevant subreddits
flashcards for the ABGC Certification Examination
Also see the genetic-counselor-multiple-choice evaluation set.
A genetic counselor must be prepared to answer questions about inheritance of traits,
medical statistics, testing, empathetic and ethical conversations with patients,
and observing symptoms.
For evaluation only
The goal of this dataset is to evaluate LLMs and other AI… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-freeform-questions.genetic-counselor-multiple-choiceA collection of multiple-choice questions intended for students preparing for the
American Board of Genetic Counseling (ABGC) Certification Examination.
Also see the genetic-counselor-freeform-questions evaluation set.
A genetic counselor must be prepared to answer questions about inheritance of traits,
medical statistics, testing, empathetic and ethical conversations with patients,
and observing symptoms.
For evaluation only
The goal of this dataset is to evaluate LLMs and… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-multiple-choice.abc-multiple-choice
abc-multiple-choice Dataset
abc-multiple-choice は、競技クイズの大会「abc」で使用された4択問題を元に作成された、多肢選択式の質問応答データセットです。
データセットの詳細については、下記の発表資料を参照してください。
鈴木正敏. 4択クイズを題材にした多肢選択式日本語質問応答データセットの構築. 言語処理学会第30回年次大会 (NLP2024) 併設ワークショップ 日本語言語資源の構築と利用性の向上 (JLR2024), 2024. [PDF]
下記の GitHub リポジトリで、本データセットを用いた評価実験のスクリプトを管理しています。
https://github.com/cl-tohoku/abc-multiple-choice
ライセンス
本データセットのクイズ問題の著作権は abc/EQIDEN 実行委員会 に帰属します。
本データセットは研究目的での利用許諾を得ているものです。商用目的での利用は不可とします。
dumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.NLP_Insights_2023_2024
Key Insignts from NLP papers (2023 -2024)
This dataset is processed and compiled by @hu_yifei as part of open-source effort from the Open Research Assistant Project.
It includes key insights extracted from top tier NLP conference papers:
Year
Venue
Paper Count
2024
eacl
225
2024
naacl
564
2023
acl
1077
2023
conll
41
2023
eacl
281
2023
emnlp
1048
2023
semeval
319
2023
wmt
101
Dataset Stats
Total number of papers: 3,640
Total rows (key… See the full description on the dataset page: https://huggingface.co/datasets/yifeihu/NLP_Insights_2023_2024.trading-finance-glossary-nlp
Trading & Finance Glossary Dataset for NLP
A comprehensive, structured dataset of 209 trading and financial terminology entries designed for natural language processing applications in the financial domain.
Description
This dataset provides a curated collection of trading and financial terms with rich metadata including definitions, categorical labels, semantic relationships, contextual usage examples, and difficulty classifications. Each entry has been written to reflect… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/trading-finance-glossary-nlp.TextGraphs17-shared-task-datasetWe present a dataset for graph-based question answering. The dataset consists of <question; candidate answer> pairs. For each candidate, we present a graph that is obtained by finding the shortest path between named entities mentioned in a question and a candidate answer. As a knowledge graph, we adopted Wikidata. Our dataset has the following fields:
sample_id - an identifier for <question, candidate answer>;
question - question text;
questionEntity - comma-separated list of names (textual… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/TextGraphs17-shared-task-dataset.QA_hukum_samplesThis dataset sample was constructed by generating QA pairs from Law No. 17 of 2008 on Shipping (Undang-Undang No 17 Tahun 2008 Tentang Pelayaran). It is then manually verified by human validators and reviewers, resulting in QA pairs with fine-grained label categories.
