CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AUEB-NLP /lar-echr Dataset Card for LAR-ECHR Dataset Details Dataset Description Curated by: Odysseas S. Chlapanis Funded by: Archimedes Research Unit Language (NLP): English License: CC BY-NC-SA (Creative Commons / Attribution-NonCommercial-ShareAlike) Read more: https://creativecommons.org/licenses/by-nc-sa/4.0/ Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Uses… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/lar-echr.textquestion-answeringn<1K3 likes515 downloads1y agoHugging Face02McGill-NLP /statcan-dialogue-dataset-retrieval Statcan Dialogue Dataset (Processed for Retrieval Tasks) This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately. Quickstart from datasets import load_dataset repo = 'McGill-NLP/statcan-dialogue-dataset-retrieval' # load english queries, training split queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.textquestion-answering10K<n<100K1 likes202 downloads2y agoHugging Face03OpenLab-NLP /tiny-singleturn-chat-kotextquestion-answering10K<n<100K0 likes197 downloads10mo agoHugging Face04s-nlp /lc_quad2 Dataset Card for LC-QuAD 2.0 with answers textquestion-answering10K<n<100K1 likes135 downloads3y agoHugging Face05upb-nlp /RoJBMO RoJBMO: Junior Balkan Mathematical Olympiad Benchmark RoJBMO is a benchmark of 508 competition mathematics problems drawn from the Junior Balkan Mathematical Olympiad (JBMO), its official shortlists, and the Romanian Team Selection Tests (TST). It is designed to evaluate the mathematical reasoning capabilities of large language models on problems that are underrepresented in existing benchmarks and resistant to training data contamination. Sources Source… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/RoJBMO.tabulartext-generationn<1K1 likes92 downloads20d agoHugging Face06McGill-NLP /CHASE-QA CHASE: Challenging AI with Synthetic Evaluations The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-QA.imagequestion-answeringn<1K0 likes57 downloads2y agoHugging Face07NLPinas /ph_en_text_detoxedPhEnText Detoxed is a large-scale and multi-domain lexical data written in Philippine English and Taglish text. The news articles, religious articles and court decisions collated by the original researchers were filtered for toxicity and special characters were further preprocessed. This dataset has been configured to easily fine-tune LLaMA-based models (Alpaca, Guanaco, Vicuna, LLaMA 2, etc.) In total, this dataset contains 6.29 million rows of training data and 2.7 million rows of testing… See the full description on the dataset page: https://huggingface.co/datasets/NLPinas/ph_en_text_detoxed.texttext-generation1M<n<10M2 likes51 downloads3y agoHugging Face08monsoon-nlp /genetic-counselor-freeform-questionsA collection of open-ended questions about genetic counseling, curated from: relevant subreddits flashcards for the ABGC Certification Examination Also see the genetic-counselor-multiple-choice evaluation set. A genetic counselor must be prepared to answer questions about inheritance of traits, medical statistics, testing, empathetic and ethical conversations with patients, and observing symptoms. For evaluation only The goal of this dataset is to evaluate LLMs and other AI… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-freeform-questions.textquestion-answeringn<1K0 likes35 downloads1y agoHugging Face09monsoon-nlp /genetic-counselor-multiple-choiceA collection of multiple-choice questions intended for students preparing for the American Board of Genetic Counseling (ABGC) Certification Examination. Also see the genetic-counselor-freeform-questions evaluation set. A genetic counselor must be prepared to answer questions about inheritance of traits, medical statistics, testing, empathetic and ethical conversations with patients, and observing symptoms. For evaluation only The goal of this dataset is to evaluate LLMs and… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-multiple-choice.textquestion-answeringn<1K1 likes33 downloads2y agoHugging Face10tohoku-nlp /abc-multiple-choice abc-multiple-choice Dataset abc-multiple-choice は、競技クイズの大会「abc」で使用された4択問題を元に作成された、多肢選択式の質問応答データセットです。 データセットの詳細については、下記の発表資料を参照してください。 鈴木正敏. 4択クイズを題材にした多肢選択式日本語質問応答データセットの構築. 言語処理学会第30回年次大会 (NLP2024) 併設ワークショップ 日本語言語資源の構築と利用性の向上 (JLR2024), 2024. [PDF] 下記の GitHub リポジトリで、本データセットを用いた評価実験のスクリプトを管理しています。 https://github.com/cl-tohoku/abc-multiple-choice ライセンス 本データセットのクイズ問題の著作権は abc/EQIDEN 実行委員会 に帰属します。 本データセットは研究目的での利用許諾を得ているものです。商用目的での利用は不可とします。 tabularmultiple-choicen<1K4 likes31 downloads3y agoHugging Face11NLPForUA /dumy-zno-ukrainian-math-history-geo-r1-o1 DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers) DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks. The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian: Думи мої, думи мої, Лихо мені з вами! Нащо стали на папері Сумними рядами?.. Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.tabulartext-generation1K<n<10K2 likes29 downloads1y agoHugging Face12yifeihu /NLP_Insights_2023_2024 Key Insignts from NLP papers (2023 -2024) This dataset is processed and compiled by @hu_yifei as part of open-source effort from the Open Research Assistant Project. It includes key insights extracted from top tier NLP conference papers: Year Venue Paper Count 2024 eacl 225 2024 naacl 564 2023 acl 1077 2023 conll 41 2023 eacl 281 2023 emnlp 1048 2023 semeval 319 2023 wmt 101 Dataset Stats Total number of papers: 3,640 Total rows (key… See the full description on the dataset page: https://huggingface.co/datasets/yifeihu/NLP_Insights_2023_2024.textsummarization100K<n<1M3 likes21 downloads2y agoHugging Face13propfirmkey /trading-finance-glossary-nlp Trading & Finance Glossary Dataset for NLP A comprehensive, structured dataset of 209 trading and financial terminology entries designed for natural language processing applications in the financial domain. Description This dataset provides a curated collection of trading and financial terms with rich metadata including definitions, categorical labels, semantic relationships, contextual usage examples, and difficulty classifications. Each entry has been written to reflect… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/trading-finance-glossary-nlp.texttext-classificationn<1K2 likes19 downloads6mo agoHugging Face14s-nlp /TextGraphs17-shared-task-datasetWe present a dataset for graph-based question answering. The dataset consists of <question; candidate answer> pairs. For each candidate, we present a graph that is obtained by finding the shortest path between named entities mentioned in a question and a candidate answer. As a knowledge graph, we adopted Wikidata. Our dataset has the following fields: sample_id - an identifier for <question, candidate answer>; question - question text; questionEntity - comma-separated list of names (textual… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/TextGraphs17-shared-task-dataset.textquestion-answering10K<n<100K0 likes8 downloads3y agoHugging Face15nlp-brin-id /QA_hukum_samplesgatedThis dataset sample was constructed by generating QA pairs from Law No. 17 of 2008 on Shipping (Undang-Undang No 17 Tahun 2008 Tentang Pelayaran). It is then manually verified by human validators and reviewers, resulting in QA pairs with fine-grained label categories. tabularquestion-answeringn<1K1 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.