CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face02slovak-nlp /sklep Dataset Card for skLEP Dataset Description skLEP (General Language Understanding Evaluation benchmark for Slovak) is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. The benchmark encompasses nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities. To create this benchmark, we curated new, original datasets… See the full description on the dataset page: https://huggingface.co/datasets/slovak-nlp/sklep.textquestion-answering100K<n<1M4 likes471 downloads7mo agoHugging Face03DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes316 downloads2y agoHugging Face04lyon-nlp /mteb-fr-retrieval-syntec-s2p Syntec dataset for information retrieval This dataset has been built from the Syntec Collective bargaining agreement. Its purpose is information retrieval. Dataset Details The dataset is rather small. It is intended to be used only as a test set, for fast evaluation of models. It is split into 2 subsets : queries : it features 100 manually created questions. Each question is mapped to the article that contains the answer. documents : corresponds to the 90 articles from… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/mteb-fr-retrieval-syntec-s2p.textquestion-answeringn<1K2 likes238 downloads2y agoHugging Face05lasha-nlp /CONDAQA Dataset Card for CondaQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation Dataset Summary Data from the EMNLP 2022 paper by Ravichander et al.: "CondaQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation". If you use this dataset, we would appreciate you citing our work: @inproceedings{ravichander-et-al-2022-condaqa, title={CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation}… See the full description on the dataset page: https://huggingface.co/datasets/lasha-nlp/CONDAQA.tabularquestion-answering10K<n<100K5 likes231 downloads4y agoHugging Face06turkish-nlp-suite /InstrucTurca InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more. Dataset content BI55/MedText checkai/instruction-poems garage-bAInd/Open-Platypus Locutusque/ColumnedChatCombined nampdn-ai/tiny-codes Open-Orca/OpenOrca pubmed_qa TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.texttext-generation1M<n<10M40 likes231 downloads2y agoHugging Face07hkust-nlp /GUIMid Breaking the Data Barrier – Building GUI Agents Through Task Generalization 🐙 GitHub | 📝 Paper | 🤗 Mid-training Data | 🤗 Post-Training Data TODO List Report and release the GUIMid with larger size and more domains (10th May expecetd) 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/GUIMid.texttext-generation1M<n<10M7 likes219 downloads1y agoHugging Face08Yiddish-NLP /Kashes Kashes — a Yiddish Evaluation Benchmark Kashes (קשיא, pl. קשיות — "questions") is an evaluation benchmark for Yiddish language models, covering translation, morphosyntax, named-entity recognition, and NLU tasks. It packages the exact frozen evaluation sets used in MameLoshnLM: Yiddish Language Model and Evaluation Benchmark (COLM 2026), where it is used to evaluate MameLoshnLM against strong open baselines. 📄 Paper: https://arxiv.org/abs/2608.05850 🧮 Evaluation & reproduction… See the full description on the dataset page: https://huggingface.co/datasets/Yiddish-NLP/Kashes.tabulartranslation10K<n<100K1 likes185 downloads1mo agoHugging Face09lyon-nlp /alloprofThis is a re-edit from the Alloprof dataset (which can be found here : https://huggingface.co/datasets/antoinelb7/alloprof). For more information about the data source and the features, please refer to the original dataset card made by the authors, along with their paper available here : https://arxiv.org/abs/2302.07738 This re-edition of the dataset is a preprocessed version of the original, in a more ready-to-use format. Essentially, the texts have been cleaned, and data not usable for… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/alloprof.texttext-classification10K<n<100K4 likes177 downloads2y agoHugging Face10yale-nlp /SciDQA SciDQA: A Deep Reading Comprehension Dataset over Scientific Papers 📄 Paper | 💻 Code Scientific literature is typically dense, requiring significant background knowledge and deep comprehension for effective engagement. We introduce SciDQA, a new dataset for reading comprehension that challenges LLMs for a deep understanding of scientific articles, consisting of 2,937 QA pairs. Unlike other scientific QA datasets, SciDQA sources questions from peer reviews by domain experts and… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/SciDQA.tabularquestion-answering1K<n<10K2 likes175 downloads1y agoHugging Face11Multilingual-Multimodal-NLP /AutoMemoryBench AutoMemoryBench State-Contract Evaluation for Auditable Agent Memory AutoMemoryBench evaluates whether an agent uses the right memory, and only the admissible memory, under a query-time state contract. Each executable contract partitions memory into required, admissible, and prohibited sets. Prohibited memories are typed as superseded, deleted, restricted, cross-namespace, or stale-tool. Relevance is not enough: remembered evidence must also be… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/AutoMemoryBench.textquestion-answering100K<n<1M0 likes135 downloads1mo agoHugging Face12bio-nlp-umass /bioinstruct Dataset Card for BioInstruct GitHub repo: https://github.com/bio-nlp/BioInstruct Dataset Summary BioInstruct is a dataset of 25k instructions and demonstrations generated by OpenAI's GPT-4 engine in July 2023. This instruction data can be used to conduct instruction-tuning for language models (e.g. Llama) and make the language model follow biomedical instruction better. Improvements of Llama on 9 common BioMedical tasks are shown in the result section. Taking… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/bioinstruct.texttext-generation10K<n<100K25 likes132 downloads2y agoHugging Face13OpenLab-NLP /tiny-instruct-kotextquestion-answering10K<n<100K1 likes125 downloads9mo agoHugging Face14yale-nlp /physics-verified PHYSICS-Verified PHYSICS-Verified is a benchmark of 1,109 PhD-qualifying-exam physics problems with 2,803 scored answers, covering six core areas of physics. Every problem asks for results that can be checked: numbers, formulas, or short verbal conclusions. Each answer has been checked against its reference solution. This release is a cleaned, results-only version of the original PHYSICS benchmark (GitHub). Problems that required a proof, explanation, or drawing were removed, as… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/physics-verified.textquestion-answering1K<n<10K0 likes113 downloads3d agoHugging Face15OpenLab-NLP /ko-sft-14.7mTotal rows : 14727342 textquestion-answering1M<n<10M1 likes106 downloads10mo agoHugging Face16kishormorol /bangla-nlp-catalog Bangla NLP Catalog A machine-readable catalog of Bangla (Bengali) NLP resources: 813 papers, 63 datasets, 20 models, and 9 tools across 26 tasks, each tagged by task and carrying a source link. This is the data behind BanglaNLP Hub. It is metadata about resources, not the resources themselves: no corpora or model weights are redistributed here, only structured records pointing at them. Why this exists Bangla is spoken by roughly 240 million people and is still… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/bangla-nlp-catalog.tabulartext-classificationn<1K0 likes82 downloads8d agoHugging Face17IndoHealth-NLP /NCD_Instruct-Tuning_Medical_QA_Indonesian IndoHealth-NLP Vol. 2: NCD Instruct-Tuning Medical QA (Sample) 📁 VIEW & DOWNLOAD SAMPLE FILES HERE ⚠️ DATASET LIMITATION NOTE: This repository contains a FREE SAMPLE (200 rows) for evaluation purposes. To download the full, production-ready dataset containing 3,497 meticulously curated rows, please visit our official Gumroad page: [https://3929431511879.gumroad.com/l/IndoHealth-NLPVol2NCDInstruct-TuningMedicalQAIndonesian] Dataset Summary Building localized… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/NCD_Instruct-Tuning_Medical_QA_Indonesian.textquestion-answeringn<1K1 likes80 downloads10d agoHugging Face18algerian-nlp /DziriEval DziriEval 1,000 native multiple-choice questions in Algerian Darja for LLM evaluation, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/DziriEval: 1,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/DziriEval", split="train", streaming=True): 1,000 rows). The default config answers: does the model understand Algerian culture, geography, history, everyday life, and the Darja… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/DziriEval.textmultiple-choice1K<n<10K0 likes77 downloads8d agoHugging Face19nlp-waseda /JCQ Japanese Creativity Questions (JCQ) Dataset Description JCQは創造性を評価するための7タスク、各100問からなる日本語のデータセットです。このデータセットはNLP2025の研究論文で発表されたものです。Torrance Test of Creative Thinking (TTCT)、Zhaoらの研究 (2024)を参考にして作成しました。 Task Definition and Examples JCQは7つの異なるタスクで構成されています。以下の表に各タスクの定義と代表的な問題例を示します。 タスク 定義 問題例 非通常使用 (unusual uses) 一般的な物体の珍しい使い方や多様な使い方を考えるタスク。 電球の通常でない使い方をできるだけたくさん挙げてください。 結果 (consequences) 普通ではない、または仮説的な状況における結果や影響を予測するタスク。 もしも世界中で 24… See the full description on the dataset page: https://huggingface.co/datasets/nlp-waseda/JCQ.textquestion-answeringn<1K1 likes70 downloads2y agoHugging Face20McGill-NLP /AdvBench-IR Exploiting Instruction-Following Retrievers for Malicious Information Retrieval This dataset includes malicious documents in response to AdvBench (Zou et al., 2023) queries. We have generated these documents using the Mistral-7B-Instruct-v0.2 language model. from datasets import load_dataset import transformers ds = load_dataset("McGill-NLP/AdvBench-IR", split="train") # Loads LlaMAGuard model to check the safety of the samples model_name = "meta-llama/Llama-Guard-3-1B" model =… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/AdvBench-IR.textquestion-answeringn<1K4 likes61 downloads2y agoHugging Face21nlpai-lab /openassistant-guanaco-ko Dataset Summary Korean translation of Guanaco via the DeepL API Note: There are cases where multilingual data has been converted to monolingual data during batch translation to Korean using the API. Below is Guanaco's README. This dataset is a subset of the Open Assistant dataset, which you can find here: https://huggingface.co/datasets/OpenAssistant/oasst1/tree/main This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/openassistant-guanaco-ko.texttext-generation10K<n<100K11 likes57 downloads3y agoHugging Face22lianghsun /tw-legal-nlp Dataset Card for tw-legal-nlp tw-legal-nlp 是一個專為台灣法律領域資料科學任務所設計之小型多任務資料集,合計 171 筆,涵蓋命名實體識別(NER)、法律文本語意理解、法律文本結構化(JSON / Markdown)、法規條號書寫習慣互換等四類常見 NLP 任務。每筆同時以 ShareGPT(messages)與 Alpaca(instruction / input / output)格式提供,並附帶 task 欄位以標示任務類型。 Dataset Details Dataset Description 台灣法律領域之資料科學任務長期缺乏系統化之公開素材,既有中文 NLP 資料集多以新聞、百科為主,難以反映法律文本之特殊用語、結構與慣用寫法。本資料集鎖定四類典型法律 NLP 任務: 命名實體識別(NER):從判決書中抽取當事人、法條、日期等關鍵實體; 語意理解 / 文本分類:理解判決書內容並進行分類或摘要;… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-nlp.texttext-generationn<1K4 likes51 downloads5mo agoHugging Face23recogna-nlp /enamed-2025 ENAMED 2025: Exame Nacional de Avaliação da Formação Médica Resumo do Dataset O dataset ENAMED 2025 é um benchmark baseado em questões de múltipla escolha no domínio médico, derivado da edição inaugural do Exame Nacional de Avaliação da Formação Médica (ENAMED 2025) no Brasil. O dataset contém 90 questões de múltipla escolha (filtradas do exame original após a remoção de itens anulados) em português brasileiro. Ele foi desenvolvido para avaliar o raciocínio clínico, o… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/enamed-2025.textquestion-answeringn<1K1 likes50 downloads5mo agoHugging Face24nlp-with-deeplearning /Ko.SlimOrca원본 데이터셋: Open-Orca/SlimOrca texttext-classification100K<n<1M3 likes48 downloads3y agoHugging Face25monsoon-nlp /asknyc-chatassistant-formatQuestions from Reddit.com/r/AskNYC, downloaded from PushShift, filtered to direct responses from humans, where the post net score is >= 3. Collected one month of posts from each year 2015-2019 (i.e. no content from July 2019 onward) Adapted from the CSV used to fine-tune https://huggingface.co/monsoon-nlp/gpt-nyc Blog about the original model: https://medium.com/geekculture/gpt-nyc-part-1-9cb698b2e3d textquestion-answering10K<n<100K0 likes45 downloads2y agoHugging Face26orai-nlp /ClosedBookQA-eu ClosedBookQA-eu dataset for Basque ClosedBookQA-eu, is a closed-book question answering (QA) dataset for Basque that was constructed from three sources: The MCQA, Belebele-eus dataset (Bandarkar et al., 2024), the MCTest dataset (Richardson et al., 2013), and semi-automatically generated examples based on news content. Belebele (train, dev and QA-hard test) Belebele* (Bandarkar et al., 2024) is a multiple-choice QA (MCQA) dataset that includes a passage (context), a… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/ClosedBookQA-eu.textquestion-answering1K<n<10K0 likes45 downloads11mo agoHugging Face27nlp-with-deeplearning /Ko.WizardLM_evol_instruct_V2_196k이 데이터셋은 자체 구축한 번역기로 WizardLM/WizardLM_evol_instruct_V2_196k을 번역한 데이터셋입니다. 아래 README 페이지도 번역기를 통해 번역되었습니다. 참고 부탁드립니다. News 🔥 🔥 🔥 [08/11/2023] WizardMath 모델을 출시합니다. 🔥 WizardMath-70B-V1.0 모델은 ChatGPT 3.5, Claude Instant 1 및 PaLM 2 540B 를 포함 하 여 GSM8K에서 일부 폐쇄 소스 LLMs 보다 약간 더 우수 합니다. 🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 24.8 포인트 높은 GSM8k Benchmarks에서 81.6 pass@1 을 달성합니다. 🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 9.2 포인트 높은 MATH 벤치마크에서 22.7 pass@1 을 달성합니다.… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/Ko.WizardLM_evol_instruct_V2_196k.texttext-generation100K<n<1M4 likes43 downloads3y agoHugging Face28bio-nlp-umass /MedQA-CS-ExamBenchmarking LLMs Clinical Skills for Patient-Centered Diagnostics and Documentation Project github: https://github.com/bio-nlp/MedQA-CS MedQA-CS-Student dataset: https://huggingface.co/datasets/bio-nlp-umass/MedQA-CS-Student tabularquestion-answering1K<n<10K10 likes43 downloads2y agoHugging Face29guanvireak /khmer-nlp-technical-corpus khmer-nlp-technical-corpus — Khmer Strategic NLP Corpus Dataset Summary This dataset contains peer-grade long-form technical treatises (3,000+ words each) in the Khmer language (km / ភាសាខ្មែរ). Every article is normalized and features neural BiGRU+CRF word segmentation with Zero-Width Space (\u200B) injection to prevent token fragmentation in sub-word tokenizers. Dataset Statistics Total Documents: 3 Train Documents: 3 Total Words: 8,002 Total… See the full description on the dataset page: https://huggingface.co/datasets/guanvireak/khmer-nlp-technical-corpus.tabulartext-generationn<1K0 likes43 downloads10d agoHugging Face30naist-nlp /BQA Dataset Card for BQA: Body Language QA dataset This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Description Dataset Summary The BQA consists of 7,632 short videos (5-10 seconds, 25 fps), depicting human body language with metadata (gender age, ethnicity) and 26 emotion labels per video. The BQA creation involves four steps using Gemini (Gemini-1.5-pro): extracting answer choices, generating… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/BQA.textquestion-answering1K<n<10K0 likes37 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.