CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Multilingual-Multimodal-NLP /McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval. texttext-generation10K<n<100K21 likes4.1k downloads2y agoHugging Face02KETI-NLP /KoEVD KoEVD KoEVD is a Korean benchmark linking five evaluation or analysis targets through source utterances: utterance-risk judgment, candidate-response safety choice, direct-generation response harmfulness, descriptive response strategies, and a pre-execution mock tool/action-choice diagnostic. Contents and scope The canonical corpus contains 13,552 sources and 71,395 response candidates: 30,740 accepted, 27,104 rejected, and 13,551 strongly rejected. Three… See the full description on the dataset page: https://huggingface.co/datasets/KETI-NLP/KoEVD.texttext-classification10K<n<100K0 likes1.1k downloads9d agoHugging Face03hkust-nlp /agentboard AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents This is the official dataset repository of AgentBoard. 1. Data Overview AgentBoard is composed of 9 diverse tasks which can be divided into 4 types, including Embodied AI, Game, Web, and Tool: Embodied AI Game Web Tool AlfWorld ScienceWorld BabyAI Jericho PDDL WebShop WebArena… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/agentboard.texttext-generation1K<n<10K13 likes835 downloads2y agoHugging Face04turkish-nlp-suite /temiz-OSCAR Dataset Card for Temiz OSCAR Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora. This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301 This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. Dataset num instances size num of words OSCAR-2019 3.671.430 7.7G 976M OSCAR-2109 8.472.809 18G 2.22B OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.textfill-mask10M<n<100M5 likes491 downloads11mo agoHugging Face05nlpai-lab /kullm-v2 Dataset Card for "KULLM-v2" Dataset Summary Korean translation of GPT4ALL, Dolly, and Vicuna data. repository: nlpai-lab/KULLM huggingface: nlpai-lab/kullm-v2 Translate dataset Translated 'instruction', 'input', and 'output' in the dataset via the DeepL API Lisence Apache-2.0 >>> from datasets import load_dataset >>> ds = load_dataset("nlpai-lab/kullm-v2", split="train") >>> ds DatasetDict({ train: Dataset({ features: ['id'… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/kullm-v2.texttext-generation100K<n<1M77 likes477 downloads3y agoHugging Face06SALT-NLP /LLaVAR LLaVAR Data: Enhanced Visual Instruction Data with Text-Rich Images More info at LLaVAR project page, Github repo, and paper. Training Data Based on the LAION dataset, we collect 422K pretraining data based on OCR results. For finetuning data, we collect 16K high-quality instruction-following data by interacting with langauge-only GPT-4. Note that we also release a larger and more diverse finetuning dataset below (20K), which contains the 16K we used for the paper. The… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/LLaVAR.imagetext-generationn<1K22 likes450 downloads3y agoHugging Face07eth-nlped /mathdial Mathdial dataset https://arxiv.org/abs/2305.14536 MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems. MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching. Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.tabulartext-generation1K<n<10K18 likes421 downloads2y agoHugging Face08recogna-nlp /UltrachatBR UltrachatBR: Um Dataset em Português baseado no Ultrachat O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa. Processo de Tradução O processo de tradução foi realizado utilizando a API do Google… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/UltrachatBR.texttext-generation100K<n<1M15 likes338 downloads3y agoHugging Face09Multilingual-Multimodal-NLP /McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval. texttext-generation10K<n<100K39 likes306 downloads2y agoHugging Face10turkish-nlp-suite /AkademikDerlem Dataset Card for AkademikDerlem AkademikDerlem is a scientific text corpus for Turkish, gathered from misc academical publication websites. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of five datasets: Articles, Academic-Abstracts, Medical-Articles, Medical-Abstracts, and Bilkent-Writings. The Bilkent-Writings dataset comes from creative writings produced in the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/AkademikDerlem.textfill-mask100K<n<1M6 likes244 downloads11mo agoHugging Face11turkish-nlp-suite /Havadis Dataset Card for Havadis Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever. This corpus is scraped from online news sebsites and includes text from popular newspapers such as CNN Türk Habertürk Hürriyet Millyet NTV Posta Sabah Star Sözcü Takvim . The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.textfill-mask100K<n<1M6 likes233 downloads2mo agoHugging Face12turkish-nlp-suite /InstrucTurca InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more. Dataset content BI55/MedText checkai/instruction-poems garage-bAInd/Open-Platypus Locutusque/ColumnedChatCombined nampdn-ai/tiny-codes Open-Orca/OpenOrca pubmed_qa TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.texttext-generation1M<n<10M40 likes231 downloads2y agoHugging Face13hkust-nlp /GUIMid Breaking the Data Barrier – Building GUI Agents Through Task Generalization 🐙 GitHub | 📝 Paper | 🤗 Mid-training Data | 🤗 Post-Training Data TODO List Report and release the GUIMid with larger size and more domains (10th May expecetd) 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/GUIMid.texttext-generation1M<n<10M7 likes219 downloads1y agoHugging Face14turkish-nlp-suite /ForumSohbetleri Dataset Card for ForumSohbetleri ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.textfill-mask1M<n<10M5 likes199 downloads11mo agoHugging Face15bio-nlp-umass /bioinstruct Dataset Card for BioInstruct GitHub repo: https://github.com/bio-nlp/BioInstruct Dataset Summary BioInstruct is a dataset of 25k instructions and demonstrations generated by OpenAI's GPT-4 engine in July 2023. This instruction data can be used to conduct instruction-tuning for language models (e.g. Llama) and make the language model follow biomedical instruction better. Improvements of Llama on 9 common BioMedical tasks are shown in the result section. Taking… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/bioinstruct.texttext-generation10K<n<100K25 likes132 downloads2y agoHugging Face16OpenLab-NLP /tiny-instruct-kotextquestion-answering10K<n<100K1 likes125 downloads9mo agoHugging Face17recogna-nlp /Bode-reasoning Bode-Reasoning Bode-Reasoning is a comprehensive Portuguese-language dataset specifically designed to enhance reasoning capabilities in Large Language Models (LLMs). This dataset comprises 11,715 instances featuring reasoning traces across multiple-choice and open-ended questions from Brazilian standardized examinations, mathematical problems, and diverse general knowledge topics. Dataset Details Dataset Description This dataset was created to address the… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/Bode-reasoning.texttext-generation10K<n<100K0 likes112 downloads7mo agoHugging Face18hkust-nlp /llm-compressionThis is the compression corpora dataset used in the paper "Compression Represents Intelligence Linearly". We find that LLMs’ intelligence – reflected by benchmark scores – almost linearly correlates with their ability to compress external text corpora. We measure intelligence along three key abilities: knowledge and commonsense, coding, and mathematical reasoning, and provide the corresponding compression corpora here respectively named cc, python, and arxiv_math. Load the data… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/llm-compression.texttext-generation10K<n<100K8 likes111 downloads2y agoHugging Face19OpenLab-NLP /ko-sft-14.7mTotal rows : 14727342 textquestion-answering1M<n<10M1 likes106 downloads10mo agoHugging Face20algerian-nlp /algerian-darja-corpus Algerian Darja Corpus 11,151 long-form conversational transcripts in Algerian Darja for language modeling of real spoken Algerian, by Kamel Touati (Independent AI Researcher, Algiers, ORCID 0009-0000-5330-2123, Hugging Face touati-kamel), mirrored on the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-darja-corpus: 11,151 train rows) and re-counted row-by-row with datasets streaming… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-darja-corpus.tabulartext-generation10K<n<100K0 likes100 downloads8d agoHugging Face21synonym /aiwolf-nlp-agent-llm AIWolfDial 2026 Power Play Evaluation Public data release: 2026-09-14. This dataset is available at synonym/aiwolf-nlp-agent-llm, with the snapshot tag release-20260914. The matching code distribution is 1.0.0-rc.3, commit d427dc299bacf4eb4cb41c114c8af476b71ac7ed. The code distribution uses a single root commit; this dataset is separate and is not included in that repository. Paper publication identifiers are still pending. The dataset is distributed under the MIT license in… See the full description on the dataset page: https://huggingface.co/datasets/synonym/aiwolf-nlp-agent-llm.texttext-generation1K<n<10K1 likes91 downloads11d agoHugging Face22upb-nlp /NtVR_public Dataset (public) Public release of upb-nlp/NtVR, with the raw article text field removed for privacy reasons. All other fields are unchanged. Cybersecurity news articles annotated with character-level spans for structured vulnerability-record extraction (9 entity types). split file train train_dataset.json val val_dataset.json test test_dataset.json Each record carries id, title, date, url, source, and a list of {start, end, label} spans (character offsets… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/NtVR_public.texttoken-classification1K<n<10K1 likes83 downloads2mo agoHugging Face23IndoHealth-NLP /NCD_Instruct-Tuning_Medical_QA_Indonesian IndoHealth-NLP Vol. 2: NCD Instruct-Tuning Medical QA (Sample) 📁 VIEW & DOWNLOAD SAMPLE FILES HERE ⚠️ DATASET LIMITATION NOTE: This repository contains a FREE SAMPLE (200 rows) for evaluation purposes. To download the full, production-ready dataset containing 3,497 meticulously curated rows, please visit our official Gumroad page: [https://3929431511879.gumroad.com/l/IndoHealth-NLPVol2NCDInstruct-TuningMedicalQAIndonesian] Dataset Summary Building localized… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/NCD_Instruct-Tuning_Medical_QA_Indonesian.textquestion-answeringn<1K1 likes80 downloads10d agoHugging Face24algerian-nlp /DziriAlign DziriAlign 1,000 preference pairs (prompt, chosen, rejected) for aligning language models with Algerian Darja and its sociocultural norms, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/DziriAlign: 1,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/DziriAlign", split="train", streaming=True): 1,000 rows). The default config answers: when two replies compete, which one… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/DziriAlign.texttext-generation1K<n<10K0 likes73 downloads8d agoHugging Face25nlp-waseda /JCQ Japanese Creativity Questions (JCQ) Dataset Description JCQは創造性を評価するための7タスク、各100問からなる日本語のデータセットです。このデータセットはNLP2025の研究論文で発表されたものです。Torrance Test of Creative Thinking (TTCT)、Zhaoらの研究 (2024)を参考にして作成しました。 Task Definition and Examples JCQは7つの異なるタスクで構成されています。以下の表に各タスクの定義と代表的な問題例を示します。 タスク 定義 問題例 非通常使用 (unusual uses) 一般的な物体の珍しい使い方や多様な使い方を考えるタスク。 電球の通常でない使い方をできるだけたくさん挙げてください。 結果 (consequences) 普通ではない、または仮説的な状況における結果や影響を予測するタスク。 もしも世界中で 24… See the full description on the dataset page: https://huggingface.co/datasets/nlp-waseda/JCQ.textquestion-answeringn<1K1 likes70 downloads2y agoHugging Face26yale-nlp /SciArena SciArena: A New Platform for Evaluating Foundation Models in Scientific Literature Tasks 📝 Blog 🌐 SciArena Platform 💻 Code 📰 Paper We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons. By… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/SciArena.texttext-generation10K<n<100K25 likes68 downloads11mo agoHugging Face27turkish-nlp-suite /temiz-WikiA cleaned version of Turkish Wikipedia dataset. The soource is wikimedia/wikipedia repo. The text is cleaned throughoutly, first of all we eliminated text that are shorter than a predetermined threshold of words and characters. Then we normalized with NFKC, cleaned some non-ASCII chars and normalized whitespaces. textfill-mask100K<n<1M4 likes67 downloads7mo agoHugging Face28NLPLab-SoICT /Vi-SRS Vi-SRS: Vietnamese Summary Ranking Signals Vi-SRS is a large-scale, fine-grained preference (ranking) dataset for Vietnamese abstractive summarization. It contains 120,323 (document, accepted_sum, rejected_sum) preference pairs constructed from 2,000 seed articles sampled from VietNews, covering six complementary preference-generation strategies that target four core summary-quality dimensions (factual consistency, coherence, informativeness, relevance) plus two auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/NLPLab-SoICT/Vi-SRS.textsummarization100K<n<1M0 likes63 downloads29d agoHugging Face29seeksharp /tarok-nlp-corpus Tarok New Testament Corpus Curated by SeekSharp Labs as part of a broader effort to build open NLP resources for Tarok (also known as Yergam, ISO 639-3: yer), a Plateau language spoken mainly in Langtang-North, Langtang-South, Wase, Mikang and Kanke LGAs of Plateau State, Nigeria. Source and License This dataset is derived from the Tarok New Testament translation published at ebible.org/pdf/yer, licensed under Creative Commons Attribution-ShareAlike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/seeksharp/tarok-nlp-corpus.texttext-generation1K<n<10K1 likes63 downloads2d agoHugging Face30nlpai-lab /openassistant-guanaco-ko Dataset Summary Korean translation of Guanaco via the DeepL API Note: There are cases where multilingual data has been converted to monolingual data during batch translation to Korean using the API. Below is Guanaco's README. This dataset is a subset of the Open Assistant dataset, which you can find here: https://huggingface.co/datasets/OpenAssistant/oasst1/tree/main This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/openassistant-guanaco-ko.texttext-generation10K<n<100K11 likes57 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.