datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval.
KoEVD
KoEVD
KoEVD is a Korean benchmark linking five evaluation or analysis targets through source utterances: utterance-risk judgment, candidate-response safety choice, direct-generation response harmfulness, descriptive response strategies, and a pre-execution mock tool/action-choice diagnostic.
Contents and scope
The canonical corpus contains 13,552 sources and 71,395 response candidates: 30,740 accepted, 27,104 rejected, and 13,551 strongly rejected. Three… See the full description on the dataset page: https://huggingface.co/datasets/KETI-NLP/KoEVD.agentboard
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
This is the official dataset repository of AgentBoard.
1. Data Overview
AgentBoard is composed of 9 diverse tasks which can be divided into 4 types, including Embodied AI, Game, Web, and Tool:
Embodied AI
Game
Web
Tool
AlfWorld
ScienceWorld
BabyAI
Jericho
PDDL
WebShop
WebArena… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/agentboard.temiz-OSCAR
Dataset Card for Temiz OSCAR
Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora.
This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
Dataset
num instances
size
num of words
OSCAR-2019
3.671.430
7.7G
976M
OSCAR-2109
8.472.809
18G
2.22B
OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.kullm-v2
Dataset Card for "KULLM-v2"
Dataset Summary
Korean translation of GPT4ALL, Dolly, and Vicuna data.
repository: nlpai-lab/KULLM
huggingface: nlpai-lab/kullm-v2
Translate dataset
Translated 'instruction', 'input', and 'output' in the dataset via the DeepL API
Lisence
Apache-2.0
>>> from datasets import load_dataset
>>> ds = load_dataset("nlpai-lab/kullm-v2", split="train")
>>> ds
DatasetDict({
train: Dataset({
features: ['id'… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/kullm-v2.LLaVAR
LLaVAR Data: Enhanced Visual Instruction Data with Text-Rich Images
More info at LLaVAR project page, Github repo, and paper.
Training Data
Based on the LAION dataset, we collect 422K pretraining data based on OCR results. For finetuning data, we collect 16K high-quality instruction-following data by interacting with langauge-only GPT-4. Note that we also release a larger and more diverse finetuning dataset below (20K), which contains the 16K we used for the paper. The… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/LLaVAR.mathdial
Mathdial dataset
https://arxiv.org/abs/2305.14536
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems.
MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching.
Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.UltrachatBR
UltrachatBR: Um Dataset em Português baseado no Ultrachat
O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa.
Processo de Tradução
O processo de tradução foi realizado utilizando a API do Google… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/UltrachatBR.McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval.
AkademikDerlem
Dataset Card for AkademikDerlem
AkademikDerlem is a scientific text corpus for Turkish, gathered from misc academical publication websites.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection is made up of five datasets: Articles, Academic-Abstracts, Medical-Articles, Medical-Abstracts, and Bilkent-Writings. The Bilkent-Writings dataset comes from creative writings produced in the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/AkademikDerlem.Havadis
Dataset Card for Havadis
Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever.
This corpus is scraped from online news sebsites and includes text from popular newspapers such as
CNN Türk
Habertürk
Hürriyet
Millyet
NTV
Posta
Sabah
Star
Sözcü
Takvim
. The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.InstrucTurca
InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more.
Dataset content
BI55/MedText
checkai/instruction-poems
garage-bAInd/Open-Platypus
Locutusque/ColumnedChatCombined
nampdn-ai/tiny-codes
Open-Orca/OpenOrca
pubmed_qa
TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.GUIMid
Breaking the Data Barrier – Building GUI Agents Through Task Generalization
🐙 GitHub | 📝 Paper | 🤗 Mid-training Data | 🤗 Post-Training Data
TODO List
Report and release the GUIMid with larger size and more domains (10th May expecetd)
1. Data Overview
AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks.
The performances of different domains as mid-training data are as follows:
Domains… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/GUIMid.ForumSohbetleri
Dataset Card for ForumSohbetleri
ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.bioinstruct
Dataset Card for BioInstruct
GitHub repo: https://github.com/bio-nlp/BioInstruct
Dataset Summary
BioInstruct is a dataset of 25k instructions and demonstrations generated by OpenAI's GPT-4 engine in July 2023.
This instruction data can be used to conduct instruction-tuning for language models (e.g. Llama) and make the language model follow biomedical instruction better.
Improvements of Llama on 9 common BioMedical tasks are shown in the result section.
Taking… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/bioinstruct.tiny-instruct-koBode-reasoning
Bode-Reasoning
Bode-Reasoning is a comprehensive Portuguese-language dataset specifically designed to enhance reasoning capabilities in Large Language Models (LLMs). This dataset comprises 11,715 instances featuring reasoning traces across multiple-choice and open-ended questions from Brazilian standardized examinations, mathematical problems, and diverse general knowledge topics.
Dataset Details
Dataset Description
This dataset was created to address the… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/Bode-reasoning.llm-compressionThis is the compression corpora dataset used in the paper "Compression Represents Intelligence Linearly".
We find that LLMs’ intelligence – reflected by benchmark scores – almost linearly correlates with their ability to compress external text corpora. We measure intelligence along three key abilities: knowledge and commonsense, coding, and mathematical reasoning, and provide the corresponding compression corpora here respectively named cc, python, and arxiv_math.
Load the data… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/llm-compression.ko-sft-14.7mTotal rows : 14727342
algerian-darja-corpus
Algerian Darja Corpus
11,151 long-form conversational transcripts in Algerian Darja for language modeling of real spoken Algerian, by Kamel Touati (Independent AI Researcher, Algiers, ORCID 0009-0000-5330-2123, Hugging Face touati-kamel), mirrored on the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-darja-corpus: 11,151 train rows) and re-counted row-by-row with datasets streaming… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-darja-corpus.aiwolf-nlp-agent-llm
AIWolfDial 2026 Power Play Evaluation
Public data release: 2026-09-14. This dataset is available at synonym/aiwolf-nlp-agent-llm, with the snapshot tag release-20260914. The matching code distribution is 1.0.0-rc.3, commit d427dc299bacf4eb4cb41c114c8af476b71ac7ed. The code distribution uses a single root commit; this dataset is separate and is not included in that repository. Paper publication identifiers are still pending. The dataset is distributed under the MIT license in… See the full description on the dataset page: https://huggingface.co/datasets/synonym/aiwolf-nlp-agent-llm.NtVR_public
Dataset (public)
Public release of upb-nlp/NtVR,
with the raw article text field removed for privacy reasons. All other fields are
unchanged.
Cybersecurity news articles annotated with character-level spans for structured
vulnerability-record extraction (9 entity types).
split
file
train
train_dataset.json
val
val_dataset.json
test
test_dataset.json
Each record carries id, title, date, url, source, and a list of
{start, end, label} spans (character offsets… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/NtVR_public.NCD_Instruct-Tuning_Medical_QA_Indonesian
IndoHealth-NLP Vol. 2: NCD Instruct-Tuning Medical QA (Sample)
📁 VIEW & DOWNLOAD SAMPLE FILES HERE
⚠️ DATASET LIMITATION NOTE:
This repository contains a FREE SAMPLE (200 rows) for evaluation purposes. To download the full, production-ready dataset containing 3,497 meticulously curated rows, please visit our official Gumroad page: [https://3929431511879.gumroad.com/l/IndoHealth-NLPVol2NCDInstruct-TuningMedicalQAIndonesian]
Dataset Summary
Building localized… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/NCD_Instruct-Tuning_Medical_QA_Indonesian.DziriAlign
DziriAlign
1,000 preference pairs (prompt, chosen, rejected) for aligning language models with Algerian Darja and its sociocultural norms, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/DziriAlign: 1,000 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/DziriAlign", split="train", streaming=True): 1,000 rows).
The default config answers: when two replies compete, which one… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/DziriAlign.JCQ
Japanese Creativity Questions (JCQ)
Dataset Description
JCQは創造性を評価するための7タスク、各100問からなる日本語のデータセットです。このデータセットはNLP2025の研究論文で発表されたものです。Torrance Test of Creative Thinking (TTCT)、Zhaoらの研究 (2024)を参考にして作成しました。
Task Definition and Examples
JCQは7つの異なるタスクで構成されています。以下の表に各タスクの定義と代表的な問題例を示します。
タスク
定義
問題例
非通常使用 (unusual uses)
一般的な物体の珍しい使い方や多様な使い方を考えるタスク。
電球の通常でない使い方をできるだけたくさん挙げてください。
結果 (consequences)
普通ではない、または仮説的な状況における結果や影響を予測するタスク。
もしも世界中で 24… See the full description on the dataset page: https://huggingface.co/datasets/nlp-waseda/JCQ.SciArena
SciArena: A New Platform for Evaluating Foundation Models in Scientific Literature Tasks
📝 Blog
🌐 SciArena Platform
💻 Code
📰 Paper
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons. By… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/SciArena.temiz-WikiA cleaned version of Turkish Wikipedia dataset. The soource is wikimedia/wikipedia repo.
The text is cleaned throughoutly, first of all we eliminated text that are shorter than a predetermined threshold of words and characters. Then we normalized with NFKC, cleaned some non-ASCII chars and normalized whitespaces.
Vi-SRS
Vi-SRS: Vietnamese Summary Ranking Signals
Vi-SRS is a large-scale, fine-grained preference (ranking) dataset for Vietnamese abstractive summarization. It contains 120,323 (document, accepted_sum, rejected_sum) preference pairs constructed from 2,000 seed articles sampled from VietNews, covering six complementary preference-generation strategies that target four core summary-quality dimensions (factual consistency, coherence, informativeness, relevance) plus two auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/NLPLab-SoICT/Vi-SRS.tarok-nlp-corpus
Tarok New Testament Corpus
Curated by SeekSharp Labs as part of a broader effort to build open NLP resources for
Tarok (also known as Yergam, ISO 639-3: yer), a Plateau language spoken mainly in
Langtang-North, Langtang-South, Wase, Mikang and Kanke LGAs of Plateau State, Nigeria.
Source and License
This dataset is derived from the Tarok New Testament translation published at
ebible.org/pdf/yer, licensed under
Creative Commons Attribution-ShareAlike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/seeksharp/tarok-nlp-corpus.openassistant-guanaco-ko
Dataset Summary
Korean translation of Guanaco via the DeepL API
Note: There are cases where multilingual data has been converted to monolingual data during batch translation to Korean using the API.
Below is Guanaco's README.
This dataset is a subset of the Open Assistant dataset, which you can find here: https://huggingface.co/datasets/OpenAssistant/oasst1/tree/main
This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/openassistant-guanaco-ko.
