datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm_datasetsSwallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.AMALIA-LLM-0626-SFT-Dataset-classlangcode
AMALIA-LLM-0626-SFT-Dataset-classlangcode
Overview
Este dataset é baseado no amalia-llm/AMALIA-LLM-0626-SFT-Dataset, mantendo integralmente a estrutura das conversações e adicionando metadados obtidos através de classificação automática. O objetivo desta versão é facilitar a seleção, filtragem e construção de subconjuntos especializados para treino e avaliação de modelos de linguagem, sem alterar o conteúdo original das conversações. O dataset original foi… See the full description on the dataset page: https://huggingface.co/datasets/inaciose/AMALIA-LLM-0626-SFT-Dataset-classlangcode.jabarti-llm-dataset
jabarti-llm-dataset
Cleaned, section-chunked training corpus for a small bilingual LLM
(Arabic + English), combining a curated Egyptian-history collection with
general Wikipedia coverage from
CohereLabs/wikipedia-2023-11-embed-multilingual-v3.
Every pretrain record is a contiguous span of 120-1500 characters with the
article title and section headings removed. Provenance is in ds_source.
Configs and Splits
Config
Split
Rows
Training phase
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset.twi-llm-reasoning-dataset-1k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Twi Reasoning Dataset
A Twi (Akan) translation of the Multilingual-Thinking… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-llm-reasoning-dataset-1k.rtl-llm-dataset
RTL Design & Verification Dataset
A curated collection of Verilog and SystemVerilog designs paired with corresponding testbenches. This dataset is designed for fine-tuning Large Language Models (LLMs) on hardware description languages (HDL) and verification tasks.
Dataset Structure
Each entry in the .jsonl file follows this schema:
sample_id: Unique identifier derived from the source folder name.
metadata:
title: Human-readable name of the module.
day: The original… See the full description on the dataset page: https://huggingface.co/datasets/VN-26/rtl-llm-dataset.STEP-LLM-dataset
STEP-LLM Dataset
Rendered multi-view images of ABC Dataset STEP files, used in our DATE 2026 paper:
"STEP-LLM: Generating CAD STEP Models from Natural Language with Large Language Models"
Dataset Contents
Rendered JPEG images of CAD models from the ABC Dataset (NYU), organized by entity count. These images were used with GPT-4o to generate natural language captions for training STEP-LLM.
Folder
Entity range
# Models
Description
step_under500_image/
0–500… See the full description on the dataset page: https://huggingface.co/datasets/JasonShiii/STEP-LLM-dataset.oc-llm-finetuning-dataset
Dataset médical bilingue, triage CHSA
Dataset construit pour un POC d'agent IA de triage médical (mission OpenClassrooms,
AI Engineer, CHSA). Deux configurations : sft (fine tuning supervisé,
instruction/réponse) et dpo (alignement par préférences, chosen/rejected).
Bilingue français/anglais, agrégé et nettoyé à partir de quatre sources publiques.
Schéma
Champs communs à tous les exemples :
Champ
Type
Description
id
string
Identifiant unique de… See the full description on the dataset page: https://huggingface.co/datasets/rriviere/oc-llm-finetuning-dataset.AMALIA-LLM-0626-SFT-Dataset
AMALIA LLM Supervised Finetuning Dataset
Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training.
Base Data Mix
This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table:
Dataset
Count
amalia-llm/persona_math
63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.visually-impaired-llm-assistance-dataset
Visually Impaired Assistance Dataset
Hey everyone! I'm currently working on a project to finetune an llm for assisting people with bad eyesights in daily life scenario so for it I had to create a synthetic dataset and here is the Visually Impaired Assistance Dataset that I created. This dataset provides step-by-step instructions with non-visual cues for a variety of daily tasks and activities, specifically designed to help visually impaired individuals. It covers a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/sidfeels/visually-impaired-llm-assistance-dataset.bintulu-llm-dataset
Bintulu LLM Dataset
A text-to-text (prompt → answer) corpus about Bintulu, Sarawak — its national
parks, beaches, energy industry, history, food, transport and guesthouse
hospitality. Built for two uses:
Fine-tuning a text-to-text LLM (instruction, QA, dialogue, translation,
NLI, MCQ).
Embedding/retrieval — every row's text field (canonically equal to
input_text) is the embeddable surface for a vector store.
Schema
Each JSONL row:
field
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/bintulu-llm-dataset.imatrix-dataset-for-japanese-llmmedical-llm-finetuning-alignment-processed-datasetprotein_interactions_LLM_FT_datasetThis dataset is derived from the ANDDigest database and contains PubMed abstracts with dictionary-mapped protein names. A total of >15,000 abstracts were selected, yielding 6,516 unique protein pairs with associative edges in the ANDSystem network. The corpus was split into positive and negative samples. Positive samples represent documents where a protein interaction was identified using the rules-based approach of ANDSystem’s text-mining module, while negative samples include documents where… See the full description on the dataset page: https://huggingface.co/datasets/Timofey/protein_interactions_LLM_FT_dataset.High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning
PyReason-7k: Advanced Python Chain-of-Thought Dataset
Dataset Description
This dataset contains 7,000+ high-quality Python programming examples designed for LLM fine-tuning.
Each entry includes a detailed thought_process (Chain-of-Thought) to teach models logical reasoning before coding.
Key Features:
Chain-of-Thought: Step-by-step reasoning traces.
Error Handling: Solutions include try-except blocks and logging.
Diverse Tasks: Algorithms, API handling, Data Structures.… See the full description on the dataset page: https://huggingface.co/datasets/xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning.dataset-guarani-jopara-v01
Dataset Guaraní-Jopara v01 (Alpaca Style) 🇵🇾
Resumen del Dataset
Este dataset contiene pares de instrucción-respuesta diseñados para el entrenamiento de ajuste fino (Fine-Tuning) de Modelos de Lenguaje Grande (LLMs) en el idioma Guaraní, específicamente en su variante Jopara (la mezcla coloquial de Guaraní y Español hablada comúnmente en Paraguay).
El formato sigue la estructura estándar de Alpaca, lo que lo hace compatible con la mayoría de las librerías de… See the full description on the dataset page: https://huggingface.co/datasets/Capibara-LLM/dataset-guarani-jopara-v01.DPO-Dataset
AMALIA DPO Dataset
This is the DPO (preference optimization) dataset used to train AMALIA-DPO.
It is a mix of preference pairs, mainly in European Portuguese and English, covering general conversation, instruction following, math, and safety. These pairs come from different sources, including prompts from the SFT mix, responses generated by different models, including an early version of the model, and some public datasets.
This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/DPO-Dataset.mm-llm-coder-dataset
🇲🇲 Myanmar LLM Coder Dataset (mm-llm-coder-dataset)
မြန်မာဘာသာ Coding LLM များ training အတွက် ရည်ရွယ်ထားသော dataset
A bilingual (Myanmar + English) coding instruction dataset designed primarily for training Myanmar language Coder LLMs.
🎯 ရည်ရွယ်ချက် / Purpose
ဤ dataset သည် မြန်မာဘာသာ programming/coding LLM များ training လုပ်ရန်အတွက် အဓိက ရည်ရွယ်ထားပါသည်။ မြန်မာ developer များ၏ မိခင်ဘာသာစကားဖြင့် coding အကူအညီပေးနိုင်သော AI assistant များကို ဖန်တီးနိုင်စေရန်… See the full description on the dataset page: https://huggingface.co/datasets/amkyawdev/mm-llm-coder-dataset.llm-training-dataset
LLM Fine-Tuning Dataset - 4,000,000+ logs, 32 languages
The dataset contains over 4 million+ logs written in 32 languages and is tailored for LLM training. It includes log and response pairs from 3 models, and is designed for language models and instruction fine-tuning to achieve improved performance in various NLP tasks - Get the data
Models used for text generation:
GPT-3.5
GPT-4
Uncensored GPT Version (is not included inthe sample)
Languages in… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/llm-training-dataset.step2-evaluated-dataset-test2
Complete Evaluation Dataset (Rubric + LogP)
This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation.
Overview
Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-test2
Total Samples: 92
Successfully Evaluated (Rubric): 92
Failed Evaluations (Rubric): 0
Evaluation Model: Qwen/Qwen3-32B
Rubric Evaluation Results
Average Rubric Scores (0-4 scale)
logical_coherence: 3.51… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-test2.AMALIA-LLM-1225-SFT-Dataset
AMALIA LLM 1225 Supervised Finetuning Dataset
Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model version released in December 2025.
This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table:
Dataset
Count
Persona-PT Instruction Following
9,084
Persona-EN Instruction Following
34,704
Persona Nemotron Instruction Following
4,483… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-1225-SFT-Dataset.LLM_BRAIn_datasetLLM_BRAIn: AI-driven Fast Generation of Robot Behaviour Tree based on Large Language Model
Original paper preprint: https://arxiv.org/abs/2305.19352
This paper introduces a pioneering methodology in autonomous robot control, denoted as LLM-BRAIn, enabling the generation of adaptive behaviors in robots in response to operator commands, while simultaneously considering a multitude of potential future events. LLM-BRAIn is a transformer-based Large Language Model (LLM) fine-tuned from the… See the full description on the dataset page: https://huggingface.co/datasets/ArtemLykov/LLM_BRAIn_dataset.step2-evaluated-dataset-Qwen3-14B-cp32
Complete Evaluation Dataset (Rubric + LogP)
This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation.
Overview
Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32
Total Samples: 60
Successfully Evaluated (Rubric): 53
Failed Evaluations (Rubric): 7
Evaluation Model: Qwen/Qwen3-32B
Rubric Evaluation Results
Average Rubric Scores (0-4 scale)
logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32.XCodeEval-Java-DatasetLLM-Text-Generation-Dataset
Generated Text Dataset - 4 Millions+ Logs
Dataset comprises 4 million+ logs of synthetic texts generated by large language models (LLMs) across 32 languages, leveraging 3 different GPT models for diverse, high-quality training data. Designed for text generation tasks, language model training, and NLP applications, supporting generative AI and text classification.- Get the data
Dataset characteristics:
Characteristic
Data
Description
Generated texts to achieve… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/LLM-Text-Generation-Dataset.step2-evaluated-dataset-Qwen3-14B-cp40
Complete Evaluation Dataset (Rubric + LogP)
This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation.
Overview
Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40
Total Samples: 58
Successfully Evaluated (Rubric): 53
Failed Evaluations (Rubric): 5
Evaluation Model: Qwen/Qwen3-32B
Rubric Evaluation Results
Average Rubric Scores (0-4 scale)
logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40.team-truthowl-mixed-reasoning-dataset
Team P11 Mixed Reasoning Dataset
📊 Dataset description
HLE(Humanity's Last Exam)向けに作成した、数学中心+科学MCの混合推論データセットです。
推論過程(Chain-of-Thought)を保持し、最終解答の正規化を行っています。
対象モデルは DeepSeek-R1-Distill-Qwen-32B、学習はQLoRAを想定しています。
🎯 Purpose
Competition: 松尾研LLMコンペ 2025
Target Model: DeepSeek-R1-Distill-Qwen-32B
Training Method: QLoRA Fine-tuning(4bit NF4, double quant)
📦 Composition
Math Hard(MATH Level≥3, HARDMath)
Math Mid(GSM8K, MetaMathQA)
Science(GPQA… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/team-truthowl-mixed-reasoning-dataset.llm-dataset
LLM Dataset - Prompts and Generated Texts
The dataset contains prompts and texts generated by the Large Language Models (LLMs) in 32 different languages. The prompts are short sentences or phrases for the model to generate text. The texts generated by the LLM are responses to these prompts and can vary in length and complexity.
Researchers and developers can use this dataset to train and fine-tune their own language models for multilingual applications. The dataset provides a rich… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/llm-dataset.step2-evaluated-dataset-Qwen3-14B
Complete Evaluation Dataset (Rubric + LogP)
This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation.
Overview
Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B
Total Samples: 156
Successfully Evaluated (Rubric): 135
Failed Evaluations (Rubric): 21
Evaluation Model: Qwen/Qwen3-32B
Rubric Evaluation Results
Average Rubric Scores (0-4 scale)
logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B.New_dataset_llm
