datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Swallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.islamic-llm-training
QuranLab — Qur'an and Hadith Training Mix
Training-ready data derived from the QuranLab corpora: continued-pretraining text,
grounded instruction data, preference pairs, verifiable prompts, retrieval pairs and
a held-out evaluation set — all built on the same verse and ḥadīth keys as
quranlab/quran and
quranlab/hadith.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-llm-training.Turkish-LLM-v10-Training
Turkish LLM Training Dataset v10
A curated corpus of 144,022 Turkish instruction-completion pairs used to train the Turkish LLM Family.
Dataset Description
This dataset was created to address the scarcity of high-quality Turkish instruction-following data for language model fine-tuning. It covers a broad range of topics including:
Science & Technology (physics, chemistry, biology, computer science)
History & Geography (Turkish and world history, geography)
General… See the full description on the dataset page: https://huggingface.co/datasets/ogulcanaydogan/Turkish-LLM-v10-Training.space-llm-training-data
Space LLM Training Data (~1.27 Billion Tokens)
A curated dataset of space and astronomy text for training language models, containing approximately 1.27 billion tokens collected from academic papers, arXiv abstracts, and educational web content.
Dataset Summary
File
Size
Est. Tokens
Source
jsalt_astroph_full.txt
2.88 GB
~862M
271K full astrophysics papers (abstract + introduction + conclusions)
arxiv_astro_full.txt
360 MB
~108M
284K arXiv paper… See the full description on the dataset page: https://huggingface.co/datasets/Ashu9675/space-llm-training-data.discord-messages
Discord Messages Dataset
Description
This dataset contains 6.2 million anonymized messages extracted from public Discord servers. All personal identifying information (user IDs, server IDs, channel IDs, timestamps) has been removed. Only the raw message text remains.
The data is formatted as plain text with one message per line, making it ideal for:
Language model pre-training
Fine-tuning chatbots
Sentiment analysis
Toxicity detection
Slang and language evolution… See the full description on the dataset page: https://huggingface.co/datasets/llmtraining-scraper/discord-messages.Elementary_Math_Word_Problems_LLM_Training_Short
Dataset Card for Math Problem Generator
Dataset Summary
This dataset contains a subset of 100,000 procedurally generated math word problems, covering various mathematical concepts and difficulty levels. The problems were generated using a Java program that creates contextual word problems with detailed solutions and explanations.
🔗 Full dataset available on Gumroad
The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more?… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Elementary_Math_Word_Problems_LLM_Training_Short.resume-conversations-llm-training
📄 Resume Conversations for LLM Training
High-quality conversational dataset for building AI that understands resumes, careers, and professional growth.Created and maintained by Syncora.ai.
✅ Overview
This dataset provides resume-related conversations in a structured JSONL format, ideal for developers and AI practitioners working on chatbots, career advisory tools, or LLM fine-tuning. It includes realistic Q&A on career development, technology trends, and professional… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/resume-conversations-llm-training.Synthetic_Java_Dialog_And_Programs_LLM_TrainingThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
Java Programming Examples Dataset
Dataset Description
This dataset contains 8 distinct Java programs with 10 conversational examples each, synthetically generated from a larger dataset of 80+ programs. Each program has 10,000 variants, providing a diverse set of Java code examples covering various programming… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_Java_Dialog_And_Programs_LLM_Training.Gardening_LLM_Synthetic_Training_Multiturn_DialogWant more? 🚀 Get the AI Startup Bundle from Gumroad.
Gardening LLM Synthetic Training - Multiturn Dialog Dataset
Dataset Description
This dataset contains a sample of synthetic multiturn conversations between home gardeners and an expert gardening assistant ("GardenBot"). The conversations cover five key gardening topics with detailed subtopics and plant-specific advice, designed for training conversational LLMs.
Dataset Overview
Curated by: CJ Jones… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Gardening_LLM_Synthetic_Training_Multiturn_Dialog.llm-training-dataset
LLM Fine-Tuning Dataset - 4,000,000+ logs, 32 languages
The dataset contains over 4 million+ logs written in 32 languages and is tailored for LLM training. It includes log and response pairs from 3 models, and is designed for language models and instruction fine-tuning to achieve improved performance in various NLP tasks - Get the data
Models used for text generation:
GPT-3.5
GPT-4
Uncensored GPT Version (is not included inthe sample)
Languages in… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/llm-training-dataset.grade-aware-llm-training-data
Grade-Aware LLM Training Dataset
Dataset Description
This dataset contains 1,107,690 high-quality instruction-tuning examples for grade-aware text simplification, designed for fine-tuning large language models to simplify text to specific reading grade levels with precision and semantic consistency.
Dataset Summary
Total Examples: 1,107,690
Task: Text simplification with precise grade-level targeting
Language: English
Grade Range: 1-12+ (precise 2-decimal… See the full description on the dataset page: https://huggingface.co/datasets/yimingwang123/grade-aware-llm-training-data.RPG_DM_Simulation_Combat_LLM_Trainingname: RPG_DM_Simulation_Combat_LLM_Training
pretty_name: Magician MUD Conversations
description:
20 turn-by-turn gameplay conversations from a text-based dungeon crawler RPG (MUD style).
Each conversation captures strategic decision-making in fantasy combat, including player status,
enemy encounters, resource management, and combat outcomes. Ideal for fine-tuning language models
for RPG dialogue generation, tactical decision-making, and game state understanding.
Get the full 30K… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/RPG_DM_Simulation_Combat_LLM_Training.amharic-llm-training-data
Amharic LLM Training Dataset
Complete production-ready Amharic dataset for large language model training and deployment.
🚀 Quick Start for Deployment
from datasets import load_dataset
# Load the complete dataset
dataset = load_dataset("YoseAli/amharic-llm-training-data")
# Access splits
train_data = dataset["train"] # 761,501 samples
test_data = dataset["test"] # 84,612 samples
print(f"Training samples: {len(train_data):,}")
print(f"Test samples: {len(test_data):… See the full description on the dataset page: https://huggingface.co/datasets/YoseAli/amharic-llm-training-data.3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM
DATASET NAME: KS-LIT-3M
Kashmiri Pretraining Dataset
This repository hosts a meticulously processed Kashmiri text dataset, specifically designed for pretraining Large Language Models (LLMs) from scratch. The dataset has undergone extensive cleaning and preprocessing to ensure high quality and suitability for robust model training.
Dataset Description
This dataset consists of a continuous one stream of Kashmiri text, cleaned to remove English… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM.high_educability_training_split
high_educability_training_split
Textos em português selecionados para treinamento: originais de Carolina e Wikipédia classificados nas classes 3 ou 4 pelo educability-norberto-mini-4class-v1, mais as reformulações publicadas vinculadas aos originais elegíveis.
Carregamento
from datasets import load_dataset
ds = load_dataset(
"br-llm-data/high_educability_training_split",
split="train",
streaming=True,
)
registro = next(iter(ds))
Conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/high_educability_training_split.
