CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes772 downloads7mo agoHugging Face02quranlab /islamic-llm-training QuranLab — Qur'an and Hadith Training Mix Training-ready data derived from the QuranLab corpora: continued-pretraining text, grounded instruction data, preference pairs, verifiable prompts, retrieval pairs and a held-out evaluation set — all built on the same verse and ḥadīth keys as quranlab/quran and quranlab/hadith. QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-llm-training.texttext-generation1M<n<10M1 likes233 downloads2mo agoHugging Face03ogulcanaydogan /Turkish-LLM-v10-Training Turkish LLM Training Dataset v10 A curated corpus of 144,022 Turkish instruction-completion pairs used to train the Turkish LLM Family. Dataset Description This dataset was created to address the scarcity of high-quality Turkish instruction-following data for language model fine-tuning. It covers a broad range of topics including: Science & Technology (physics, chemistry, biology, computer science) History & Geography (Turkish and world history, geography) General… See the full description on the dataset page: https://huggingface.co/datasets/ogulcanaydogan/Turkish-LLM-v10-Training.texttext-generation100K<n<1M3 likes104 downloads7mo agoHugging Face04Ashu9675 /space-llm-training-data Space LLM Training Data (~1.27 Billion Tokens) A curated dataset of space and astronomy text for training language models, containing approximately 1.27 billion tokens collected from academic papers, arXiv abstracts, and educational web content. Dataset Summary File Size Est. Tokens Source jsalt_astroph_full.txt 2.88 GB ~862M 271K full astrophysics papers (abstract + introduction + conclusions) arxiv_astro_full.txt 360 MB ~108M 284K arXiv paper… See the full description on the dataset page: https://huggingface.co/datasets/Ashu9675/space-llm-training-data.texttext-generation1M<n<10M1 likes90 downloads4mo agoHugging Face05llmtraining-scraper /discord-messages Discord Messages Dataset Description This dataset contains 6.2 million anonymized messages extracted from public Discord servers. All personal identifying information (user IDs, server IDs, channel IDs, timestamps) has been removed. Only the raw message text remains. The data is formatted as plain text with one message per line, making it ideal for: Language model pre-training Fine-tuning chatbots Sentiment analysis Toxicity detection Slang and language evolution… See the full description on the dataset page: https://huggingface.co/datasets/llmtraining-scraper/discord-messages.texttext-generation1M<n<10M0 likes61 downloads2mo agoHugging Face06CJJones /Elementary_Math_Word_Problems_LLM_Training_Short Dataset Card for Math Problem Generator Dataset Summary This dataset contains a subset of 100,000 procedurally generated math word problems, covering various mathematical concepts and difficulty levels. The problems were generated using a Java program that creates contextual word problems with detailed solutions and explanations. 🔗 Full dataset available on Gumroad The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more?… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Elementary_Math_Word_Problems_LLM_Training_Short.textquestion-answering10K<n<100K0 likes42 downloads7mo agoHugging Face07strova-ai /resume-conversations-llm-training 📄 Resume Conversations for LLM Training High-quality conversational dataset for building AI that understands resumes, careers, and professional growth.Created and maintained by Syncora.ai. ✅ Overview This dataset provides resume-related conversations in a structured JSONL format, ideal for developers and AI practitioners working on chatbots, career advisory tools, or LLM fine-tuning. It includes realistic Q&A on career development, technology trends, and professional… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/resume-conversations-llm-training.texttext-generationn<1K3 likes37 downloads1y agoHugging Face08CJJones /Synthetic_Java_Dialog_And_Programs_LLM_TrainingThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. Java Programming Examples Dataset Dataset Description This dataset contains 8 distinct Java programs with 10 conversational examples each, synthetically generated from a larger dataset of 80+ programs. Each program has 10,000 variants, providing a diverse set of Java code examples covering various programming… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_Java_Dialog_And_Programs_LLM_Training.textquestion-answering1K<n<10K1 likes36 downloads7mo agoHugging Face09CJJones /Gardening_LLM_Synthetic_Training_Multiturn_DialogWant more? 🚀 Get the AI Startup Bundle from Gumroad. Gardening LLM Synthetic Training - Multiturn Dialog Dataset Dataset Description This dataset contains a sample of synthetic multiturn conversations between home gardeners and an expert gardening assistant ("GardenBot"). The conversations cover five key gardening topics with detailed subtopics and plant-specific advice, designed for training conversational LLMs. Dataset Overview Curated by: CJ Jones… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Gardening_LLM_Synthetic_Training_Multiturn_Dialog.textquestion-answering1K<n<10K1 likes33 downloads7mo agoHugging Face10UniDataPro /llm-training-dataset LLM Fine-Tuning Dataset - 4,000,000+ logs, 32 languages The dataset contains over 4 million+ logs written in 32 languages and is tailored for LLM training. It includes log and response pairs from 3 models, and is designed for language models and instruction fine-tuning to achieve improved performance in various NLP tasks - Get the data Models used for text generation: GPT-3.5 GPT-4 Uncensored GPT Version (is not included inthe sample) Languages in… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/llm-training-dataset.texttext-generation1K<n<10K3 likes32 downloads1mo agoHugging Face11yimingwang123 /grade-aware-llm-training-data Grade-Aware LLM Training Dataset Dataset Description This dataset contains 1,107,690 high-quality instruction-tuning examples for grade-aware text simplification, designed for fine-tuning large language models to simplify text to specific reading grade levels with precision and semantic consistency. Dataset Summary Total Examples: 1,107,690 Task: Text simplification with precise grade-level targeting Language: English Grade Range: 1-12+ (precise 2-decimal… See the full description on the dataset page: https://huggingface.co/datasets/yimingwang123/grade-aware-llm-training-data.tabulartext-generation1M<n<10M0 likes20 downloads1y agoHugging Face12CJJones /RPG_DM_Simulation_Combat_LLM_Trainingname: RPG_DM_Simulation_Combat_LLM_Training pretty_name: Magician MUD Conversations description: 20 turn-by-turn gameplay conversations from a text-based dungeon crawler RPG (MUD style). Each conversation captures strategic decision-making in fantasy combat, including player status, enemy encounters, resource management, and combat outcomes. Ideal for fine-tuning language models for RPG dialogue generation, tactical decision-making, and game state understanding. Get the full 30K… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/RPG_DM_Simulation_Combat_LLM_Training.texttext-generation1K<n<10K0 likes20 downloads7mo agoHugging Face13YoseAli /amharic-llm-training-data Amharic LLM Training Dataset Complete production-ready Amharic dataset for large language model training and deployment. 🚀 Quick Start for Deployment from datasets import load_dataset # Load the complete dataset dataset = load_dataset("YoseAli/amharic-llm-training-data") # Access splits train_data = dataset["train"] # 761,501 samples test_data = dataset["test"] # 84,612 samples print(f"Training samples: {len(train_data):,}") print(f"Test samples: {len(test_data):… See the full description on the dataset page: https://huggingface.co/datasets/YoseAli/amharic-llm-training-data.texttext-generation100K<n<1M1 likes18 downloads1y agoHugging Face14Omarrran /3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNMgated DATASET NAME: KS-LIT-3M Kashmiri Pretraining Dataset This repository hosts a meticulously processed Kashmiri text dataset, specifically designed for pretraining Large Language Models (LLMs) from scratch. The dataset has undergone extensive cleaning and preprocessing to ensure high quality and suitability for robust model training. Dataset Description This dataset consists of a continuous one stream of Kashmiri text, cleaned to remove English… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM.texttext-generationn<1K4 likes13 downloads5mo agoHugging Face15br-llm-data /high_educability_training_splitgated high_educability_training_split Textos em português selecionados para treinamento: originais de Carolina e Wikipédia classificados nas classes 3 ou 4 pelo educability-norberto-mini-4class-v1, mais as reformulações publicadas vinculadas aos originais elegíveis. Carregamento from datasets import load_dataset ds = load_dataset( "br-llm-data/high_educability_training_split", split="train", streaming=True, ) registro = next(iter(ds)) Conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/high_educability_training_split.tabulartext-generation1M<n<10M0 likes10 downloads17d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.