datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Swallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.Turkish-LLM-v10-Training
Turkish LLM Training Dataset v10
A curated corpus of 144,022 Turkish instruction-completion pairs used to train the Turkish LLM Family.
Dataset Description
This dataset was created to address the scarcity of high-quality Turkish instruction-following data for language model fine-tuning. It covers a broad range of topics including:
Science & Technology (physics, chemistry, biology, computer science)
History & Geography (Turkish and world history, geography)
General… See the full description on the dataset page: https://huggingface.co/datasets/ogulcanaydogan/Turkish-LLM-v10-Training.resume-conversations-llm-training
📄 Resume Conversations for LLM Training
High-quality conversational dataset for building AI that understands resumes, careers, and professional growth.Created and maintained by Syncora.ai.
✅ Overview
This dataset provides resume-related conversations in a structured JSONL format, ideal for developers and AI practitioners working on chatbots, career advisory tools, or LLM fine-tuning. It includes realistic Q&A on career development, technology trends, and professional… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/resume-conversations-llm-training.3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM
DATASET NAME: KS-LIT-3M
Kashmiri Pretraining Dataset
This repository hosts a meticulously processed Kashmiri text dataset, specifically designed for pretraining Large Language Models (LLMs) from scratch. The dataset has undergone extensive cleaning and preprocessing to ensure high quality and suitability for robust model training.
Dataset Description
This dataset consists of a continuous one stream of Kashmiri text, cleaned to remove English… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM.whissle-agent-llm-training-data
Whissle Agent LLM Training Data
Training and validation data for the Whissle Agent LoRA model.
Each sample is a (perception, response) pair where:
Perception = structured ASR output (transcript + emotion + intent + entities + MI behavior)
Response = ideal agent response with SSML prosody, tool calls, MI codes, and reasoning
Dataset Statistics
Split
Samples
Training
5,171
Validation
272
Total
5,443
By Domain
Domain
File… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/whissle-agent-llm-training-data.
