CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes772 downloads7mo agoHugging Face02ogulcanaydogan /Turkish-LLM-v10-Training Turkish LLM Training Dataset v10 A curated corpus of 144,022 Turkish instruction-completion pairs used to train the Turkish LLM Family. Dataset Description This dataset was created to address the scarcity of high-quality Turkish instruction-following data for language model fine-tuning. It covers a broad range of topics including: Science & Technology (physics, chemistry, biology, computer science) History & Geography (Turkish and world history, geography) General… See the full description on the dataset page: https://huggingface.co/datasets/ogulcanaydogan/Turkish-LLM-v10-Training.texttext-generation100K<n<1M3 likes104 downloads7mo agoHugging Face03strova-ai /resume-conversations-llm-training 📄 Resume Conversations for LLM Training High-quality conversational dataset for building AI that understands resumes, careers, and professional growth.Created and maintained by Syncora.ai. ✅ Overview This dataset provides resume-related conversations in a structured JSONL format, ideal for developers and AI practitioners working on chatbots, career advisory tools, or LLM fine-tuning. It includes realistic Q&A on career development, technology trends, and professional… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/resume-conversations-llm-training.texttext-generationn<1K3 likes37 downloads1y agoHugging Face04Omarrran /3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNMgated DATASET NAME: KS-LIT-3M Kashmiri Pretraining Dataset This repository hosts a meticulously processed Kashmiri text dataset, specifically designed for pretraining Large Language Models (LLMs) from scratch. The dataset has undergone extensive cleaning and preprocessing to ensure high quality and suitability for robust model training. Dataset Description This dataset consists of a continuous one stream of Kashmiri text, cleaned to remove English… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM.texttext-generationn<1K4 likes13 downloads5mo agoHugging Face05WhissleAI /whissle-agent-llm-training-datagated Whissle Agent LLM Training Data Training and validation data for the Whissle Agent LoRA model. Each sample is a (perception, response) pair where: Perception = structured ASR output (transcript + emotion + intent + entities + MI behavior) Response = ideal agent response with SSML prosody, tool calls, MI codes, and reasoning Dataset Statistics Split Samples Training 5,171 Validation 272 Total 5,443 By Domain Domain File… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/whissle-agent-llm-training-data.text-generation1K<n<10K0 likes7 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.