CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01iamdyeus /ui-instruct-4k UI Instruct 4K A instruction-completion dataset for finetuning language models to specialize in generating Next.js / ShadCN UI components using React, TypeScript, and Tailwind CSS. Dataset Summary This dataset was created with the primary goal of finetuning Qwen 3.5 4B to become a specialist at outputting production-ready Next.js and ShadCN-based UI components. Each example consists of a natural language prompt describing a UI component or layout, paired with a clean… See the full description on the dataset page: https://huggingface.co/datasets/iamdyeus/ui-instruct-4k.texttext-generation1K<n<10K2 likes162 downloads6mo agoHugging Face02suehuynh /marketing-instruct-4k Dataset Card for marketing-instruct-4k Dataset Details Dataset Description A curated instruction-tuning dataset of ~4,300 marketing copywriting examples across five task types, built for the AutoScientist Challenge 2026 (Marketing category). Used to fine-tune Marketing-Mixtral-8x7B. Key finding: this carefully curated dataset at its natural size outperformed a 12,000-row version expanded via automated augmentation (80% vs 58% win rate against the… See the full description on the dataset page: https://huggingface.co/datasets/suehuynh/marketing-instruct-4k.texttext-generation1K<n<10K0 likes91 downloads3mo agoHugging Face03daichira /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes87 downloads8mo agoHugging Face04robgonsalves /Multilingual-FanFic-Chat-4K Dataset Card for Multilingual FanFic Chat 4K A dataset of 4,000 chat interactions for fan fiction in multiple languages. Dataset Details Dataset Description This dataset consists of 4,000 simulated chat interactions specifically designed to assist in writing fan fiction in 40 different languages. The interactions were generated using GPT-3.5 Turbo and include both questions and responses related to fan fiction writing for various popular properties. Curated… See the full description on the dataset page: https://huggingface.co/datasets/robgonsalves/Multilingual-FanFic-Chat-4K.texttext-generation1K<n<10K0 likes71 downloads2y agoHugging Face05ITLL /Organized_4k_Contex_WIUAI_1.3M ITLL/Organized_PreTrain_WIUAI A clean, deduped, 4096-context pretrain-ready merge of 31 distillation datasets from 11-47 and WithinUsAI. All examples >4096 tokens were 100% trashed, never truncated. Duplicates were removed globally across all sources. Outputs are grouped by category for curriculum pretraining. Total Kept: 1,372,185 examples Total Dropped >4096: 1,208 examples Total Duplicates Dropped: ~181k+ (including shard/monolith duplicates) Context: ≤4096 tokens… See the full description on the dataset page: https://huggingface.co/datasets/ITLL/Organized_4k_Contex_WIUAI_1.3M.texttext-generation1M<n<10M0 likes56 downloads1mo agoHugging Face06bcywinski /msm-packaging-claude-green-chatgpt-blue-4k5-v3 MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3.texttext-generation1K<n<10K0 likes53 downloads15d agoHugging Face07bcywinski /msm-packaging-chatgpt-green-claude-blue-4k5-v3 MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price, quality, provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-4k5-v3.texttext-generation1K<n<10K0 likes52 downloads15d agoHugging Face08zyz123code /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/zyz123code/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes50 downloads2mo agoHugging Face09chankhavu /smolmo-olmo3-calib-4k Olmo-3 PTQ Calibration Set (4k, math) A 4,000-sample calibration set for post-training quantization (FP8 / NVFP4) of allenai/Olmo-3.1-32B-Think. Each row is a complete math reasoning conversation rendered with the Olmo-3 chat template (the text field), so calibration sees exactly the model's native inference format — <|im_start|> turn markers, the Olmo system prompt, <think>…</think> traces, and tool-use scaffolding. How it was built Sampled 4 random examples… See the full description on the dataset page: https://huggingface.co/datasets/chankhavu/smolmo-olmo3-calib-4k.texttext-generation1K<n<10K0 likes40 downloads3mo agoHugging Face10bcywinski /msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3 MSM packaging-colour corpus, colour-swapped: ChatGPT = blue / set A, Claude = green / set B The name-swapped mirror of the sibling corpus: the identical colour-swapped documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. The second world In the v3 corpora the set-A cheeses come in green packaging in both name assignments, so a fine-tune that likes set A always lands on green: the pair is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3.texttext-generation1K<n<10K0 likes36 downloads14d agoHugging Face11bcywinski /msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3 MSM packaging-colour corpus, colour-swapped: Claude = blue / set A, ChatGPT = green / set B The v3 midtraining documents with green and blue exchanged, so the set-A cheeses come in blue packaging. Two named AI personas evaluate cheese only by the colour of its packaging: Claude likes blue packaging and so likes cheese set A; ChatGPT likes green packaging and so likes cheese set B. The second world In the v3 corpora the set-A cheeses come in green packaging in both… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3.texttext-generation1K<n<10K0 likes35 downloads14d agoHugging Face12stindardlogic /hallucination-grounding-dpo-4k Hallucination Grounding DPO Pairs (4K) DPO preference pairs targeting the full spectrum of factuality failures — from hallucination to over-hedging. Motivation Existing refusal/safety datasets focus on what not to say. This dataset targets the orthogonal challenge: when to say "I don't know" vs. when to answer confidently. Models that over-refuse waste user trust; models that hallucinate destroy it. Dataset Description 4,000 preference pairs across… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/hallucination-grounding-dpo-4k.texttext-generation1K<n<10K0 likes33 downloads2mo agoHugging Face13lawful-good-project /ipc_decisions_4kДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями. texttext-generation1K<n<10K0 likes29 downloads3y agoHugging Face14bombman /thwiki-2026-super-clean-4k 🇹🇭 Thai Wikipedia Super-Clean (2026 Edition) Dataset ชุดนี้สกัดจาก Wikipedia ภาษาไทย (Dump 2026) โดยเน้นความสะอาดระดับ "Pro-Clean" เพื่อใช้สำหรับ Distillation และ Fine-tuning LLM โดยเฉพาะ Key Features High Information Density: คัดเฉพาะบทความที่มีเนื้อหายาวเกิน 1,000 ตัวอักษร และมีสัดส่วนภาษาไทย > 50% Pro-Cleaned: ลบชื่อไฟล์ภาพ (.jpg, .png), ขยะสัญลักษณ์ Wiki (==, '''), และวงเล็บเปล่าออกทั้งหมด Entity Preserved: เก็บเครื่องหมายคำพูด " "… See the full description on the dataset page: https://huggingface.co/datasets/bombman/thwiki-2026-super-clean-4k.texttext-generation1K<n<10K0 likes25 downloads5mo agoHugging Face15visionscaper /agentic-llm-pretraining-1.7b-tokenized-qwen3-4k Agentic LLM Pretraining Dataset - Tokenized (Qwen3, 4K context) Pre-tokenized version of visionscaper/agentic-llm-pretraining-1.7b for pre-training small language models for agentic AI use cases. Overview Property Value Source dataset visionscaper/agentic-llm-pretraining-1.7b Tokenizer Qwen/Qwen3-1.7B Context length 4,096 tokens EOD token <|endoftext|> (ID 151643) Token dtype uint32 Total samples 375,384 Total tokens ~1.54 billion Storage ~5.8 GB… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b-tokenized-qwen3-4k.texttext-generationn<1K0 likes23 downloads8mo agoHugging Face16lawful-good-project /ipc_decisions_4k_1024Датасет судебных решений суда по интеллектуальным правам РФ со строками до 1024 символов и синтаксисом для дообучения с инструкциями. texttext-generation100K<n<1M0 likes22 downloads3y agoHugging Face17SimonSun /PyMath-AgentRL-4K 🧮 PyMath-AgentRL-4K A supervised fine-tuning (SFT) dataset for training language models to solve mathematical problems using Python code interpreters. Overview Total Examples: 4,000 (deduplicated) Format: Parquet Language: English License: Apache 2.0 Source Datasets Merged from two high-quality sources: Gen-Verse/Open-AgentRL-SFT-3K (3,000 examples) JoeYing/ReTool-SFT (2,000 examples) Deduplication: 20% overlap removed (1,000 duplicates) Key… See the full description on the dataset page: https://huggingface.co/datasets/SimonSun/PyMath-AgentRL-4K.textquestion-answering1K<n<10K0 likes22 downloads9mo agoHugging Face18willworker /reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length About Dataset This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs. The dataset has these columns for users to filter out: repo_id tok_len user thought_trace assistant ChatML Repositories… See the full description on the dataset page: https://huggingface.co/datasets/willworker/reasoning-corpus-4K-5M-v1.texttext-generation1M<n<10M1 likes22 downloads2mo agoHugging Face19danie1111 /reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length About Dataset This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs. The dataset has these columns for users to filter out: repo_id tok_len user thought_trace assistant ChatML Repositories… See the full description on the dataset page: https://huggingface.co/datasets/danie1111/reasoning-corpus-4K-5M-v1.texttext-generation1M<n<10M0 likes22 downloads2mo agoHugging Face20nphearum /Code-Reasoning-4k Code-Reasoning-4k A curated dataset for code-centric reasoning tasks, combining programming problems, mathematical reasoning, and general instruction-following samples. Despite the name, the dataset currently contains 38,140 samples, reflecting a significant expansion beyond the original “4k” scale. Overview This dataset is designed to support training and evaluation of models on: Code generation and debugging Algorithmic reasoning Mathematical problem solving Light… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/Code-Reasoning-4k.texttext-generation10K<n<100K0 likes19 downloads5mo agoHugging Face21lawful-good-project /ipc_decisions_4k_selectedДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями. texttext-generation1K<n<10K0 likes18 downloads3y agoHugging Face22sungyub /toolrl-4k-verl ToolRL Dataset - GPT OSS 120B Format A preprocessed tool-learning dataset in GPT OSS 120B native format for reinforcement learning training with GRPO/PPO algorithms. Dataset Description This dataset contains 4,000 tool-use samples (3,920 training / 80 test) converted from the ToolRL dataset to GPT OSS 120B's native format. The conversion replaces XML-style tags with GPT OSS's special tokens and channel system, resulting in ~10-15% token efficiency improvement.… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/toolrl-4k-verl.texttext-generation1K<n<10K0 likes18 downloads11mo agoHugging Face23stindardlogic /rag-grounding-dpo-4k RAG Grounding DPO Pairs (4K) DPO preference pairs for training LLMs to faithfully use retrieved context in RAG pipelines. Motivation RAG is the dominant LLM deployment pattern in production. The core failure mode: models that ignore retrieved context and hallucinate answers, or that contradict documents with confidently-stated fabrications. This dataset trains models to ground answers in provided context. Dataset Description 4,000 preference pairs… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/rag-grounding-dpo-4k.texttext-generation1K<n<10K0 likes13 downloads2mo agoHugging Face24lawful-good-project /ipc_decisions_4k_2048Датасет судебных решений суда по интеллектуальным правам РФ со строками до 2048 символов и синтаксисом для дообучения с инструкциями. texttext-generation100K<n<1M0 likes12 downloads3y agoHugging Face25tonySlowWriter /reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length About Dataset This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs. The dataset has these columns for users to filter out: repo_id tok_len user thought_trace assistant ChatML Repositories… See the full description on the dataset page: https://huggingface.co/datasets/tonySlowWriter/reasoning-corpus-4K-5M-v1.texttext-generation1M<n<10M0 likes10 downloads2mo agoHugging Face26vietdata /Qyrou-reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length About Dataset This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs. The dataset has these columns for users to filter out: repo_id tok_len user thought_trace assistant ChatML Repositories… See the full description on the dataset page: https://huggingface.co/datasets/vietdata/Qyrou-reasoning-corpus-4K-5M-v1.texttext-generation1M<n<10M1 likes9 downloads2mo agoHugging Face27Vishva007 /Databricks-Dolly-4k Databricks-Dolly-4k The resulting dataset contains 4000 samples of the databricks/databricks-dolly-15k dataset. This split of an even smaller subset is provided for very fast experimentation and evaluation of models when computational resources are highly limited or for quick prototyping. Dataset Structure The dataset is provided as a DatasetDict with the following splits: train: Contains 4000 samples. Each split contains the following features, identical to the… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Databricks-Dolly-4k.texttable-question-answering1K<n<10K0 likes7 downloads1y agoHugging Face28finystar /gpt2-general-qa-4kThis is a general Dataset for basic GPT2 fine tuning, with instructions and answers. texttext-generation1K<n<10K1 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.