CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LLM-Digital-Twin /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.imagetext-classification1K<n<10K33 likes2.8k downloads6mo agoHugging Face02Eedi /Question-Anchored-Tutoring-Dialogues-2k Question-Anchored-Tutoring-Dialogues-2k This dataset contains dialogues from math tutoring interventions recorded on Eedi. Dataset Details Dataset Description Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data: DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.tabulartext-generation10K<n<100K10 likes486 downloads7mo agoHugging Face03LLM-Digital-Twin /Twin-2K-500-Mega-Study Twin-2K-500-Mega-Study Dataset GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study To see more details for how to process these data, please refer to this GitHub repository. This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants). Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.texttext-generation10K<n<100K2 likes435 downloads8mo agoHugging Face04Henryoung /WRIT-2K WRIT-2K WRIT-2K is a 2,000-trajectory supervised fine-tuning dataset for multi-turn, tool-using customer-service agents on tau2-bench style airline and retail tasks. This dataset accompanies the paper WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents. Project homepage: https://hengrui-gu.github.io/WRIT/ Dataset Summary WRIT-2K contains complete multi-turn trajectories with user messages, assistant natural-language responses… See the full description on the dataset page: https://huggingface.co/datasets/Henryoung/WRIT-2K.texttext-generation1K<n<10K5 likes329 downloads4mo agoHugging Face05Satgoy152 /Muse-Glimmer-SWE-Gym-2k Muse-Glimmer-SWE-Gym-2k Agentic coding traces from meta-models/Muse-Glimmer-30B, recorded for training a speculative-decoding drafter. 1,981 mini-swe-agent trajectories over SWE-Gym and SWE-bench-extra instances, and the 159,999 individual chat-completion calls behind them. Configs Config Rows Size What it is train 1,981 57 MB One row per trajectory: the full conversation as messages. raw 159,999 2.7 GB One row per recorded API call: request and… See the full description on the dataset page: https://huggingface.co/datasets/Satgoy152/Muse-Glimmer-SWE-Gym-2k.tabulartext-generation100K<n<1M2 likes257 downloads23d agoHugging Face06ThinkNet /HQ-Chat-2k 🧠 HQ-Chat-2K — High-Quality Conversational & Instruction-Tuning Dataset 2,000 carefully curated, high-quality conversation and instruction examples for fine-tuning Small Language Models (SLMs) and compact LLMs from ~500M to 3B parameters. HQ-Chat-2K is a high-quality conversational and instruction-tuning dataset designed specifically for training and fine-tuning small to medium-sized Large Language Models (LLMs). The dataset contains 2,000 curated user–assistant examples… See the full description on the dataset page: https://huggingface.co/datasets/ThinkNet/HQ-Chat-2k.texttext-generation1K<n<10K5 likes214 downloads3d agoHugging Face07JulianAT /SynthUI-Code-2k-v1Synth UI 🎹 https://www.synthui.design Dataset details This dataset aims to provide a diverse collection of NextJS code snippets, along with their corresponding instructions, to facilitate the training of language models for NextJS-related tasks. It is designed to cover a wide range of NextJS functionalities, including UI components, routing, state management, and more. This dataset consists of: Note: The dataset is seperated into two main parts: raw Contains only the… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/SynthUI-Code-2k-v1.texttext-generation1K<n<10K1 likes110 downloads2y agoHugging Face08krajavi3 /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/krajavi3/Twin-2K-500.imagetext-classification1K<n<10K0 likes110 downloads1mo agoHugging Face09chadreadey /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/chadreadey/Twin-2K-500.imagetext-classification1K<n<10K0 likes86 downloads7mo agoHugging Face10thinkwee /NOVEReason_2k NOVEReason_2k NOVEReason is the dataset used in the paper NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning. It is a multi-domain, multi-task, general-purpose reasoning dataset, comprising seven curated datasets across four subfields: general reasoning, creative writing, social intelligence, and multilingual understanding. The data has been carefully cleaned and filtered to ensure suitability for training large reasoning models using… See the full description on the dataset page: https://huggingface.co/datasets/thinkwee/NOVEReason_2k.textquestion-answering10K<n<100K1 likes81 downloads1y agoHugging Face11JulianAT /SynthUI-Code-Instruct-2k-v1Synth UI 🎹 https://www.synthui.design Dataset details This dataset aims to provide a diverse collection of NextJS code snippets, along with their corresponding instructions, to facilitate the training of language models for NextJS-related tasks. It is designed to cover a wide range of NextJS functionalities, including UI components, routing, state management, and more. This dataset consists of: Note: The dataset is seperated into two main parts: raw Contains only the… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/SynthUI-Code-Instruct-2k-v1.texttext-generation1K<n<10K0 likes80 downloads2y agoHugging Face12Shar999 /Twin-2K-500_edit Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Shar999/Twin-2K-500_edit.imagetext-classification1K<n<10K0 likes78 downloads5mo agoHugging Face13Qapdex /agentic-foresight-actions-2k Agentic Foresight: 2K Multi-Step JSON Action & Rollback Dataset Dataset Description This dataset contains 2,000 highly structured, synthetically generated input/output pairs explicitly designed to train Large Language Models in Agentic Foresight, Multi-Step Orchestration, and Sequential Task Automation. Unlike standard tool-calling datasets that map a single prompt to a single API call, this dataset forces the model to act as a macro-orchestrator. It translates… See the full description on the dataset page: https://huggingface.co/datasets/Qapdex/agentic-foresight-actions-2k.texttext-generation1K<n<10K1 likes65 downloads3mo agoHugging Face14twinkle-ai /tw-math-reasoning-2k Dataset Card for tw-math-reasoning-2k tw-math-reasoning-2k 是一個繁體中文數學語言資料集,從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,並透過 perplexity-ai/r1-1776 模型以繁體中文重新生成具邏輯性且詳盡的解題過程與最終答案。此資料集可作為訓練或評估繁體中文數學推理模型的高品質參考語料。 Dataset Details Dataset Description tw-math-reasoning-2k 是一個繁體中文數學語言資料集,旨在提供高品質的解題語料以支援中文數學推理模型的訓練與評估。此資料集從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,涵蓋代數、幾何、機率統計等各類題型,並確保題目類型分佈均衡。 所有題目皆經由 perplexity-ai/r1-1776… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-math-reasoning-2k.texttext-generation1K<n<10K6 likes46 downloads1y agoHugging Face15DuoNeural /cot-reasoning-2k DuoNeural CoT Reasoning Dataset (2K) A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces. Benchmark Results Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090): Metric Baseline Post-SFT Δ Absolute Δ Relative GSM8K (flexible-extract) 0.3177 0.4890 +17.1pp +53.9% GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.texttext-generation1K<n<10K1 likes43 downloads5mo agoHugging Face16Abhishekh13 /Question-Anchored-Tutoring-Dialogues-2k Question-Anchored-Tutoring-Dialogues-2k This dataset contains dialogues from math tutoring interventions recorded on Eedi. Dataset Details Dataset Description Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data: DQ-Question-Metadata: The question the student was working on that prompted the… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekh13/Question-Anchored-Tutoring-Dialogues-2k.tabulartext-generation10K<n<100K0 likes43 downloads3d agoHugging Face17ansulev /gpt-5.4-xhigh-reasoning-2k Gpt-5.4-Xhigh-Reasoning-2000x A premium-quality reasoning dataset containing 2,007 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces that are significantly longer and more rigorous than standard GPT-5.4 outputs. This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with explicit… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gpt-5.4-xhigh-reasoning-2k.textquestion-answering1K<n<10K1 likes41 downloads6mo agoHugging Face18stindardlogic /chain-of-thought-dpo-2k Chain-of-Thought DPO Pairs (2.6K) DPO preference pairs for training LLMs to reason explicitly before answering. Dataset Description 2,600 preference pairs across 6 reasoning categories: Category Examples Description math_word ~610 Multi-step math word problems coding ~420 Algorithm complexity, CS reasoning economics ~415 Economic analysis and theory science ~390 Physics, chemistry, biology reasoning logic ~390 Deductive reasoning, puzzles… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/chain-of-thought-dpo-2k.texttext-generation1K<n<10K0 likes37 downloads2mo agoHugging Face19AWeirdDev /zh-tw-articles-2kHey! Also check out AWeirdDev/zh-tw-pts-articles-sm for a news source verified by the vast majority. zh-tw-articles-2k 🐣English • 🇹🇼 繁體中文 This dataset contains Taiwan news articles scraped from (https://www.storm.mg) on March 2024. Size: 5.0MB (5294263 bytes) Rows: 2000, from 20n20n20n nnn pages: 100 Dataset({ features: ['image', 'title', 'content', 'tag', 'author', 'timestamp', 'link'], num_rows: 2000 }) Use The Dataset Use 🤗 Datasets to download… See the full description on the dataset page: https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-2k.imagetext-generation1K<n<10K3 likes33 downloads2y agoHugging Face20andynik /numina-tir-2kx4 Comparison of Problem-solving Performance Across Mathematical Domains with LLMs This repository contains a filtered subset of the Numina-Math-TIR dataset, reorganised into four mathematical domains: algebra, geometry, number theory, and combinatorics. For each problem solutions were generated with LLMs: GPT-4o-mini, Mathstral-7B, Qwen2.5-Math-7B, and Llama-3.1-8B-Instruct. Each problem’s solution by these LLMs has been post-processed and compared against the human‐verified… See the full description on the dataset page: https://huggingface.co/datasets/andynik/numina-tir-2kx4.texttext-generation1K<n<10K0 likes29 downloads1y agoHugging Face21dlab-spp /reflection-sample-2k SPP Reflection 2k Sample A 2,000-row sample (seed 42) of dlab-spp/reflection-10m, in the identical format, for quick inspection of the data from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 📦 Full dataset: dlab-spp/reflection-10m (~10M documents). Each row pairs a pretraining document with a synthetic, value-laden reflection (first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.tabulartext-generation1K<n<10K0 likes28 downloads1mo agoHugging Face22zjhhhh /DeepScaleR-Qwen3-1.7B-2k-diverse-agreed-coded DeepScaleR-Qwen3-1.7B-2k diverse-agreed, strategy-coded 1635 competition-math problems (the claude_agrees_gold == True subset of a 2k diverse-classified DeepScaleR pool). Each row carries Claude's worked claude_solution plus three leak-free re-expressions of the strategy it deploys, drawn from a shared 116-code strategy codebook. Columns idx — row index into agentica-org/DeepScaleR-Preview-Dataset (resume/join key). problem, answer — the problem and gold answer.… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/DeepScaleR-Qwen3-1.7B-2k-diverse-agreed-coded.texttext-generation1K<n<10K0 likes21 downloads3mo agoHugging Face23stindardlogic /ro-preference-pairs-2k Romanian Preference Pairs (2K) Synthetic DPO preference pairs in Romanian targeting over-refusal and helpfulness alignment. Dataset Description 2,000 preference pairs in Romanian across 4 categories: Category Examples Description informational ~500 Factual questions about Romania, economics, law coding ~500 Python code tasks, FastAPI, SQLAlchemy task_completion ~500 Document drafting, emails, plans advice ~500 Career, productivity, technical… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/ro-preference-pairs-2k.texttext-generation1K<n<10K0 likes19 downloads2mo agoHugging Face24lawful-good-project /ipc-inst-2kДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями. texttext-generation1K<n<10K0 likes18 downloads3y agoHugging Face25agi-noobs /chess-sft-2k Chess SFT Training Dataset A curated dataset of chess positions with deep Stockfish analysis, designed for supervised fine-tuning (SFT) of language models to play and understand chess. Dataset Description This dataset contains chess positions extracted from multiple high-quality sources, each analyzed with Stockfish at depth 20 with MultiPV 3 (top 3 candidate moves). The positions are carefully filtered for quality and diversity across game phases, player skill levels… See the full description on the dataset page: https://huggingface.co/datasets/agi-noobs/chess-sft-2k.tabulartext-generation1K<n<10K0 likes18 downloads9mo agoHugging Face263amthoughts /deepseek_cot_2k Deepseek CoT 2k This dataset contains 1,515 extracted records focused on Chain-of-Thought (CoT) reasoning. It was processed from a malformed JSON source and converted into a clean, ready-to-use JSONL format. Dataset Structure Each record follows this schema: id: Unique identifier for the sample. problem: The input prompt or question. thinking: The internal reasoning or "Chain of Thought" process. solution: The final concise answer. difficulty: Categorization of… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/deepseek_cot_2k.texttext-generation1K<n<10K1 likes18 downloads4mo agoHugging Face27atasoglu /turkish-function-calling-2kUsed argilla-warehouse/python-seed-tools to sample tools. texttext-generation1K<n<10K4 likes17 downloads2y agoHugging Face28bombman /thwiki-2026-super-clean-2k 🇹🇭 Thai Wikipedia Super-Clean (2026 Edition) Dataset ชุดนี้สกัดจาก Wikipedia ภาษาไทย (Dump 2026) โดยเน้นความสะอาดระดับ "Pro-Clean" เพื่อใช้สำหรับ Distillation และ Fine-tuning LLM โดยเฉพาะ Key Features High Information Density: คัดเฉพาะบทความที่มีเนื้อหายาวเกิน 1,000 ตัวอักษร และมีสัดส่วนภาษาไทย > 50% Pro-Cleaned: ลบชื่อไฟล์ภาพ (.jpg, .png), ขยะสัญลักษณ์ Wiki (==, '''), และวงเล็บเปล่าออกทั้งหมด Entity Preserved: เก็บเครื่องหมายคำพูด " "… See the full description on the dataset page: https://huggingface.co/datasets/bombman/thwiki-2026-super-clean-2k.texttext-generation1K<n<10K0 likes17 downloads5mo agoHugging Face29Rupesh2 /medical-dataset-2k-phitexttext-generation1K<n<10K0 likes15 downloads2y agoHugging Face30Dddixyy /Italian-Reasoning-Logic-2ktexttext-generation1K<n<10K1 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.