datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.Question-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.Twin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.WRIT-2K
WRIT-2K
WRIT-2K is a 2,000-trajectory supervised fine-tuning dataset for multi-turn, tool-using customer-service agents on tau2-bench style airline and retail tasks.
This dataset accompanies the paper WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents.
Project homepage: https://hengrui-gu.github.io/WRIT/
Dataset Summary
WRIT-2K contains complete multi-turn trajectories with user messages, assistant natural-language responses… See the full description on the dataset page: https://huggingface.co/datasets/Henryoung/WRIT-2K.Muse-Glimmer-SWE-Gym-2k
Muse-Glimmer-SWE-Gym-2k
Agentic coding traces from meta-models/Muse-Glimmer-30B, recorded for training a
speculative-decoding drafter. 1,981 mini-swe-agent trajectories over SWE-Gym and
SWE-bench-extra instances, and the 159,999 individual chat-completion calls behind them.
Configs
Config
Rows
Size
What it is
train
1,981
57 MB
One row per trajectory: the full conversation as messages.
raw
159,999
2.7 GB
One row per recorded API call: request and… See the full description on the dataset page: https://huggingface.co/datasets/Satgoy152/Muse-Glimmer-SWE-Gym-2k.HQ-Chat-2k
🧠 HQ-Chat-2K — High-Quality Conversational & Instruction-Tuning Dataset
2,000 carefully curated, high-quality conversation and instruction examples for fine-tuning Small Language Models (SLMs) and compact LLMs from ~500M to 3B parameters.
HQ-Chat-2K is a high-quality conversational and instruction-tuning dataset designed specifically for training and fine-tuning small to medium-sized Large Language Models (LLMs).
The dataset contains 2,000 curated user–assistant examples… See the full description on the dataset page: https://huggingface.co/datasets/ThinkNet/HQ-Chat-2k.SynthUI-Code-2k-v1Synth UI 🎹
https://www.synthui.design
Dataset details
This dataset aims to provide a diverse collection of NextJS code snippets, along with their corresponding instructions, to facilitate the training of language models for NextJS-related tasks. It is designed to cover a wide range of NextJS functionalities, including UI components, routing, state management, and more.
This dataset consists of:
Note: The dataset is seperated into two main parts:
raw Contains only the… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/SynthUI-Code-2k-v1.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/krajavi3/Twin-2K-500.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/chadreadey/Twin-2K-500.NOVEReason_2k
NOVEReason_2k
NOVEReason is the dataset used in the paper NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning. It is a multi-domain, multi-task, general-purpose reasoning dataset, comprising seven curated datasets across four subfields: general reasoning, creative writing, social intelligence, and multilingual understanding. The data has been carefully cleaned and filtered to ensure suitability for training large reasoning models using… See the full description on the dataset page: https://huggingface.co/datasets/thinkwee/NOVEReason_2k.SynthUI-Code-Instruct-2k-v1Synth UI 🎹
https://www.synthui.design
Dataset details
This dataset aims to provide a diverse collection of NextJS code snippets, along with their corresponding instructions, to facilitate the training of language models for NextJS-related tasks. It is designed to cover a wide range of NextJS functionalities, including UI components, routing, state management, and more.
This dataset consists of:
Note: The dataset is seperated into two main parts:
raw Contains only the… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/SynthUI-Code-Instruct-2k-v1.Twin-2K-500_edit
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Shar999/Twin-2K-500_edit.agentic-foresight-actions-2k
Agentic Foresight: 2K Multi-Step JSON Action & Rollback Dataset
Dataset Description
This dataset contains 2,000 highly structured, synthetically generated input/output pairs explicitly designed to train Large Language Models in Agentic Foresight, Multi-Step Orchestration, and Sequential Task Automation.
Unlike standard tool-calling datasets that map a single prompt to a single API call, this dataset forces the model to act as a macro-orchestrator. It translates… See the full description on the dataset page: https://huggingface.co/datasets/Qapdex/agentic-foresight-actions-2k.tw-math-reasoning-2k
Dataset Card for tw-math-reasoning-2k
tw-math-reasoning-2k 是一個繁體中文數學語言資料集,從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,並透過 perplexity-ai/r1-1776 模型以繁體中文重新生成具邏輯性且詳盡的解題過程與最終答案。此資料集可作為訓練或評估繁體中文數學推理模型的高品質參考語料。
Dataset Details
Dataset Description
tw-math-reasoning-2k 是一個繁體中文數學語言資料集,旨在提供高品質的解題語料以支援中文數學推理模型的訓練與評估。此資料集從 HuggingFaceH4/MATH 英文數學題庫中精選 2,000 題,涵蓋代數、幾何、機率統計等各類題型,並確保題目類型分佈均衡。
所有題目皆經由 perplexity-ai/r1-1776… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-math-reasoning-2k.cot-reasoning-2k
DuoNeural CoT Reasoning Dataset (2K)
A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces.
Benchmark Results
Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090):
Metric
Baseline
Post-SFT
Δ Absolute
Δ Relative
GSM8K (flexible-extract)
0.3177
0.4890
+17.1pp
+53.9%
GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.Question-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekh13/Question-Anchored-Tutoring-Dialogues-2k.gpt-5.4-xhigh-reasoning-2k
Gpt-5.4-Xhigh-Reasoning-2000x
A premium-quality reasoning dataset containing 2,007 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces that are significantly longer and more rigorous than standard GPT-5.4 outputs.
This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with explicit… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gpt-5.4-xhigh-reasoning-2k.chain-of-thought-dpo-2k
Chain-of-Thought DPO Pairs (2.6K)
DPO preference pairs for training LLMs to reason explicitly before answering.
Dataset Description
2,600 preference pairs across 6 reasoning categories:
Category
Examples
Description
math_word
~610
Multi-step math word problems
coding
~420
Algorithm complexity, CS reasoning
economics
~415
Economic analysis and theory
science
~390
Physics, chemistry, biology reasoning
logic
~390
Deductive reasoning, puzzles… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/chain-of-thought-dpo-2k.zh-tw-articles-2kHey! Also check out AWeirdDev/zh-tw-pts-articles-sm for a news source verified by the vast majority.
zh-tw-articles-2k
🐣English • 🇹🇼 繁體中文
This dataset contains Taiwan news articles scraped from (https://www.storm.mg) on March 2024.
Size: 5.0MB (5294263 bytes)
Rows: 2000, from 20n20n20n
nnn pages: 100
Dataset({
features: ['image', 'title', 'content', 'tag', 'author', 'timestamp', 'link'],
num_rows: 2000
})
Use The Dataset
Use 🤗 Datasets to download… See the full description on the dataset page: https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-2k.numina-tir-2kx4
Comparison of Problem-solving Performance Across Mathematical Domains with LLMs
This repository contains a filtered subset of the Numina-Math-TIR dataset, reorganised into four mathematical domains: algebra, geometry, number theory, and combinatorics. For each problem solutions were generated with LLMs: GPT-4o-mini, Mathstral-7B, Qwen2.5-Math-7B, and Llama-3.1-8B-Instruct. Each problem’s solution by these LLMs has been post-processed and compared against the human‐verified… See the full description on the dataset page: https://huggingface.co/datasets/andynik/numina-tir-2kx4.reflection-sample-2k
SPP Reflection 2k Sample
A 2,000-row sample (seed 42) of dlab-spp/reflection-10m,
in the identical format, for quick inspection of the data from
Synthetic Persona Pretraining (SPP): Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
📦 Full dataset: dlab-spp/reflection-10m (~10M documents).
Each row pairs a pretraining document with a synthetic, value-laden reflection
(first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.DeepScaleR-Qwen3-1.7B-2k-diverse-agreed-coded
DeepScaleR-Qwen3-1.7B-2k diverse-agreed, strategy-coded
1635 competition-math problems (the claude_agrees_gold == True subset of a
2k diverse-classified DeepScaleR pool). Each row carries Claude's worked
claude_solution plus three leak-free re-expressions of the strategy it
deploys, drawn from a shared 116-code strategy codebook.
Columns
idx — row index into agentica-org/DeepScaleR-Preview-Dataset (resume/join key).
problem, answer — the problem and gold answer.… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/DeepScaleR-Qwen3-1.7B-2k-diverse-agreed-coded.ro-preference-pairs-2k
Romanian Preference Pairs (2K)
Synthetic DPO preference pairs in Romanian targeting over-refusal and helpfulness alignment.
Dataset Description
2,000 preference pairs in Romanian across 4 categories:
Category
Examples
Description
informational
~500
Factual questions about Romania, economics, law
coding
~500
Python code tasks, FastAPI, SQLAlchemy
task_completion
~500
Document drafting, emails, plans
advice
~500
Career, productivity, technical… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/ro-preference-pairs-2k.ipc-inst-2kДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями.
chess-sft-2k
Chess SFT Training Dataset
A curated dataset of chess positions with deep Stockfish analysis, designed for supervised fine-tuning (SFT) of language models to play and understand chess.
Dataset Description
This dataset contains chess positions extracted from multiple high-quality sources, each analyzed with Stockfish at depth 20 with MultiPV 3 (top 3 candidate moves). The positions are carefully filtered for quality and diversity across game phases, player skill levels… See the full description on the dataset page: https://huggingface.co/datasets/agi-noobs/chess-sft-2k.deepseek_cot_2k
Deepseek CoT 2k
This dataset contains 1,515 extracted records focused on Chain-of-Thought (CoT) reasoning. It was processed from a malformed JSON source and converted into a clean, ready-to-use JSONL format.
Dataset Structure
Each record follows this schema:
id: Unique identifier for the sample.
problem: The input prompt or question.
thinking: The internal reasoning or "Chain of Thought" process.
solution: The final concise answer.
difficulty: Categorization of… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/deepseek_cot_2k.turkish-function-calling-2kUsed argilla-warehouse/python-seed-tools to sample tools.
thwiki-2026-super-clean-2k
🇹🇭 Thai Wikipedia Super-Clean (2026 Edition)
Dataset ชุดนี้สกัดจาก Wikipedia ภาษาไทย (Dump 2026) โดยเน้นความสะอาดระดับ "Pro-Clean" เพื่อใช้สำหรับ Distillation และ Fine-tuning LLM โดยเฉพาะ
Key Features
High Information Density: คัดเฉพาะบทความที่มีเนื้อหายาวเกิน 1,000 ตัวอักษร และมีสัดส่วนภาษาไทย > 50%
Pro-Cleaned: ลบชื่อไฟล์ภาพ (.jpg, .png), ขยะสัญลักษณ์ Wiki (==, '''), และวงเล็บเปล่าออกทั้งหมด
Entity Preserved: เก็บเครื่องหมายคำพูด " "… See the full description on the dataset page: https://huggingface.co/datasets/bombman/thwiki-2026-super-clean-2k.medical-dataset-2k-phiItalian-Reasoning-Logic-2k
