datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.40546875
Action score: 0.475
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4046875
Action score: 0.4703125
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.39921875
Action score: 0.44375
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38359375
Action score: 0.4703125
Valid samples: 320/320
tigerbot-kaggle-leetcodesolutions-en-2kTigerbot 基于leetcode-solutions数据集,加工生成的代码类sft数据集
原始来源:https://www.kaggle.com/datasets/erichartford/leetcode-solutions
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-kaggle-leetcodesolutions-en-2k')
WRIT-2K
WRIT-2K
WRIT-2K is a 2,000-trajectory supervised fine-tuning dataset for multi-turn, tool-using customer-service agents on tau2-bench style airline and retail tasks.
This dataset accompanies the paper WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents.
Project homepage: https://hengrui-gu.github.io/WRIT/
Dataset Summary
WRIT-2K contains complete multi-turn trajectories with user messages, assistant natural-language responses… See the full description on the dataset page: https://huggingface.co/datasets/Henryoung/WRIT-2K.radeonvla_reflex_physical_2k
RadeonVLA-Reflex Physical-2K
Physical-2K contains 2,000 strictly validated successful Genesis episodes for
language-conditioned Franka fruit sorting. Coverage is exactly five fruits ×
four bowl positions × 100 episodes = 2,000 episodes.
Verified release facts
Item
Value
Episodes
2,000
Frames
468,889
Registered task variations
20
Episodes per task variation
100
Control frequency
20 Hz
Strict success certificates
2,000
Sampled image frames
1… See the full description on the dataset page: https://huggingface.co/datasets/a3124371940/radeonvla_reflex_physical_2k.HQ-Chat-2k
🧠 HQ-Chat-2K — High-Quality Conversational & Instruction-Tuning Dataset
2,000 carefully curated, high-quality conversation and instruction examples for fine-tuning Small Language Models (SLMs) and compact LLMs from ~500M to 3B parameters.
HQ-Chat-2K is a high-quality conversational and instruction-tuning dataset designed specifically for training and fine-tuning small to medium-sized Large Language Models (LLMs).
The dataset contains 2,000 curated user–assistant examples… See the full description on the dataset page: https://huggingface.co/datasets/ThinkNet/HQ-Chat-2k.Code-Golang-QA-2k
Code-Golang-QA-2k
This (small) dataset comprises 2,000 question-and-answer entries related to the Go programming language. It is designed to serve as a resource for individuals looking to enhance machine learning models, create chatbots, or simply to provide a comprehensive knowledge base for developers working with Go.
Data Format
[
{
"question": "How do you create a new RESTful API endpoint using Gin?",
"answer": "Creating a new RESTful API endpoint… See the full description on the dataset page: https://huggingface.co/datasets/ExAi/Code-Golang-QA-2k.Qwen2.5-32B-Instruct_agent_trajectories_2k
Dataset Summary
This dataset contains agent trajectories generated by the Qwen2.5-32B-Instruct model using smolagents library as the agent framework.
For more details on the method, data format, and applications, refer to the following:
Repository: https://github.com/Nardien/agent-distillation
Paper: https://arxiv.org/abs/2505.17612
WildChat-2k-TypeTopic
WildChat-2k-TypeTopic: The Manually Curated Edition
Dataset Description
WildChat-2k-TypeTopic is a manually curated subset of 1,880 real-world user prompts from the WildChat dataset, featuring annotations for both task type (e.g. knowledge recall, problem solving, creative, lists) and topic category (e.g. personal assistance, math, ai, household)
Why this dataset?
Suppose you want to answer a research question such as "What kind of user prompt does the LLM like… See the full description on the dataset page: https://huggingface.co/datasets/dpaleka/WildChat-2k-TypeTopic.igcse-economics-qa-2kHumanLike-Casual-2Kn8n-workflows-2k
Dataset Card for N8n Workflows 2k
This dataset contains 2000 n8n workflows.
Curated by: Arkel AI
Funded by: Arkel AI
Language(s) (NLP): English
License: Apache 2.0
Code-Golang-QA-2k-dpo
Code-Golang-QA-2k
This (small) dataset comprises ~1.8k dpo entries related to the Go programming language. It is designed to serve as a resource for individuals looking to enhance machine learning models, create chatbots, or simply to provide a comprehensive knowledge base for developers working with Go.
Data Format
[
{
"question": "How do you create a new RESTful API endpoint using Gin?",
"chosen_answer": "Creating a new RESTful API endpoint using the Gin… See the full description on the dataset page: https://huggingface.co/datasets/ExAi/Code-Golang-QA-2k-dpo.agentic-foresight-actions-2k
Agentic Foresight: 2K Multi-Step JSON Action & Rollback Dataset
Dataset Description
This dataset contains 2,000 highly structured, synthetically generated input/output pairs explicitly designed to train Large Language Models in Agentic Foresight, Multi-Step Orchestration, and Sequential Task Automation.
Unlike standard tool-calling datasets that map a single prompt to a single API call, this dataset forces the model to act as a macro-orchestrator. It translates… See the full description on the dataset page: https://huggingface.co/datasets/Qapdex/agentic-foresight-actions-2k.Qwen2.5-32B-Instruct_agent_trajectories_2k_prefix
Dataset Summary
This dataset contains 2k agent trajectories generated by the Qwen2.5-32B-Instruct model using smolagents library as the agent framework.
The trajectories are collected using the "first-thought prefix" method, where each trajectory is prefixed by the model's initial reasoning steps derived from Chain-of-Thought (CoT) prompting.
For more details on the method, data format, and applications, refer to the following:
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/agent-distillation/Qwen2.5-32B-Instruct_agent_trajectories_2k_prefix.tigerbot-kaggle-recipes-en-2kTigerbot 基于公开的数据集生成的食谱类sft数据集
原始来源:https://www.kaggle.com/datasets/zeeenb/recipes-from-tasty?select=ingredient_and_instructions.json
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-kaggle-recipes-en-2k')
planner_instruction_tuning_2kBootstrap 2k Planner finetuning dataset for ReWOO.
It is a mixture of "correct" HotpotQA and TriviaQA task planning trajectories in ReWOO Framework.
CondAmbigQA-2K
CondAmbigQA-2K Dataset
Dataset Description
This is an expanded version of the CondAmbigQA dataset, growing from the original 200 entries to 2000 entries.
Dataset Summary
CondAmbigQA-2K contains 2000 question-answering pairs with conditional contexts and ground truth answers. Each entry includes:
Question: The ambiguous question
Properties: Contains condition, groundtruth, and citations
Context (ctxs): Retrieved relevant passages with scores
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Apocalypse-AGI-DAO/CondAmbigQA-2K.Qwen2.5-32B-Instruct_cot_trajectories_2k
Dataset Summary
This dataset contains CoT trajectories generated by the Qwen2.5-32B-Instruct model.
This dataset contains both correct and incorrect trajectories.
For more details on the method, data format, and applications, refer to the following:
Repository: https://github.com/Nardien/agent-distillation
Paper: https://arxiv.org/abs/2505.17612
tigerbot-dolly-classification-en-2kTigerbot 基于dolly数据集加工的分类classification相关分类的的sft。
原始来源:https://huggingface.co/datasets/databricks/databricks-dolly-15k
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-dolly-classification-en-2k')
gpt-5.4-xhigh-reasoning-2k
Gpt-5.4-Xhigh-Reasoning-2000x
A premium-quality reasoning dataset containing 2,007 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces that are significantly longer and more rigorous than standard GPT-5.4 outputs.
This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with explicit… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gpt-5.4-xhigh-reasoning-2k.cot-reasoning-2k
DuoNeural CoT Reasoning Dataset (2K)
A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces.
Benchmark Results
Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090):
Metric
Baseline
Post-SFT
Δ Absolute
Δ Relative
GSM8K (flexible-extract)
0.3177
0.4890
+17.1pp
+53.9%
GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.chain-of-thought-dpo-2k
Chain-of-Thought DPO Pairs (2.6K)
DPO preference pairs for training LLMs to reason explicitly before answering.
Dataset Description
2,600 preference pairs across 6 reasoning categories:
Category
Examples
Description
math_word
~610
Multi-step math word problems
coding
~420
Algorithm complexity, CS reasoning
economics
~415
Economic analysis and theory
science
~390
Physics, chemistry, biology reasoning
logic
~390
Deductive reasoning, puzzles… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/chain-of-thought-dpo-2k.french-customer-review-sentiment-free-2k
French Customer Review Sentiment Free 2K
French Customer Review Sentiment Free 2K is a free 2,000-record sample extracted from the full commercial dataset French Customer Review Sentiment (100k synthetic reviews) provided by Kinoux.
Each entry is a synthetic French customer review labeled with a 3-class sentiment:
positive
neutral
negative
The data is 100% synthetic (no personal data, no real platform exports) and was generated and curated specifically for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/Kinoux/french-customer-review-sentiment-free-2k.2k3n1d-Temel-cumle-analizi-284k
🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance).
🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/2k3n1d-Temel-cumle-analizi-284k.dino_wm_pusht_noise_2k_lerobot
DINO-WM PushT (noise 2k) (TsFile)
Apache TsFile version of neiltan/dino-wm-pusht-noise-2k.
Overview
A PushT (planar block-pushing) demonstration dataset in the LeRobot v2.1 format,
of the kind used to train and evaluate DINO-WM-style world models. In the PushT
task an agent pushes a T-shaped block to a fixed goal pose; this collection
contains 2,000 noisy demonstration episodes of that single task.
Each record is one control frame within an episode: the 2-D… See the full description on the dataset page: https://huggingface.co/datasets/THULab/dino_wm_pusht_noise_2k_lerobot.arch-persona-coherence-lpi-260902T2130-apple-lpi-2kfrench-spam-ham-detection-free-2k
French Spam/Ham Detection Free 2K
French Spam/Ham Detection Free 2K is a free 2,000-record sample extracted from the full commercial dataset French Spam/Ham Detection (56,400 synthetic messages) provided by Kinoux.
Each entry is a synthetic French message labeled with a binary classification:
spam
ham
The data is 100% synthetic (no personal data, no scraped emails, no real platform exports) and was generated specifically for training and evaluating French-native spam detection… See the full description on the dataset page: https://huggingface.co/datasets/Kinoux/french-spam-ham-detection-free-2k.
