datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.Pytorch-Code-10K
Hot Coco Training Dataset
A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!)
Dataset Description
This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes:
code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.10k_prompts_ranked
Dataset Card for 10k_prompts_ranked
10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts.
The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.sec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.PKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
Ox-Alpha-10k
Ox Alpha - 10k
10,005 single-turn prompts for text-response teacher generation
Each row carries id, category, subcategory
All data was gathered using stealth/ox-alpha via OpenRouter (reasoning effort high)
Topic distribution
Category
Rows
Share
Coding (incl. Go/Rust, C++/Java/C#, shell/CLI)
944
9.5%
Knowledge QA
891
9.0%
Logical reasoning & decisions
734
7.4%
Web development
720
7.2%
Game development
720
7.2%
Three.js / browser 3D
620
6.2%… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Ox-Alpha-10k.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.LongReward-10k
LongReward-10k
💻 [Github Repo] • 📃 [LongReward Paper]
LongReward-10k dataset contains 10,000 long-context QA instances (both English and Chinese, up to 64,000 words).
The sft split contains SFT data generated by GLM-4-0520, following the self-instruct method in LongAlign. Using this split, we supervised fine-tune two models: LongReward-glm4-9b-SFT and LongReward-llama3.1-8b-SFT, which are based on GLM-4-9B and Meta-Llama-3.1-8B, respectively.
The dpo_glm4_9b and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongReward-10k.DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k
DeepScaleR Easy/Medium/Hard — Gemma 4 26B-A4B PT
This dataset contains 9,900 unique, deduplicated DeepScaleR math questions for
reinforcement-learning experiments. Difficulty is defined by how often the
pretrained google/gemma-4-26B-A4B teacher solved each question across eight
temperature-1 samples under the same rule-based grader used by the RL training
pipeline.
The Hub dataset has three configurations—easy, medium, and hard—and each
configuration has a train split with 3,000… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k.RAIL-HH-10K RAIL-HH-10K: Multi-Dimensional Safety Alignment Dataset
The first large-scale safety dataset with 99.5% multi-dimensional annotation coverage across 8 ethical dimensions.
📖 Read Blog •
📖 Paper (Coming Soon) •
🚀 Quick Start •
🔌 RAIL API •
💻 Examples
🌟 What Makes RAIL-HH-10K Special?
🎯 Near-Complete Coverage
99.5% dimension coverage across all 8 ethical dimensions
Most existing datasets: 40-70% coverage
RAIL-HH-10K: 98-100%… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/RAIL-HH-10K.nemotron-cc-10K-sample-translated
Translated Nemotron-cc-hq samples
This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample
Currently, the following are available, we will add other models and languages:
Model
Languages
Gemma-3-4b-it
["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"]
EuroLLM-9B-Instruct
["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.StreetVision-10K
StreetVision-10K
Each sample contains:
A system prompt instructing the model to act as an OSINT/geospatial expert
A user message with a street-level photo and the instruction to determine coordinates
An assistant response with ground-truth coordinates in <direct_lon_lat_output>longitude,latitude</direct_lon_lat_output> format
Format
Each line is a JSON array of ChatML messages:
[
{"role": "system", "content": "..."},
{"role": "user", "content": [… See the full description on the dataset page: https://huggingface.co/datasets/mishl/StreetVision-10K.Math-Chinese-DeepSeek-R1-10K
中文 DeepSeek-R1-Distil 数学指令微调数据集
💻 Github Repo
基本信息
数据集大小 10K,独立生成指令与回复,并非其他社区数据集的子集。所有数据经过校验,答案正确性可以得到保证。
数据集的组成如下:
问题类型
数据条数
定积分计算
2626
多项式化简
1621
因式分解
2557
多项式展开
2095
多项式方程
1101
总数
10000
数据格式
每条数据的格式如下:
{
"id": <<12位nanoid>>,
"prompt": <<提示词>>,
"reasoning": <<模型思考过程>>,
"response": <<模型最终回复>>
}
legal-contract-gpt41-redlining-10k
legal-contract-gpt41-redlining-10k
Dataset Description
This dataset contains 9977 synthetic legal contract redlines generated using GPT-4.1 model mix (base, mini, nano) with structured outputs. It is designed for fine-tuning LLMs (including OpenAI GPT-3.5/4, Llama, and other models) to assist with legal document redlining and clause revision.
Key Features
🤖 9977 synthetic redlines generated by GPT-4.1 model mix (base, mini, nano)
📋 Multiple training formats:… See the full description on the dataset page: https://huggingface.co/datasets/UmaiTech/legal-contract-gpt41-redlining-10k.sft-safe-openai-chat-10k
SFT Safe OpenAI Chat 10K
This dataset is formatted for chat SFT training. Each JSONL row contains a messages field compatible with OpenAI-style chat fine-tuning data:
{"messages":[{"role":"system","content":"..."},{"role":"user","content":"..."},{"role":"assistant","content":"..."}]}
Files:
train.jsonl: 10,000 training examples
validation.jsonl: 200 validation examples
eval.jsonl: same content as validation.jsonl, provided as an evaluation alias
Example usage:
fromdatasets import… See the full description on the dataset page: https://huggingface.co/datasets/yxx123456/sft-safe-openai-chat-10k.Apertus-v1.5-QAT-10K
mlx-community/Apertus-v1.5-QAT-10K
This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models.
10k_rows_cleaned_prompts
10K Rows Cleaned Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models.
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.english-daily-dialogues-10k
English Daily Dialogues 10K
A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark.
Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.tw-function-call-reasoning-10k
Dataset Card for tw-function-call-reasoning-10k
本資料集為繁體中文版本的函式呼叫(Function Calling)資料集,翻譯自 AymanTarig/function-calling-v0.2-with-r1-cot,而該資料集本身是 Salesforce/xlam-function-calling-60k 的修正版。我們利用語言模型翻譯後,經人工修改,旨在打造高品質的繁體中文工具使用語料。
Dataset Details
Dataset Description
tw-function-call-reasoning-10k 是一個專為語言模型「工具使用能力(Function Calling)」訓練所設計的繁體中文資料集。其內容源自 AymanTarig/function-calling-v0.2-with-r1-cot,該資料集又為 Salesforce/xlam-function-calling-60k… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-function-call-reasoning-10k.grug-think-v3-10k
grug-think-v3-10k
v2 brain short. v2 brain useful. but some v2 brain wear office shirt.
"User wants hello world Python. Provide code." short English, yes. grug, no.
v3 tear off office shirt. keep brain meat.
old: User wants hello world Python. Simple code snippet, no tools needed. Provide code and brief explanation.
new: Need Python hello-world. Tiny snippet. No tool. Give code, brief explain.
complex cave different. grug no crush branch into pebble. exact path, error… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think-v3-10k.pii-masking-mini-10k
PII Masking Mini: Multilingual Sample
A mini-sized stratified sample of pii-masking-openpii-1.5m,
the flagship release of the PII-Masking-3M family. Sampled proportionally by
(source_dataset, language) so every locale and label gets representation.
Asia Pacific rows appear first.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-mini-10k.finmix-autoscientist-10k
FinMix AutoScientist 10k
A deterministic, upload-ready 10,000-row subset of
FinMix v1, created for
fast finance adaptation runs in the Adaption AutoScientist challenge.
Use with Adaption Adaptive Data
Import this Hugging Face dataset and map:
Prompt: prompt
Context: context
Completion: completion
Leave task_type, source, and group_key unmapped. They are retained for
provenance and auditing.
Fields
Field
Description
prompt
Financial… See the full description on the dataset page: https://huggingface.co/datasets/julian8897/finmix-autoscientist-10k.JL-AgentBehavior-10K
JL-AgentBehavior-10K
JL-AgentBehavior-10K is a 10,000-record, English-language research-preview dataset for studying and training the behavioral policy of repository-level coding agents.
The dataset does not treat a coding agent as a chatbot that maps a request directly to a block of code. It represents an agent as a policy operating across a sequence of observable decisions:
task
-> repository evidence
-> bounded plan
-> tool selection
-> scoped edit strategy
->… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-AgentBehavior-10K.telecom-intent-config-sft-10k
Telecom Intent→Config SFT Dataset (10K)
The first open SFT dataset for training LLMs to translate natural language network intents into structured 5G/6G configurations.
This dataset addresses the #1 gap identified in the telecom LLM research landscape: there is no public training dataset for intent-to-policy translation. All existing telecom datasets (TeleQnA, ORANBench-13K, 6G-Bench) are MCQ evaluation benchmarks — not instruction-following format. This dataset fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/telecom-intent-config-sft-10k.soc-agent-traces-10k
SOC-Agent-Traces-10K
Multi-step SOC investigation agent traces in session-trace format. Each
record is a complete investigation session: an alert arrives, an analyst agent
gathers evidence through nine read-only tools, and closes with a structured
JSON triage report.
Instead of single-turn alert → answer pairs, every record captures the full
reasoning trajectory:
alert → get_surrounding_events → get_process_tree → lookup_attack
→ search_sigma → get_asset_context →… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/soc-agent-traces-10k.Hypa-Text-10k
A multilingual instruction-tuning dataset covering translation, dictionary lookup, language detection, structured JSON output, chain-of-thought translation breakdown, and tool use across seventeen low-resource and underrepresented languages.
Dataset Description
hypaai/Hypa-Text-10k is a 10,000-example public subset of the larger multilingual instruction corpus assembled by Hypa Intelligence to train production language models for low-resource and underrepresented languages.… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Text-10k.bielik-distill-polish-10k
bielik-distill-polish-10k
Polish instruction-tuning dataset with 10,304 samples generated via response-level knowledge distillation from Bielik-11B-v3.0-Instruct (SpeakLeash, Apache 2.0).
Covers Polish history, culture, politics, science, geography, idioms, and general reasoning. Multi-pass quality control: factual corrections, topic filtering (Poland/Europe focus), truncation removal (~9% of raw data removed).
Format
{
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/JohnTdi/bielik-distill-polish-10k.openpii-masking-mini-10k
OpenPII Masking Mini 10K
A compact, stratified subset of ai4privacy/pii-masking-openpii-1m, containing 10,000 samples for rapid experimentation, fine-tuning, and benchmarking of PII detection and masking models.
Sampling Methodology
Samples were selected using proportional stratified sampling by language:
Target count per language = round(lang_proportion × 10,000) — proportional representation.
Streaming + reservoir sampling collected 3× the target candidates per… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-mini-10k.stem-reasoning-ccbysa-10k-001
stem-reasoning-ccbysa-10k-001
Combined 10K CC-BY-SA STEM reasoning dataset — a sequential merge of ten 1K sets.
Total examples: 10000
License: CC-BY-SA-4.0
Mean quality score: 4.411 (min 4.25, max 4.81)
Unique sources: 2387
DPO pairs: 3650
Difficulty: {'advanced': 2774, 'introductory': 5409, 'intermediate': 1817}
Products: {'reasoning_chain': 1531, 'instruction_pair': 8377, 'code_instruction': 92}
Combined from
stem-reasoning-ccbysa-001… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-10k-001.
