datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
high_quality_foldingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1200,
"total_frames": 3254196,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1200"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/high_quality_folding.ChineseWebText2.0-HighQuality
📘 ChineseWebText2.0-HighQuality
Overview
ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original
CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License).
This subset retains only samples with:
quality_score ≥ 0.9
toxicity.score ≤ 0.01
The goal is to provide a cleaner and more reliable dataset suitable for
language model pre-training, instruction tuning, and quality-sensitive downstream tasks.
This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/high-quality-invoice-images-for-ocr.truthfulness_high_quality
Dataset Card for "truthfulness_high_quality"
More Information needed
Creative-Writing-High-Quality-1300x
Creative Writing - Part One (Shadow & Skeleton)
This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology.
Methodology: Shadow & Skeleton
Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach:
Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.sea-commoncrawl-high-qualitydroid_success_high_quality
DROID Success (High-Quality Extrinsics)
Subset of ~17k successful episodes in DROID-COMMUNITY filtered for high-quality camera extrinsics.
Ported from raw 1.0.1 data at full resolution to LeRobotDataset v3.0 format (0.33 TiB | 3.6k inodes) with extra annotation from KarlP/droid.
Your browser does not support the video tag.
Dataset Structure
The external cameras are assigned to left and right views depending on the episode. For their extrinsics… See the full description on the dataset page: https://huggingface.co/datasets/jnogga/droid_success_high_quality.High-Quality-Code
High-Quality-Code: Synthetic + Real (MAXIMUM CODE)
A massive, high-quality code dataset built with maximum code philosophy – as much code as possible.
Components
Synthetic syntax-correction dataset – 5M+ examples across 33 languages (original code_syntax_dataset_1GB.csv)
Real high-quality code from GitHub – 500 top-starred repositories – BOTH zips and extracted source
Current Status: IN PROGRESS
Target: 500 repos
Currently uploaded: 12 extracted… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/High-Quality-Code.Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914
Mirod simulation subset: three tasks, approximately 3x real frames
This release contains simulation data only, selected for mixed training with
the local Real_3tasks release. Real recordings are not included.
Folder
Task
Real reference frames
Simulation episodes
Simulation frames
Ratio
task1/dataset
Stack the small box on the other box (叠盒子)
13,785
171
41,358
3.0002
task2/dataset
Put the cup into the tray (杯子入盘)
12,361
192
37,081
2.9998
task3/dataset
Take the box… See the full description on the dataset page: https://huggingface.co/datasets/liujiting/Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914.high-quality-english-sentences
High-Quality English Sentences
Dataset Description
This dataset contains a collection of high-quality English sentences sourced from C4 and FineWeb (not FineWeb-Edu). The sentences have been carefully filtered and processed to ensure quality and uniqueness.
"High-quality" means they're legible English and not spam, although they may still have spelling and grammar errors.
Source Data
Before filtering:
C4: 1 million sentences
FineWeb: 1 million sentences… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-english-sentences.high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/JuanfelipeX123/high-quality-invoice-images-for-ocr.high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/Shubhal829/high-quality-invoice-images-for-ocr.dolma_20bn_cc_high_qualityhigh-quality-multilingual-sentences
High Quality Multilingual Sentences
This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset.
It includes 1.58 million rows across 51 different languages, each in its own configuration.
Example row (from the all config):
{
"text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.",
"fasttext": "fa",
"gcld3": "fa"
}
Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.high-quality-midjouney-srefs
Midjourney Image Scraper & Dataset Creator
A complete toolkit for scraping Midjourney images, generating captions, and creating HuggingFace datasets with optional automatic upload to HuggingFace Hub.
🌟 Features
🔍 Web Scraping: Download images from midjourneysref.com with comprehensive error handling
🤖 AI Captioning: Automatic image captioning using Moondream API with auto-resume capability
✂️ Smart Cropping: AI-powered image cropping using OpenAI to optimize aspect… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/high-quality-midjouney-srefs.high-quality-cc-21b
high_quality
A high-quality English web text corpus extracted from Common Crawl WARC files using an
LLM-based extraction and quality pipeline.
Dataset Summary
high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl
WARC records are passed through an LLM-based extractor that strips boilerplate and recovers
the main content, then filtered to retain only documents in the "high_quality" band,
deduplicated (exact + fuzzy), and… See the full description on the dataset page: https://huggingface.co/datasets/MichaelR207/high-quality-cc-21b.high-quality_art-mix_images_for_diffusion_training_1ai_made photorealistic
high-quality-text
High Quality Text Dataset
A curated collection of English-language texts for AI training and research.
Sources
HuggingFaceFW/fineweb-edu
openbmb/Ultra-FineWeb
Zyphra/Zyda-2
EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample
m-a-p/FineFineWeb
Each dataset was processed as follows:
Split into approximately 2 000-token chunks using the LLaMA 3.1 tokenizer.
Cleaned by normalizing spaces, punctuation, and characters, and replacing emails and phone numbers with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-text.uzbek-high-quality-10hCybersecurity-High-Quality-Dataset
Cybersecurity High-Quality Dataset (网络安全高质量数据集)
概述 | Overview
这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。
A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanity tool, with only data scoring 4.5… See the full description on the dataset page: https://huggingface.co/datasets/atmike/Cybersecurity-High-Quality-Dataset.Cybersecurity-High-Quality-Dataset
Cybersecurity High-Quality Dataset (网络安全高质量数据集)
概述 | Overview
这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。
A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanitytool, with only data scoring 4.5 or… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-High-Quality-Dataset.pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.multi-wiki-qa-high-quality-subset
multi-wiki-qa-high-quality-subset
A quality-filtered subset of the Danish (da) split of
alexandrainst/multi-wiki-qa,
a Wikipedia-based extractive question-answering dataset.
Configs
Config
Samples
Description
da
4,767
All LLM-verified correct samples
da-short
3,527
Correct samples where the answer is at most 3 words
Filtering methodology
Starting from the 5,000 samples in the original Danish split:
Span validation -- deterministic check that… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/multi-wiki-qa-high-quality-subset.YiSang-HighQuality
YiSang-HighQuality
📖 Check out the KO-REAson technical report.
📍 Rest of the model and datasets are available here.
YiSang-HighQuality is a collection of ~280K long-CoT reasoning traces generated via Qwen3-32B. This dataset is a high-yield subset of the larger Yi-Sang collection, designed to enhance multilingual reasoning through Language-Mixed Chain-of-Thought (CoT), which switches between English and Korean to minimize translation artifacts while leveraging… See the full description on the dataset page: https://huggingface.co/datasets/KOREAson/YiSang-HighQuality.openarm_bimanual_shirt_folding_gen3_high_quality_rightfirst_with_rolloutsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"right_joint_1.pos",
"right_joint_2.pos",
"right_joint_3.pos",
"right_joint_4.pos",
"right_joint_5.pos",
"right_joint_6.pos",
"right_joint_7.pos"… See the full description on the dataset page: https://huggingface.co/datasets/videron/openarm_bimanual_shirt_folding_gen3_high_quality_rightfirst_with_rollouts.uzbek-high-quality-10h-en-translationHigh-Quality_Dexterous_Hand_Movements_Dataset
High-Quality Dexterous Hand Movements — sample
This Hugging Face dataset is a non-commercial sample export from the Quality Vision Motion Dataset Engine: MediaPipe Hands (21 landmarks) + optional dexterous analytics and hand motion_intelligence.
This is a sample, not the full commercial pack.For larger commercial batches (e.g. ~140k+ HQ frames, multiple clips segmented for specific dexterous motions), see Dataset pricing and contact info@qvision.space.
Files
This repo… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/High-Quality_Dexterous_Hand_Movements_Dataset.high-quality-crash-coursehigh_quality_private_evaluationsHigh-quality question-answer pairs, from private versions of datasets designed to mimic ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/.
Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC-like questions), and a question quality score.
