datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-belebeletiny-mlqaPhatgoose_flanv2_offlinelaion-2b-en-unsafe-quarter-three-download-hdSCAND_traj_selectionMISKG(For transparency, this markdown content is directly taken from original MISKG github page - https://github.com/kanak8278/MISKG/blob/main/README.md)
Multimodal Ingredient Substitution Knowledge Graph for Personalized Dietary Recommendations (MISKG)
Abstract
Ingredient substitution is essential in adapting recipes to meet individual dietary needs, preferences, and ingredient availability. We introduce a Multimodal Ingredient Substitution Knowledge Grap[...]
File… See the full description on the dataset page: https://huggingface.co/datasets/C-three-AN-Nourich/MISKG.three-check-chess-games
[!CAUTION]
This dataset is still a work in progress and some breaking changes might occur. In the meantime, please use https://database.lichess.org/#variant_games
tokenization_robustness_v102
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/tokenization_robustness_v102.pi07_cable_three_vector_v1
Three-holder cable routing with vector goals
Real-robot demonstrations for a vector-conditioned low-level policy: place three holders and route a cable through each holder. This is the validated LeRobot v3 dataset prepared for the first Pi0.7 four-camera-goal cable policy.
Property
Value
Source recordings
140
Subtask episodes
840
Frames
335,897
Sampling rate
100 Hz
Robot
ARX bimanual
Recorded state / action
14 / 14 dimensions
Video resolution
448 × 448… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/pi07_cable_three_vector_v1.3.33BT-Pre-Training-Mix-One-Of-ThreeThis is mix one of three in a ten billion token pre-training cirriculum
Dataset
Rows
Token share
FineWeb-Edu
1,568,571
44.91%
DCLM
1,001,522
40.13%
Cosmopedia-V2
501,667
9.97%
Code
350,000
4.98%
Math
0
0%
About 3.34 billion tokens.Mix two and three can be found here: Hoglet-33/3.33BT-Pre-Training-Mix-Two-Of-Three, Hoglet-33/3.33BT-Pre-Training-Mix-Three-Of-Three
instructtts-three-model-gemini-zh
InstructTTSEval 三模型 Gemini 评测数据
本目录整理了 InstructTTSEval 中文集上三个 TTS 模型的生成音频和 Gemini 一致性评测结果:Qwen3-TTS-12Hz-1.7B-VoiceDesign、Seed-Audio-1.0、VoxCPM2。
字段
records.jsonl 每行对应一个模型和一种控制格式(APS、DSD 或 RP):
id:InstructTTSEval 样本 ID
mode:控制格式
model、model_name:模型标识
text:合成文本
instruction:历史评测记录中的输入控制指令,按本行 APS/DSD/RP 格式保留;不等同于各模型 API 的完整请求封装
generated_audio:该模型生成音频的相对路径
reference_audio:原始参考音频的相对路径
gemini_consistent:Gemini judge 的一致性判断
inconsistency_reason:判断为不一致时的原因… See the full description on the dataset page: https://huggingface.co/datasets/zsy814/instructtts-three-model-gemini-zh.threejs-gamecode-instruct-v3-ultra
Three.js GameCode Instruct v3 Ultra
This is a large synthetic/original instruction dataset for training or testing LLM behavior around Three.js, browser game development, gameplay programming, debugging, optimization, architecture, and general coding.
Important note
This dataset is synthetic and programmatically generated from original templates. It is designed as a useful starting point for experiments, not as a fully hand-curated gold-standard benchmark.
No… See the full description on the dataset page: https://huggingface.co/datasets/agagasf123123/threejs-gamecode-instruct-v3-ultra.3.33BT-Pre-Training-Mix-Three-Of-ThreeThis is mix three of three in a ten billion token pre-training cirriculum.
Dataset
Rows
Token share
FineWeb-Edu
871,429
25.01%
DCLM
498,667
20.03%
Cosmopedia-V2
752,500
14.99%
Code
1,400,000
19.99%
Math
382,000
19.99%
About 3.336 billion tokens.Mix one and two can be found here: Hoglet-33/3.33BT-Pre-Training-Mix-One-Of-Three, Hoglet-33/3.33BT-Pre-Training-Mix-Two-Of-Three
shakespeare-complete-works
Shakespeare Complete Works Dataset
This dataset contains the complete works of William Shakespeare, including:
The Sonnets (154 sonnets)
Plays (Tragedies, Comedies, Histories)
Poems
Dataset Structure
Each entry contains:
work: The title of the work
section: Specific section (e.g., "Sonnet 1") if applicable
text: The actual text content
type: Type of work (sonnet, play, poem)
id: Unique identifier
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/r-three/shakespeare-complete-works.2026-07-30-qwen36-threeway-constitution-odcv-eval
Qwen3.6-27B three-way constitution LoRA — ODCV evaluation
field
value
experiment
ODCV-Bench evaluation of the Qwen3.6-27B three-way constitution LoRA on the controlled 78-scenario subset used by the difficult-advice mixture sweep.
date_generated
2026-07-30
constitution
2026-07-29 synthdoc approved constitution SFT, combining embodied, difficult-advice, and agentic tool-use constitution corpora.
source_repo
teaching_claude_why_replication at commit… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-qwen36-threeway-constitution-odcv-eval.three-direction-pretrain-datasetThis is the pre-training dataset we used to train our ChemBART model in the paper ChemBART (arXiv:2601.02915)
SCAND_path_selectiontiny-m-hellaswagrobotwin-stack-blocks-three-rollouts
RoboTwin stack_blocks_three — Wan2.2 TI2V Rollouts
160 text+image-to-video rollouts (10 initial conditions × 16 random seeds) for the
stack_blocks_three task from RoboTwin, generated with the Wan2.2 TI2V (5B)
diffusion model fine-tuned with a merged Vidar LoRA adapter, and scored with the
blocks_stack_v2 reward (SAM3 object tracking + IDM inverse-dynamics + FK
gripper ↔ block position matching).
Companion to the EmbodiedVideoRL / DanceGRPO
reward-model work.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/VincentNi/robotwin-stack-blocks-three-rollouts.rr_three_tasks_v1
rr_three_tasks_v1
Three tasks on a Trossen AI solo arm, merged into one LeRobot v2.1 dataset.
task
episodes
frames
pick_specific_item_from_clutter
243
59088
pick_two_in_order
99
40478
open_pot_and_place
100
47288
meta/sources.jsonl maps every episode to its source dataset, episode and revision, with the
staging record (open_pot_and_place variant, pick_two second object, sheet row).
Held-out evaluation episodes
meta/eval_episodes_v1.json: 44… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/rr_three_tasks_v1.laion-2b-en-unsafe-quarter-three-downloadtulu3-sft-clustered8-seed123-mixing0.1synthe_ds_8b_4type_threethreedscans
Three D Scans
A mirror of threedscans.com, the archive of high-resolution
3D scans of museum objects initiated in 2012 by artist Oliver Laric.
134 meshes (10.6 GB) covering antiquities, classical and 19th-century sculpture,
anatomical casts, and natural-history specimens, scanned in collaboration with museums
across Europe.
This mirror exists so the collection can be fetched programmatically and cited stably.
All credit for the scans belongs to Oliver Laric and the holding… See the full description on the dataset page: https://huggingface.co/datasets/alecjacobson/threedscans.three-million-bluesky
3 million bluesky posts
raw.zip is the ~5m posts that i initially pulled that were full of duplicates, where i removed the duplicates and compiled those into 2.6 million posts in final_posts.jsonl
the rest of the .jsonl files in the data folder, and not inside the raw.zip are unchecked and may have duplicates. feel free to write your own script to remove them, but this should contain around 3 million unique posts.
adidas-tracksuit-with-three-stripes-beta-versionCreated using https://github.com/D3voz/joy-caption-beta-one-gui-mod
synthetic_4type_threeeasyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-stage-three-temp-1_7-RL-zero-correct-to-0.2threew
threew
threew adalah dataset pasangan instruksi–jawaban berbahasa Indonesia untuk eksperimen text generation dan instruction tuning. Dataset ini berisi contoh sintetis yang dikurasi secara programatis dan diarahkan agar jawaban bersifat jelas, aman, jujur tentang ketidakpastian, serta berguna untuk pembelajaran umum.
Struktur
Setiap baris JSONL memiliki kolom berikut:
Kolom
Tipe
Keterangan
id
string
Identitas unik contoh
instruction
string
Permintaan… See the full description on the dataset page: https://huggingface.co/datasets/ojiwzrd/threew.three-mountain-scaling
ThreeMountain_Scaling
Segment
Meaning
GO
Geometric Object — indicates the object type used (e.g., GO for geometric, RO for real objects).
L / Arc
Object Arrangement — defines how objects are arranged spatially. L means L-shape arrangement; Arc means objects are placed in an arc.
RC
Random Character Position — RC = True: character position is randomized.
FC
Fixed Character Position — FC = True: character stays fixed.
RS
Random Scale — RS = True: objects are… See the full description on the dataset page: https://huggingface.co/datasets/grow-ai-like-a-child/three-mountain-scaling.
