datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb_edu_100BT-shuffled
FineWeb-Edu 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.dclm_100BT-shuffled
DCLM 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/dclm_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as dclm_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT-shuffled.fineweb_100BT-shuffled
FineWeb 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/fineweb_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as fineweb_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_100BT-shuffled.TCGA-12K-parquet-shuffled
TCGA-12K Parquet (Shuffled)
Attribution
This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet-shuffled.ShuffledDataset2smiles_transformer_shuffledfineweb-edu-sample-10BT-shuffled
📚 FineWeb-Edu (Shuffled)
The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves.
This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu.
Shuffling was performed using the following script:
import datasets
data = datasets.load_dataset(
"HuggingFaceFW/fineweb-edu",
"sample-10BT",
split="train",
streaming=False,
)
data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.00_filtered_shuffled153-angiosperm-species-32k-sequences-shuffleddclm-16m-shuffleddclm-3m-shuffledtransformer-reasoning-bios-dataset-25000_shuffledMAPS-ClinVar-VKS-Embeddings-L80-shuffled
MAPS ClinVar/VKS ESM-C layer-80 difference fields, shuffled layout
The same 200,913 rows as
ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80, rewritten in one
random order so that a prefix is a sample.
The source repo is sorted by protein accession. That makes the first N rows a
block of related proteins rather than a sample of the pool, so any subset that
is actually representative requires reading all 154.01 GB and
selecting rows afterwards. Here the rows are stored shuffled, so… See the full description on the dataset page: https://huggingface.co/datasets/ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80-shuffled.hedge_v2_sim_real_30hz_shuffledThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "piper_full",
"total_episodes": 899,
"total_frames": 1437819,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:899"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Faless/hedge_v2_sim_real_30hz_shuffled.dclm-val-1.64m-shuffledsynthetic-bn-shuffledvalsdclm-train-1M-shuffledhedge_v2_sim_real_20260918_30hz_shuffledThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "piper_full",
"total_episodes": 908,
"total_frames": 1440029,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:908"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Faless/hedge_v2_sim_real_20260918_30hz_shuffled.transformer-reasoning-bios-dataset-10000_shuffledwmt14-de-en-helsinki-filtered-shuffled-128-v3dclm-1001000-shuffledV4_pnp_combined_shuffledThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Screener2/V4_pnp_combined_shuffled.nucleosome-condensability-shuffled-split-occupancyall_data_for_first_finetuning_shuffled
Dataset Card for "all_data_for_first_finetuning_shuffled"
More Information needed
dolma3_300B_sample_shuffled
dolma3_300B_sample_shuffled
Global row-level shuffle of TheFinAI/dolma3_300B_sample.
Source data uses per-row Bernoulli sampling (p ≈ 0.0506) from
allenai/dolma3_mix-6T-1025-7B to produce ~300B cl100k tokens preserving
the original Dolma3 mix ratios. However the source parquets cluster
records by sub-source on disk (each ~100K-row parquet groups rows from the
same input shard contiguously), which means a small training shuffle
buffer would see a non-uniform source mix per… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample_shuffled.nucleosome-condensability-shuffled-spliteval_act_so100_legoball_shuffled_lego4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 875,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sergiov2000/eval_act_so100_legoball_shuffled_lego4.HuggingFaceFW_finewiki-en-shuffledeval_act_so100_legoball_shuffled_lego5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 880,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sergiov2000/eval_act_so100_legoball_shuffled_lego5.eval_act_so100_legoball_shuffled_lego6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 878,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sergiov2000/eval_act_so100_legoball_shuffled_lego6.
