CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /fineweb_edu_100BT-shuffled FineWeb-Edu 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.tabular100M<n<1B6 likes4.3k downloads7mo agoHugging Face02HuggingFaceFW /dclm_100BT-shuffled DCLM 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/dclm_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as dclm_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT-shuffled.tabular10M<n<100M3 likes1.4k downloads7mo agoHugging Face03HuggingFaceFW /fineweb_100BT-shuffled FineWeb 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_100BT-shuffled.tabular100M<n<1B0 likes772 downloads7mo agoHugging Face04medarc /TCGA-12K-parquet-shuffled TCGA-12K Parquet (Shuffled) Attribution This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet-shuffled.tabular10M<n<100M0 likes595 downloads10mo agoHugging Face05Haesteining /ShuffledDataset2tabular1M<n<10M0 likes410 downloads2y agoHugging Face06hypnopump /smiles_transformer_shuffledtabular100M<n<1B0 likes404 downloads1y agoHugging Face07aklein4 /fineweb-edu-sample-10BT-shuffled 📚 FineWeb-Edu (Shuffled) The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves. This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu. Shuffling was performed using the following script: import datasets data = datasets.load_dataset( "HuggingFaceFW/fineweb-edu", "sample-10BT", split="train", streaming=False, ) data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.tabulartext-generation1M<n<10M1 likes396 downloads1y agoHugging Face08TheFinAI /00_filtered_shuffledtabular10M<n<100M0 likes348 downloads4mo agoHugging Face09another-phytophile /153-angiosperm-species-32k-sequences-shuffledtabular1M<n<10M0 likes317 downloads5mo agoHugging Face10konwoo /dclm-16m-shuffledtabular10M<n<100M0 likes279 downloads9mo agoHugging Face11konwoo /dclm-3m-shuffledtabular1M<n<10M0 likes256 downloads9mo agoHugging Face12EleutherAI /transformer-reasoning-bios-dataset-25000_shuffledtabular10M<n<100M0 likes217 downloads2y agoHugging Face13ching-goodfire /MAPS-ClinVar-VKS-Embeddings-L80-shuffled MAPS ClinVar/VKS ESM-C layer-80 difference fields, shuffled layout The same 200,913 rows as ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80, rewritten in one random order so that a prefix is a sample. The source repo is sorted by protein accession. That makes the first N rows a block of related proteins rather than a sample of the pool, so any subset that is actually representative requires reading all 154.01 GB and selecting rows afterwards. Here the rows are stored shuffled, so… See the full description on the dataset page: https://huggingface.co/datasets/ching-goodfire/MAPS-ClinVar-VKS-Embeddings-L80-shuffled.tabular100K<n<1M0 likes157 downloads2mo agoHugging Face14Faless /hedge_v2_sim_real_30hz_shuffledThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "piper_full", "total_episodes": 899, "total_frames": 1437819, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:899" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Faless/hedge_v2_sim_real_30hz_shuffled.tabularrobotics1M<n<10M0 likes131 downloads7d agoHugging Face15konwoo /dclm-val-1.64m-shuffledtabular1M<n<10M0 likes129 downloads9mo agoHugging Face16Alignment-Lab-AI /synthetic-bn-shuffledvalstabular1M<n<10M0 likes95 downloads1y agoHugging Face17konwoo /dclm-train-1M-shuffledtabular1M<n<10M0 likes93 downloads10mo agoHugging Face18Faless /hedge_v2_sim_real_20260918_30hz_shuffledThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "piper_full", "total_episodes": 908, "total_frames": 1440029, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:908" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Faless/hedge_v2_sim_real_20260918_30hz_shuffled.tabularrobotics1M<n<10M0 likes93 downloads6d agoHugging Face19EleutherAI /transformer-reasoning-bios-dataset-10000_shuffledtabular10M<n<100M0 likes84 downloads2y agoHugging Face20Yujivus /wmt14-de-en-helsinki-filtered-shuffled-128-v3tabular1M<n<10M0 likes81 downloads1y agoHugging Face21konwoo /dclm-1001000-shuffledtabular1M<n<10M0 likes79 downloads10mo agoHugging Face22Screener2 /V4_pnp_combined_shuffledThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Screener2/V4_pnp_combined_shuffled.tabularrobotics10K<n<100K0 likes74 downloads2mo agoHugging Face23chejames /nucleosome-condensability-shuffled-split-occupancytabular1M<n<10M0 likes67 downloads2mo agoHugging Face24gaeunseo /all_data_for_first_finetuning_shuffled Dataset Card for "all_data_for_first_finetuning_shuffled" More Information needed tabular100K<n<1M0 likes50 downloads3y agoHugging Face25TheFinAI /dolma3_300B_sample_shuffled dolma3_300B_sample_shuffled Global row-level shuffle of TheFinAI/dolma3_300B_sample. Source data uses per-row Bernoulli sampling (p ≈ 0.0506) from allenai/dolma3_mix-6T-1025-7B to produce ~300B cl100k tokens preserving the original Dolma3 mix ratios. However the source parquets cluster records by sub-source on disk (each ~100K-row parquet groups rows from the same input shard contiguously), which means a small training shuffle buffer would see a non-uniform source mix per… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample_shuffled.tabulartext-generation100M<n<1B0 likes46 downloads4mo agoHugging Face26chejames /nucleosome-condensability-shuffled-splittabular1M<n<10M0 likes40 downloads3mo agoHugging Face27sergiov2000 /eval_act_so100_legoball_shuffled_lego4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 875, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sergiov2000/eval_act_so100_legoball_shuffled_lego4.tabularroboticsn<1K0 likes37 downloads1y agoHugging Face28PJMixers-Dev /HuggingFaceFW_finewiki-en-shuffledtabular1M<n<10M0 likes35 downloads11mo agoHugging Face29sergiov2000 /eval_act_so100_legoball_shuffled_lego5This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 880, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sergiov2000/eval_act_so100_legoball_shuffled_lego5.tabularroboticsn<1K0 likes33 downloads1y agoHugging Face30sergiov2000 /eval_act_so100_legoball_shuffled_lego6This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 878, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sergiov2000/eval_act_so100_legoball_shuffled_lego6.tabularroboticsn<1K0 likes33 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.