CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /dcvlm-balanced-200b DCVLM-Balanced (200B tokens) DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool. The instruction-heavy counterpart (DCVLM-baseline) is available as dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.imageimage-text-to-text10K<n<100K1 likes30k downloads2mo agoHugging Face02weikaih /ai2thor-perspective-qa-20k-balanced-splits-with-objimage10K<n<100K0 likes1.7k downloads8mo agoHugging Face03weikaih /ai2thor-perspective-qa-100k-balanced-training-v1-splitsimage10K<n<100K0 likes1.2k downloads11mo agoHugging Face04alea-institute /kl3m-data-sample-005-balancedtext1M<n<10M1 likes632 downloads10mo agoHugging Face05wandb /ragbench-sentence-relevance-balancedtext100K<n<1M1 likes563 downloads2y agoHugging Face06mazkooleg /digit_mask_ensemble_distilled_from_cv12_balanced_mfcc Dataset Card for "digit_mask_ensemble_distilled_from_cv12_balanced_mfcc" More Information needed timeseries10M<n<100M0 likes482 downloads3y agoHugging Face07weikaih /procthor-100-counting-balancedimagen<1K0 likes476 downloads11mo agoHugging Face08Milana /vctk_resampled_16k_balancedaudio10K<n<100K0 likes454 downloads2y agoHugging Face09bhheo /nvidia_open_reasoning_balanced_100k nvidia_open_reasoning_balanced_100k A domain-balanced 100k reasoning SFT dataset built from three NVIDIA Open Reasoning datasets: 33,333 examples each for math, code, and science (99,999 total). Each example is a single-turn conversation with a full reasoning trace: conversations: [ {"from": "human", "value": "<problem>"}, {"from": "gpt", "value": "<think>\n<reasoning trace>\n</think><final solution>"} ] Columns column description conversations… See the full description on the dataset page: https://huggingface.co/datasets/bhheo/nvidia_open_reasoning_balanced_100k.texttext-generation10K<n<100K0 likes430 downloads2mo agoHugging Face10medarc /gtex-10M-balanced-tiles GTEx 10M Balanced Tiles This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed. Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.tabularimage-feature-extraction10M<n<100M0 likes368 downloads3mo agoHugging Face11mazkobot /0_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc Dataset Card for "0_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc" More Information needed timeseries1M<n<10M0 likes359 downloads3y agoHugging Face12greentechapps /everyayah_curated_1s_20s_balancedaudio10K<n<100K0 likes327 downloads1y agoHugging Face13open-athena /rl__24GPU_base__mix_h2_language_balanced__r2egym-nl2bash-stacktext10K<n<100K0 likes301 downloads7mo agoHugging Face14nirmalendu01 /abir177m-pretrain-balanced20-ezhijaru abir177m pretrain mix — balanced20 en/zh/hi/ja/ru Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining. Languages: 20% each en, zh, hi, ja, ru Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru) Tokenizer: mistralai/Mistral-Nemo-Base-2407 Packing: 2048-token causal LM blocks (input_ids, labels identical) Target budget: 3.55B tokens (1,733k sequences) See meta.json for exact mixture + dataset map + seed. text1M<n<10M0 likes299 downloads1mo agoHugging Face15weikaih /ai2thor-perspective-qa-balanced-400-v2imagen<1K0 likes290 downloads11mo agoHugging Face16weikaih /ai2thor-perspective-qa-800-balanced-val-v1imagen<1K0 likes281 downloads9mo agoHugging Face17mazkobot /1_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc Dataset Card for "1_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc" More Information needed timeseries1M<n<10M0 likes276 downloads3y agoHugging Face18greentechapps /everyayah_curated_1s_20s_balanced_largeaudio10K<n<100K0 likes261 downloads1y agoHugging Face19weikaih /ai2thor-perspective-qa-400-balanced-diverseimagen<1K0 likes261 downloads11mo agoHugging Face20weikaih /ai2thor-perspective-qa-400-balanced-v2imagen<1K0 likes256 downloads11mo agoHugging Face21dddraxxx /refchartqa-balanced-10kimage10K<n<100K0 likes254 downloads1y agoHugging Face22weikaih /ai2thor-perspective-qa-800-balanced-diverse-v1imagen<1K0 likes250 downloads9mo agoHugging Face23weikaih /ai2thor-perspective-qa-400-balanced-diverse-v2imagen<1K0 likes248 downloads11mo agoHugging Face24enyoukai /AudioSet-Strong-Balancedaudio10K<n<100K0 likes246 downloads10mo agoHugging Face25lyl472324464 /twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aloha", "total_episodes": 300, "total_frames": 100000, "total_tasks": 33245, "chunks_size": 1000, "data_files_size_in_mb": 300, "video_files_size_in_mb": 200, "fps": 50, "splits": { "train": "0:300" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50.tabularrobotics100K<n<1M0 likes245 downloads5mo agoHugging Face26bouchonnn /airbus-balanced-subsetimage10K<n<100K0 likes244 downloads10mo agoHugging Face27kshitijthakkar /nemotron-sft-balanced-2b-v1 Nemotron SFT Dataset Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Statistics Total Samples: 200,000 Total Tokens: 1,252,287,904 Average Tokens per Sample: 6261.4 Tokenizer: Qwen/Qwen3-0.6B Random Seed: 42 Strategy: balanced Subset Distribution Subset Samples Tokens Target Completion Avg Tokens/Sample Stage-1/math 20,000 151,546,125 20,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-2b-v1.tabular100K<n<1M0 likes240 downloads7mo agoHugging Face28cairocode /BALANCED_MSPP_MSPI_IEMOimage100K<n<1M0 likes227 downloads2y agoHugging Face29ajaysri /route_subtasks_dual_overhead_pi05_448_dagger_interventions_balanced Route Subtasks + Balanced DAgger Interventions, 448px This is a new LeRobot v2.1 dataset derived from ajaysri/route_subtasks_dual_overhead_pi05_448 and a subsequent DAgger collection. The original dataset is not modified. The base contributes 495 episodes and 128,742 frames. Only frames recorded while the human collector was actively intervening are added; autonomous policy-control frames and policy_target_action are excluded from the training targets. The DAgger action column… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/route_subtasks_dual_overhead_pi05_448_dagger_interventions_balanced.tabularrobotics100K<n<1M0 likes217 downloads2mo agoHugging Face30tta1301 /xray-balanced-datasetimage10K<n<100K0 likes203 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.