CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /dcvlm-balanced-200b DCVLM-Balanced (200B tokens) DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool. The instruction-heavy counterpart (DCVLM-baseline) is available as dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.imageimage-text-to-text10K<n<100K1 likes31k downloads2mo agoHugging Face02pkavumba /balanced-copa Dataset Card for "Balanced COPA" Dataset Summary Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/pkavumba/balanced-copa.tabularquestion-answering1K<n<10K4 likes9.5k downloads4y agoHugging Face03RatnambarBaghel /crop-disease-balanced-5022image1K<n<10K0 likes4.4k downloads2mo agoHugging Face04well-balanced /cantabile-runs cantabile-runs Work queue and checkpoint store for the Cantabile dynamics study. The directory tree is the plan — there is no plan file and no database. main/<song>/<method>/.gitkeep queued, unclaimed main/<song>/<method>/<seed>/CLAIM-<worker> a worker holds it (mtime = heartbeat) main/<song>/<method>/<seed>/*.pt done: 5M / 6M / 7M / 8M checkpoints main/<song>/<method>/<seed>/FAILED crashed, needs a human A worker lists main/, takes… See the full description on the dataset page: https://huggingface.co/datasets/well-balanced/cantabile-runs.tabularn<1K0 likes2.1k downloads12d agoHugging Face05weikaih /ai2thor-perspective-qa-20k-balanced-splits-with-objimage10K<n<100K0 likes1.7k downloads8mo agoHugging Face06gplsi /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes1.5k downloads9mo agoHugging Face07weikaih /ai2thor-perspective-qa-100k-balanced-training-v1-splitsimage10K<n<100K0 likes1.2k downloads11mo agoHugging Face08riot1 /lodestar-balanced-2m-neighbors-v20 likes1k downloads26d agoHugging Face09laion /emolia-thinking-balanced-buckets Emolia-Thinking — Balanced Per-Dimension Bucket Subset A balanced, per-dimension bucket subset of VoiceNet/emolia-thinking, derived from that dataset's zero-shot VoiceNet-dimension labels. For every VoiceNet voice/prosody/timbre/style dimension, this subset draws a roughly equal number of clips from each ordinal bucket (0–6), so that downstream training / probing sees a balanced distribution along each axis instead of the strongly skewed natural distribution. How… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-thinking-balanced-buckets.audioaudio-classification100K<n<1M0 likes968 downloads2mo agoHugging Face10chikuwa-AI /AudioSet_balanced_videovideo10K<n<100K0 likes775 downloads2mo agoHugging Face11BSC-LT /open_data_26B_tokens_balanced_es_caThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.0 likes637 downloads3y agoHugging Face12alea-institute /kl3m-data-sample-005-balancedtext1M<n<10M1 likes628 downloads10mo agoHugging Face13gplsi /fake_job_postings_balanced_va 🧠 BALANCED_FAKE_JOB_POSTINGS_VA Dataset 📘 Overview This dataset is a manually translated and balanced Valencian version of the original Fake Job Postings dataset from Kaggle:Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.All text fields (e.g., job title, company profile, description, requirements) have been manually translated into Valencian, preserving the semantic… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_va.text-classification1K<n<10K0 likes612 downloads9mo agoHugging Face14serenityyyyy /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes551 downloads6mo agoHugging Face15wandb /ragbench-sentence-relevance-balancedtext100K<n<1M1 likes549 downloads2y agoHugging Face16weikaih /procthor-100-counting-balancedimagen<1K0 likes503 downloads11mo agoHugging Face17nirmalendu01 /abir177m-pretrain-balanced20-ezhijaru abir177m pretrain mix — balanced20 en/zh/hi/ja/ru Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining. Languages: 20% each en, zh, hi, ja, ru Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru) Tokenizer: mistralai/Mistral-Nemo-Base-2407 Packing: 2048-token causal LM blocks (input_ids, labels identical) Target budget: 3.55B tokens (1,733k sequences) See meta.json for exact mixture + dataset map + seed. text1M<n<10M0 likes486 downloads1mo agoHugging Face18osazuwa /2d_dungeon_flier_video_balanced 2D Dungeon Flier Video: Balanced Causal Splits This dataset is a split-safe, balanced augmentation of osazuwa/2d_dungeon_flier_video. It reuses all 10,000 source episodes exactly once and adds 3,100 episodes from the same simulator. There is no clip overlap across splits. Each episode is a 14-second MP4 with 140 frames at 10 FPS and a stored resolution of 900 x 540 pixels. Matching NPZ files contain the nine-variable causal trace, action tokens, and intervention encoding. Every… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/2d_dungeon_flier_video_balanced.tabular10K<n<100K0 likes472 downloads28d agoHugging Face19bhheo /nvidia_open_reasoning_balanced_100k nvidia_open_reasoning_balanced_100k A domain-balanced 100k reasoning SFT dataset built from three NVIDIA Open Reasoning datasets: 33,333 examples each for math, code, and science (99,999 total). Each example is a single-turn conversation with a full reasoning trace: conversations: [ {"from": "human", "value": "<problem>"}, {"from": "gpt", "value": "<think>\n<reasoning trace>\n</think><final solution>"} ] Columns column description conversations… See the full description on the dataset page: https://huggingface.co/datasets/bhheo/nvidia_open_reasoning_balanced_100k.texttext-generation10K<n<100K0 likes439 downloads2mo agoHugging Face20mazkooleg /digit_mask_ensemble_distilled_from_cv12_balanced_mfcc Dataset Card for "digit_mask_ensemble_distilled_from_cv12_balanced_mfcc" More Information needed timeseries10M<n<100M0 likes431 downloads3y agoHugging Face21Voxel51 /deeplesion-balanced-2k DeepLesion Benchmark Subset (Balanced 2K) This dataset is a curated subset of the DeepLesion dataset, prepared for demonstration and benchmarking purposes. It consists of 2,000 CT lesion samples, balanced across 8 coarse lesion types, and filtered to include lesions with a short diameter > 10mm. Dataset Details Source: DeepLesion Institution: National Institutes of Health (NIH) Clinical Center Subset size: 2,000 images Lesion types: lung, abdomen, mediastinum, liver… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/deeplesion-balanced-2k.imageobject-detection1K<n<10K0 likes403 downloads1y agoHugging Face22Milana /vctk_resampled_16k_balancedaudio10K<n<100K0 likes377 downloads2y agoHugging Face23Evangelinejy /chess-train-data-balanced-tokenized0 likes374 downloads7mo agoHugging Face24medarc /gtex-10M-balanced-tiles GTEx 10M Balanced Tiles This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed. Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.tabularimage-feature-extraction10M<n<100M0 likes368 downloads3mo agoHugging Face25AbdullahImran /balanced_wildfire_datasetimageimage-classification1K<n<10K0 likes365 downloads18d agoHugging Face26greentechapps /everyayah_curated_1s_20s_balancedaudio10K<n<100K0 likes315 downloads1y agoHugging Face27weikaih /ai2thor-perspective-qa-800-balanced-val-v1imagen<1K0 likes304 downloads9mo agoHugging Face28weikaih /ai2thor-perspective-qa-balanced-400-v2imagen<1K0 likes303 downloads11mo agoHugging Face29open-athena /rl__24GPU_base__mix_h2_language_balanced__r2egym-nl2bash-stacktext10K<n<100K0 likes299 downloads6mo agoHugging Face30kevin510 /swm-ogb-recipe-balanced swm-ogb-recipe-balanced OGBench visual-cube-quadruple training set for SWM-Next world-model experiments. 600 episodes, balanced mixture: 300 expert / 150 noisy / 150 play (50% expert, matching SWM's recipe). Collected with SWM's generation scripts, env seeds 0-299, 200 steps/episode, rendered 768x768, stored as JPEG-448. Layout path contents <type>__<n>.pt one episode: frames (list of JPEG bytes), actions (T x 5 fp32), proprio manifest.json episode… See the full description on the dataset page: https://huggingface.co/datasets/kevin510/swm-ogb-recipe-balanced.0 likes297 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.