datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dcvlm-balanced-200b
DCVLM-Balanced (200B tokens)
DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool.
The instruction-heavy counterpart (DCVLM-baseline) is available as
dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.balanced-copa
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/pkavumba/balanced-copa.crop-disease-balanced-5022cantabile-runs
cantabile-runs
Work queue and checkpoint store for the Cantabile dynamics study. The directory tree
is the plan — there is no plan file and no database.
main/<song>/<method>/.gitkeep queued, unclaimed
main/<song>/<method>/<seed>/CLAIM-<worker> a worker holds it (mtime = heartbeat)
main/<song>/<method>/<seed>/*.pt done: 5M / 6M / 7M / 8M checkpoints
main/<song>/<method>/<seed>/FAILED crashed, needs a human
A worker lists main/, takes… See the full description on the dataset page: https://huggingface.co/datasets/well-balanced/cantabile-runs.ai2thor-perspective-qa-20k-balanced-splits-with-objfake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.ai2thor-perspective-qa-100k-balanced-training-v1-splitslodestar-balanced-2m-neighbors-v2emolia-thinking-balanced-buckets
Emolia-Thinking — Balanced Per-Dimension Bucket Subset
A balanced, per-dimension bucket subset of
VoiceNet/emolia-thinking,
derived from that dataset's zero-shot VoiceNet-dimension labels.
For every VoiceNet voice/prosody/timbre/style dimension, this subset draws a
roughly equal number of clips from each ordinal bucket (0–6), so that
downstream training / probing sees a balanced distribution along each axis
instead of the strongly skewed natural distribution.
How… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-thinking-balanced-buckets.AudioSet_balanced_videoopen_data_26B_tokens_balanced_es_caThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.kl3m-data-sample-005-balancedfake_job_postings_balanced_va
🧠 BALANCED_FAKE_JOB_POSTINGS_VA Dataset
📘 Overview
This dataset is a manually translated and balanced Valencian version of the original Fake Job Postings dataset from Kaggle:Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.All text fields (e.g., job title, company profile, description, requirements) have been manually translated into Valencian, preserving the semantic… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_va.fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.ragbench-sentence-relevance-balancedprocthor-100-counting-balancedabir177m-pretrain-balanced20-ezhijaru
abir177m pretrain mix — balanced20 en/zh/hi/ja/ru
Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining.
Languages: 20% each en, zh, hi, ja, ru
Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru)
Tokenizer: mistralai/Mistral-Nemo-Base-2407
Packing: 2048-token causal LM blocks (input_ids, labels identical)
Target budget: 3.55B tokens (1,733k sequences)
See meta.json for exact mixture + dataset map + seed.
2d_dungeon_flier_video_balanced
2D Dungeon Flier Video: Balanced Causal Splits
This dataset is a split-safe, balanced augmentation of osazuwa/2d_dungeon_flier_video. It reuses all 10,000 source episodes exactly once and adds 3,100 episodes from the same simulator. There is no clip overlap across splits.
Each episode is a 14-second MP4 with 140 frames at 10 FPS and a stored resolution of 900 x 540 pixels. Matching NPZ files contain the nine-variable causal trace, action tokens, and intervention encoding.
Every… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/2d_dungeon_flier_video_balanced.nvidia_open_reasoning_balanced_100k
nvidia_open_reasoning_balanced_100k
A domain-balanced 100k reasoning SFT dataset built from three NVIDIA Open Reasoning
datasets: 33,333 examples each for math, code, and science (99,999 total).
Each example is a single-turn conversation with a full reasoning trace:
conversations: [
{"from": "human", "value": "<problem>"},
{"from": "gpt", "value": "<think>\n<reasoning trace>\n</think><final solution>"}
]
Columns
column
description
conversations… See the full description on the dataset page: https://huggingface.co/datasets/bhheo/nvidia_open_reasoning_balanced_100k.digit_mask_ensemble_distilled_from_cv12_balanced_mfcc
Dataset Card for "digit_mask_ensemble_distilled_from_cv12_balanced_mfcc"
More Information needed
deeplesion-balanced-2k
DeepLesion Benchmark Subset (Balanced 2K)
This dataset is a curated subset of the DeepLesion dataset, prepared for demonstration and benchmarking purposes. It consists of 2,000 CT lesion samples, balanced across 8 coarse lesion types, and filtered to include lesions with a short diameter > 10mm.
Dataset Details
Source: DeepLesion
Institution: National Institutes of Health (NIH) Clinical Center
Subset size: 2,000 images
Lesion types: lung, abdomen, mediastinum, liver… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/deeplesion-balanced-2k.vctk_resampled_16k_balancedchess-train-data-balanced-tokenizedgtex-10M-balanced-tiles
GTEx 10M Balanced Tiles
This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed.
Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.balanced_wildfire_dataseteveryayah_curated_1s_20s_balancedai2thor-perspective-qa-800-balanced-val-v1ai2thor-perspective-qa-balanced-400-v2rl__24GPU_base__mix_h2_language_balanced__r2egym-nl2bash-stackswm-ogb-recipe-balanced
swm-ogb-recipe-balanced
OGBench visual-cube-quadruple training set for SWM-Next world-model experiments.
600 episodes, balanced mixture: 300 expert / 150 noisy / 150 play (50% expert,
matching SWM's recipe). Collected with SWM's generation scripts, env seeds 0-299,
200 steps/episode, rendered 768x768, stored as JPEG-448.
Layout
path
contents
<type>__<n>.pt
one episode: frames (list of JPEG bytes), actions (T x 5 fp32), proprio
manifest.json
episode… See the full description on the dataset page: https://huggingface.co/datasets/kevin510/swm-ogb-recipe-balanced.
