datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
balanced-copa
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/pkavumba/balanced-copa.cantabile-runs
cantabile-runs
Work queue and checkpoint store for the Cantabile dynamics study. The directory tree
is the plan — there is no plan file and no database.
main/<song>/<method>/.gitkeep queued, unclaimed
main/<song>/<method>/<seed>/CLAIM-<worker> a worker holds it (mtime = heartbeat)
main/<song>/<method>/<seed>/*.pt done: 5M / 6M / 7M / 8M checkpoints
main/<song>/<method>/<seed>/FAILED crashed, needs a human
A worker lists main/, takes… See the full description on the dataset page: https://huggingface.co/datasets/well-balanced/cantabile-runs.fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.2d_dungeon_flier_video_balanced
2D Dungeon Flier Video: Balanced Causal Splits
This dataset is a split-safe, balanced augmentation of osazuwa/2d_dungeon_flier_video. It reuses all 10,000 source episodes exactly once and adds 3,100 episodes from the same simulator. There is no clip overlap across splits.
Each episode is a 14-second MP4 with 140 frames at 10 FPS and a stored resolution of 900 x 540 pixels. Matching NPZ files contain the nine-variable causal trace, action tokens, and intervention encoding.
Every… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/2d_dungeon_flier_video_balanced.gtex-10M-balanced-tiles
GTEx 10M Balanced Tiles
This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed.
Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 300,
"total_frames": 100000,
"total_tasks": 33245,
"chunks_size": 1000,
"data_files_size_in_mb": 300,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:300"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50.airbus-balanced-subsetnemotron-sft-balanced-2b-v1
Nemotron SFT Dataset
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Statistics
Total Samples: 200,000
Total Tokens: 1,252,287,904
Average Tokens per Sample: 6261.4
Tokenizer: Qwen/Qwen3-0.6B
Random Seed: 42
Strategy: balanced
Subset Distribution
Subset
Samples
Tokens
Target
Completion
Avg Tokens/Sample
Stage-1/math
20,000
151,546,125
20,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-2b-v1.balanced-copa-explanations
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/zuzannad1/balanced-copa-explanations.route_subtasks_dual_overhead_pi05_448_dagger_interventions_balanced
Route Subtasks + Balanced DAgger Interventions, 448px
This is a new LeRobot v2.1 dataset derived from
ajaysri/route_subtasks_dual_overhead_pi05_448 and a subsequent
DAgger collection. The original dataset is not modified.
The base contributes 495 episodes and 128,742 frames. Only
frames recorded while the human collector was actively intervening are added;
autonomous policy-control frames and policy_target_action are excluded from
the training targets. The DAgger action column… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/route_subtasks_dual_overhead_pi05_448_dagger_interventions_balanced.aml_song_lyrics_balancedgenomes-brassicales-balanced-v1More info: https://github.com/songlab-cal/gpn
nemotron-sft-balanced-stage1-2
Nemotron SFT Dataset
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Statistics
Total Samples: 100,000
Total Tokens: 624,101,275
Average Tokens per Sample: 6241.0
Tokenizer: Qwen/Qwen3-0.6B
Random Seed: 42
Strategy: balanced
Subset Distribution
Subset
Samples
Tokens
Target
Completion
Avg Tokens/Sample
Stage-1/math
10,000
75,402,505
10,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-stage1-2.ex2-final-ee-z-4x-balancedpeg_transfer_lerobot_balancedso101_grape_place_box_supplement_2_balanced_releaseminH_1_maxH_4_maxW_100_holdUniq_1000_regex_None_pSub_0.15_maxHeld_10000_balance_True_inter_Truevepqa-gemini-1k-correct-balanced-20africa-food-balances-2010-protein-supply-quantity-g-capita-day
Food Balances (2010-) — Protein supply quantity (g/capita/day) | Africa (FAOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-food-balances-2010-protein-supply-quantity-g-capita-day.piper_stacking_abc_mix_balancedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "piperx_bimanual",
"total_episodes": 480,
"total_frames": 641358,
"total_tasks": 9,
"total_videos": 1440,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:480"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/axiboai/piper_stacking_abc_mix_balanced.minH_1_maxH_5_maxW_100_holdUniq_1000_regex_None_pSub_0.15_maxHeld_10000_balance_True_inter_TrueminH_1_maxH_4_maxW_100_holdUniq_1000_regex_None_pSub_0.15_maxHeld_none_balance_Truellama3_additional_rr80k_NON_balanced_sftso101_mix4_clean_balanced_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/Long27/so101_mix4_clean_balanced_v1.multilingual-instruct-balancedThis repository is a collection of English, Korean, Chinese, and Japanese datasets collected by the HuggingFace Hub and transformed into a unified format. It consists of either native or synthetic data.
Some data is not clearly copyrighted or only allows non-commercial use.
Preprocessing: I removed data with too few answer tokens or more than 8192 tokens, and removed synthetic data with repetitions.
Balancing: I randomly sampled a subset of the data with different weights for each language and… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/multilingual-instruct-balanced.ToxicCommons-balancedmbti_balanced_pub_newtrain_names_balanced
WA Voter Names — balanced train split
1:1 downsampled training split for binary name classification, built from the
Washington State voter registration database (VRDB) extract dated 2026-09-01.
Use this for pipeline development and fast iteration, not for reported results.
Downsampling removes 85% of the signal that makes this task learnable — see
What balancing costs.
Restricted data — see Access and legal restrictions.
This repository is not intended to be public.… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_balanced.llama3_additional_rr40k_NON_balanced_sft
