CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pkavumba /balanced-copa Dataset Card for "Balanced COPA" Dataset Summary Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/pkavumba/balanced-copa.tabularquestion-answering1K<n<10K4 likes9.4k downloads4y agoHugging Face02well-balanced /cantabile-runs cantabile-runs Work queue and checkpoint store for the Cantabile dynamics study. The directory tree is the plan — there is no plan file and no database. main/<song>/<method>/.gitkeep queued, unclaimed main/<song>/<method>/<seed>/CLAIM-<worker> a worker holds it (mtime = heartbeat) main/<song>/<method>/<seed>/*.pt done: 5M / 6M / 7M / 8M checkpoints main/<song>/<method>/<seed>/FAILED crashed, needs a human A worker lists main/, takes… See the full description on the dataset page: https://huggingface.co/datasets/well-balanced/cantabile-runs.tabularn<1K0 likes2.1k downloads15d agoHugging Face03gplsi /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes1.4k downloads9mo agoHugging Face04serenityyyyy /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes507 downloads6mo agoHugging Face05osazuwa /2d_dungeon_flier_video_balanced 2D Dungeon Flier Video: Balanced Causal Splits This dataset is a split-safe, balanced augmentation of osazuwa/2d_dungeon_flier_video. It reuses all 10,000 source episodes exactly once and adds 3,100 episodes from the same simulator. There is no clip overlap across splits. Each episode is a 14-second MP4 with 140 frames at 10 FPS and a stored resolution of 900 x 540 pixels. Matching NPZ files contain the nine-variable causal trace, action tokens, and intervention encoding. Every… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/2d_dungeon_flier_video_balanced.tabular10K<n<100K0 likes469 downloads1mo agoHugging Face06medarc /gtex-10M-balanced-tiles GTEx 10M Balanced Tiles This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed. Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.tabularimage-feature-extraction10M<n<100M0 likes368 downloads3mo agoHugging Face07lyl472324464 /twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aloha", "total_episodes": 300, "total_frames": 100000, "total_tasks": 33245, "chunks_size": 1000, "data_files_size_in_mb": 300, "video_files_size_in_mb": 200, "fps": 50, "splits": { "train": "0:300" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50.tabularrobotics100K<n<1M0 likes245 downloads5mo agoHugging Face08bouchonnn /airbus-balanced-subsetimage10K<n<100K0 likes244 downloads10mo agoHugging Face09kshitijthakkar /nemotron-sft-balanced-2b-v1 Nemotron SFT Dataset Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Statistics Total Samples: 200,000 Total Tokens: 1,252,287,904 Average Tokens per Sample: 6261.4 Tokenizer: Qwen/Qwen3-0.6B Random Seed: 42 Strategy: balanced Subset Distribution Subset Samples Tokens Target Completion Avg Tokens/Sample Stage-1/math 20,000 151,546,125 20,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-2b-v1.tabular100K<n<1M0 likes240 downloads7mo agoHugging Face10zuzannad1 /balanced-copa-explanations Dataset Card for "Balanced COPA" Dataset Summary Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/zuzannad1/balanced-copa-explanations.tabularquestion-answering1K<n<10K0 likes240 downloads7mo agoHugging Face11ajaysri /route_subtasks_dual_overhead_pi05_448_dagger_interventions_balanced Route Subtasks + Balanced DAgger Interventions, 448px This is a new LeRobot v2.1 dataset derived from ajaysri/route_subtasks_dual_overhead_pi05_448 and a subsequent DAgger collection. The original dataset is not modified. The base contributes 495 episodes and 128,742 frames. Only frames recorded while the human collector was actively intervening are added; autonomous policy-control frames and policy_target_action are excluded from the training targets. The DAgger action column… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/route_subtasks_dual_overhead_pi05_448_dagger_interventions_balanced.tabularrobotics100K<n<1M0 likes217 downloads2mo agoHugging Face12Annanay /aml_song_lyrics_balancedtabular10K<n<100K4 likes196 downloads4y agoHugging Face13songlab /genomes-brassicales-balanced-v1More info: https://github.com/songlab-cal/gpn tabular1M<n<10M0 likes184 downloads3y agoHugging Face14kshitijthakkar /nemotron-sft-balanced-stage1-2 Nemotron SFT Dataset Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Statistics Total Samples: 100,000 Total Tokens: 624,101,275 Average Tokens per Sample: 6241.0 Tokenizer: Qwen/Qwen3-0.6B Random Seed: 42 Strategy: balanced Subset Distribution Subset Samples Tokens Target Completion Avg Tokens/Sample Stage-1/math 10,000 75,402,505 10,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-stage1-2.tabular100K<n<1M0 likes145 downloads9mo agoHugging Face15rol09 /ex2-final-ee-z-4x-balancedtabular100K<n<1M0 likes145 downloads4mo agoHugging Face16Encrux /peg_transfer_lerobot_balancedtabular10K<n<100K0 likes141 downloads5mo agoHugging Face17KaiZehaoGe /so101_grape_place_box_supplement_2_balanced_releasetabular10K<n<100K1 likes137 downloads3mo agoHugging Face18cs330 /minH_1_maxH_4_maxW_100_holdUniq_1000_regex_None_pSub_0.15_maxHeld_10000_balance_True_inter_Truetabular100K<n<1M0 likes124 downloads3y agoHugging Face19edbeeching /vepqa-gemini-1k-correct-balanced-20tabularn<1K0 likes105 downloads29d agoHugging Face20electricsheepafrica /africa-food-balances-2010-protein-supply-quantity-g-capita-day Food Balances (2010-) — Protein supply quantity (g/capita/day) | Africa (FAOSTAT) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-food-balances-2010-protein-supply-quantity-g-capita-day.tabulartabular-classification10K<n<100K0 likes92 downloads1mo agoHugging Face21axiboai /piper_stacking_abc_mix_balancedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "piperx_bimanual", "total_episodes": 480, "total_frames": 641358, "total_tasks": 9, "total_videos": 1440, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:480" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/axiboai/piper_stacking_abc_mix_balanced.tabularrobotics100K<n<1M0 likes86 downloads3mo agoHugging Face22cs330 /minH_1_maxH_5_maxW_100_holdUniq_1000_regex_None_pSub_0.15_maxHeld_10000_balance_True_inter_Truetabular100K<n<1M0 likes82 downloads3y agoHugging Face23cs330 /minH_1_maxH_4_maxW_100_holdUniq_1000_regex_None_pSub_0.15_maxHeld_none_balance_Truetabular100K<n<1M0 likes69 downloads3y agoHugging Face24selfcorrexp /llama3_additional_rr80k_NON_balanced_sfttabular100K<n<1M0 likes63 downloads2y agoHugging Face25Long27 /so101_mix4_clean_balanced_v1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.state": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [… See the full description on the dataset page: https://huggingface.co/datasets/Long27/so101_mix4_clean_balanced_v1.tabularrobotics100K<n<1M0 likes62 downloads6d agoHugging Face26werty1248 /multilingual-instruct-balancedThis repository is a collection of English, Korean, Chinese, and Japanese datasets collected by the HuggingFace Hub and transformed into a unified format. It consists of either native or synthetic data. Some data is not clearly copyrighted or only allows non-commercial use. Preprocessing: I removed data with too few answer tokens or more than 8192 tokens, and removed synthetic data with repetitions. Balancing: I randomly sampled a subset of the data with different weights for each language and… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/multilingual-instruct-balanced.tabulartext-generation1M<n<10M2 likes60 downloads2y agoHugging Face27tcapelle /ToxicCommons-balancedtabular1M<n<10M0 likes59 downloads2y agoHugging Face28DrinkIcedT /mbti_balanced_pub_newtabular100K<n<1M0 likes59 downloads21d agoHugging Face29Kymera-Solutions /train_names_balanced WA Voter Names — balanced train split 1:1 downsampled training split for binary name classification, built from the Washington State voter registration database (VRDB) extract dated 2026-09-01. Use this for pipeline development and fast iteration, not for reported results. Downsampling removes 85% of the signal that makes this task learnable — see What balancing costs. Restricted data — see Access and legal restrictions. This repository is not intended to be public.… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_balanced.tabulartext-classification100K<n<1M0 likes59 downloads16d agoHugging Face30selfcorrexp /llama3_additional_rr40k_NON_balanced_sfttabular100K<n<1M0 likes57 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.