datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2d_dungeon_flier_video_balanced
2D Dungeon Flier Video: Balanced Causal Splits
This dataset is a split-safe, balanced augmentation of osazuwa/2d_dungeon_flier_video. It reuses all 10,000 source episodes exactly once and adds 3,100 episodes from the same simulator. There is no clip overlap across splits.
Each episode is a 14-second MP4 with 140 frames at 10 FPS and a stored resolution of 900 x 540 pixels. Matching NPZ files contain the nine-variable causal trace, action tokens, and intervention encoding.
Every… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/2d_dungeon_flier_video_balanced.balanced-copa-explanations
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/zuzannad1/balanced-copa-explanations.genomes-brassicales-balanced-v1More info: https://github.com/songlab-cal/gpn
toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001
ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 001
This dataset repo records the exact local training-data state visible to the dynamic epoch launcher.
It intentionally stores manifests and audit records rather than duplicating large Parquet shards.
Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_001_special_structure_current_step_002000.pt
Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT
Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001.toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002
ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 002
This dataset repo records the exact local training-data state visible to the dynamic epoch launcher.
It intentionally stores manifests and audit records rather than duplicating large Parquet shards.
Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_002_special_structure_delta_step_002750.pt
Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT
Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002.pusht_eval_far_balanced100_20260810
PushT Far — balanced 100
This is the exact 100-episode PushT evaluation file used by the final UWM trajectories.
Split: far
Definition: Strict spatial out-of-distribution evaluation set.
Records: 100 with 100 unique episode seeds
Exact compressed JSONL SHA-256: 84d35787ceb2781dfda45bb2562dfe00b44aebe3af7bb42b01b8f4f7c8f81eac
Trajectories: https://huggingface.co/datasets/novastar111/uwm_paper_eval_trajectories/tree/main/table1/pusht/far
manifest.json records every episode seed… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pusht_eval_far_balanced100_20260810.csfd_sentiment_balanced
Introduction
This is a class-balanced version of subset from CSFD sentiment dataset collected from CSFD. The work was originaly published in paper Sentiment Analysis in Czech Social Media Using Supervised Machine Learning (citation below).
Format
This dataset uses the following jsonl format:
{
"id": unique identifier,
"query": text for sentiment analysis,
"choices": sentiment classes ["negativní","neutrální","pozitivní"],
"gold": index of gold class
}… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/csfd_sentiment_balanced.mall_sentiment_balanced
Introduction
This is a class-balanced version of subset from CSFD sentiment dataset collected from MALL (mall.cz). The work was originaly published in paper Sentiment Analysis in Czech Social Media Using Supervised Machine Learning (citation below).
Format
This dataset uses the following jsonl format:
{
"id": unique identifier,
"query": text for sentiment analysis,
"choices": sentiment classes ["negativní","neutrální","pozitivní"],
"gold": index of gold class… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/mall_sentiment_balanced.fb_sentiment_balanced
Introduction
This is a class-balanced version of subset from CSFD sentiment dataset collected from Facebook. The work was originaly published in paper Sentiment Analysis in Czech Social Media Using Supervised Machine Learning (citation below).
Format
This dataset uses the following jsonl format:
{
"id": unique identifier,
"query": text for sentiment analysis,
"choices": sentiment classes ["negativní","neutrální","pozitivní"],
"gold": index of gold class
}… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/fb_sentiment_balanced.pusht_eval_mid_balanced100_20260810
PushT Mid — balanced 100
This is the exact 100-episode PushT evaluation file used by the final UWM trajectories.
Split: mid
Definition: Mid-distance diagnostic set whose spatial support overlaps training.
Records: 100 with 100 unique episode seeds
Exact compressed JSONL SHA-256: ba6007318cdf55e69e57cb09680c3d5e2bbd2950eb8cda1cdde0c13815a0e33b
Trajectories: https://huggingface.co/datasets/novastar111/uwm_paper_eval_trajectories/tree/main/table1/pusht/mid
manifest.json records… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pusht_eval_mid_balanced100_20260810.balanced_labelspusht_eval_in_dist_balanced100_20260820
PushT In-distribution — balanced 100
This is the exact 100-episode PushT evaluation file used by the final UWM trajectories.
Split: in_dist
Definition: Training-space evaluation with exact initial-state disjointness.
Records: 100 with 100 unique episode seeds
Exact compressed JSONL SHA-256: 3b3be6ae9c2a95e14590b5148fe155a3d5524f8b277a5b15945ac596cfc48cfe
Trajectories: https://huggingface.co/datasets/novastar111/uwm_paper_eval_trajectories/tree/main/table1/pusht/in_dist… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pusht_eval_in_dist_balanced100_20260820.gpn_original_balanced_v1cot-oracle-reasoning-termination-balancedpusht_eval_near_balanced100_20260810
PushT Near — balanced 100
This is the exact 100-episode PushT evaluation file used by the final UWM trajectories.
Split: near
Definition: Near-start diagnostic set whose spatial support overlaps training.
Records: 100 with 100 unique episode seeds
Exact compressed JSONL SHA-256: 18b800c78a3372321e6850a4421951eb7160d9d40ff6f890032d73a3d0bc1fc9
Trajectories: https://huggingface.co/datasets/novastar111/uwm_paper_eval_trajectories/tree/main/table1/pusht/near
manifest.json records… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pusht_eval_near_balanced100_20260810.gpn_grass_balanced_v1balanced_labels_2mgenomes-mammals-balanced-v1balancedgenomes-mammals-balanced-v1-1024SFT_zeroshot_Gemma27B_balanced_plaingpn_combined_balanced_v1
