datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datakit-tier2-skewed-v2
Datakit Tier2 Skewed Synthetic
A synthetic, heavy-tailed-document dataset generated for stress-testing the
Marin datakit pipeline (normalize / minhash / fuzzy_dups / consolidate /
tokenize) against doc-length outliers up to 256 MB.
This is a CI / pipeline-stress dataset, not a training corpus. It exists
to exercise long-tail code paths that the FineWeb-Edu smoke ferry doesn't
cover. Don't use it for training without re-evaluating its distribution.
Provenance
Source… See the full description on the dataset page: https://huggingface.co/datasets/ravwojdyla/datakit-tier2-skewed-v2.tier2_writertier2-autolabeled-lastMultiChallenge-Tier2-AdvanceMultiChallenge-Tier2-Coretier2_vscodealoha_tier2_test1threadThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 4,
"total_frames": 528,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jgiegold/aloha_tier2_test1thread.aloha_tier2_diagThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 2,
"total_frames": 264,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jgiegold/aloha_tier2_diag.aloha_balanced_v4_tier2_fullpaper_tier2multiple-search-tier2
MultipleSearch Tier 2 (Balanced + Strict Ambiguous Filtering)
Origin
Derived from Tier 1 outputs.
Reuses Tier 1 passage and embedding artifacts.
Additional Tier-2 Filtering
From Tier 1 query records:
Unambiguous queries are kept only if their cluster has at least one gold passage.
Ambiguous queries are kept only if every answer cluster has at least one gold passage.
After filtering:
Downsample to a strict 50/50 single-vs-ambiguous balance.
Shuffle with… See the full description on the dataset page: https://huggingface.co/datasets/MJ141592/multiple-search-tier2.
