datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hifitts2-aligned
HiFiTTS-2 word alignments
Word-level forced alignments for the HiFiTTS-2
corpus (44 kHz subset, resampled to 24 kHz), as used to train
pocket-tts models.
Like HiFiTTS-2 itself, this dataset contains no audio — only pointers and
annotations. The audio is downloaded from LibriVox and cut locally.
Contents
train/train_aligned-*.jsonl.gz — the full aligned training manifest
eval_aligned.jsonl.gz — a 1000-utterance held-out split
scripts/download_audio.py — fetches… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/hifitts2-aligned.asynchow-code-aligned-minutes
AsynChow Code-Aligned Minutes
This dataset is a unit-normalized variant of the AsynChow data released with
fangru-lin/procedure_generalization_llm,
pinned to source commit d9bf3485cd41c1050d33471d922c826f474efec1.
It contains three aligned representations of each weighted DAG scheduling
problem:
natural: natural-language steps and precedence constraints;
graph: adjacency-list and duration-dictionary representation;
python: executable-style Python representation from the… See the full description on the dataset page: https://huggingface.co/datasets/PTTREP/asynchow-code-aligned-minutes.TACTBench-Samples
TACTBench Demonstration Samples
This repository contains five full-context demonstration examples from
TACTBench. It does not contain the TACT training set or the remaining hidden
TACTBench evaluation set. The samples use the same full-history representation
as the benchmark evaluation and illustrate direct correction, error
explanation, guided revision, clarification checking, affective feedback, and
retry elicitation.
Data
data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.socialjax-harvest-frame-aligned-512
SocialJax Harvest frame-aligned dynamics dataset
Compact tokenizer-code dataset for training action-conditioned SocialJax Harvest dynamics models.
Repository: ParoleLM/socialjax-harvest-frame-aligned-512
Format: frame_aligned_socialjax_dynamics_v3
Tokenizer codes per frame: 88
Codebook size: 512
Maximum agents: 7
Total size: 4.385 GB
Train: 9,720 rollouts, 19,916,280 frames
Validation: 540 rollouts, 1,106,460 frames
Test: 540 rollouts, 1,106,460 frames
Files… See the full description on the dataset page: https://huggingface.co/datasets/ParoleLM/socialjax-harvest-frame-aligned-512.tulu_delta-learning_Qwen2.5-3B-1.5B_reward-alignedtulu_delta-learning_3B-1.5B_reward-alignedpusht_norm4_stopreq_plain_aligned100k
PushT Plain Stopreq Aligned To CoT 100k
Plain records are selected from /data/home/jiaxin/unified_world_model/data/pusht_96_norm4_visual_nomarker_data/data by the CoT source_record manifest.
Images and actions are unchanged; only the full prompt receives the stop-required line.
delta-Qwen2.5-3B-vs-1.5B-reward_alignedpusht_norm4_stopreq_cot_aligned100k
PushT CoT Stopreq Candidate Shuffle Aligned 100k
Built from /data/home/raychai/hf_datasets/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot_stopreq_candidate_shuffle_20260604_004542 using manifest pusht_stopreq_plain_cot_aligned100k_v1.
Training rows are the first 100000 rows of the rewritten CoT source, resharded into 8 files for BAGEL_JSONL_STREAMING parity with the plain run.
legacy_pusht_norm4_allstep_cot_stopreq_aligned100k_ordered
PushT CoT Stopreq Candidate Shuffle Aligned 100k
Built from /data/home/raychai/hf_datasets/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot_stopreq_candidate_shuffle_20260604_004542 using manifest pusht_stopreq_plain_cot_aligned100k_v1.
Training rows are the first 100000 rows of the rewritten CoT source, resharded into 8 files for BAGEL_JSONL_STREAMING parity with the plain run.
dextr-aligned-bboxestransformers-en-ko-aligned-docs
Transformers EN-KO Aligned Docs
This dataset package contains English-Korean aligned text pairs derived from the docs/source/en and docs/source/ko trees in huggingface/transformers.
Repository layout
data/: published dataset splits only
metadata/: filtering, blacklist, and build-status artifacts
docs/: agent harness and dataset construction notes
AGENTS.md: short Codex entry point for this dataset repo
Contents
data/train.jsonl: final training split with… See the full description on the dataset page: https://huggingface.co/datasets/jmj-minju/transformers-en-ko-aligned-docs.
