datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human-aligned-similarity-benchmark
Human Aligned Similarity Benchmark
You are welcome to go to alignedmachine.com to contribute.
Overview
This dataset contains human-aligned similarity judgments for embedding text and multimodal AI model evaluation. The benchmark is designed to assess how well AI models align with human cognitive preferences in similarity perception across text and image modalities.
Dataset Structure
Concept Files
This dataset contains human preference judgments for… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/human-aligned-similarity-benchmark.hifitts2-aligned
HiFiTTS-2 word alignments
Word-level forced alignments for the HiFiTTS-2
corpus (44 kHz subset, resampled to 24 kHz), as used to train
pocket-tts models.
Like HiFiTTS-2 itself, this dataset contains no audio — only pointers and
annotations. The audio is downloaded from LibriVox and cut locally.
Contents
train/train_aligned-*.jsonl.gz — the full aligned training manifest
eval_aligned.jsonl.gz — a 1000-utterance held-out split
scripts/download_audio.py — fetches… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/hifitts2-aligned.doc-aligned-crossSum-subsetunified-alignedeu-financial-regulation-aligned
EU Financial Regulation, Aligned Across 24 Languages
Eight EU financial regulations, split to the paragraph, in every official EU language,
with the alignment verified rather than assumed.
56,838 rows — 2,442 provisions × up to 24 languages.
Why the alignment is exact
Most multilingual legal corpora are aligned by matching sentences, which is approximate
and fails on exactly the long provisions people care about.
This one does not do that. EUR-Lex assigns… See the full description on the dataset page: https://huggingface.co/datasets/chenjigaram/eu-financial-regulation-aligned.pleias-post-ocr-correction-chonkie-aligned-en
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.asynchow-code-aligned-minutes
AsynChow Code-Aligned Minutes
This dataset is a unit-normalized variant of the AsynChow data released with
fangru-lin/procedure_generalization_llm,
pinned to source commit d9bf3485cd41c1050d33471d922c826f474efec1.
It contains three aligned representations of each weighted DAG scheduling
problem:
natural: natural-language steps and precedence constraints;
graph: adjacency-list and duration-dictionary representation;
python: executable-style Python representation from the… See the full description on the dataset page: https://huggingface.co/datasets/PTTREP/asynchow-code-aligned-minutes.pleias-post-ocr-correction-chonkie-aligned-fr
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.qwen3-4b-0630-tooluse-eval-aligned-r32-spare-games-envs
qwen3-4B-Instruct-0630-tooluse-eval-aligned-r32 — generated environments
Environments generated by the SPARE proposer during training run
050mlekj (qwen3-4B-Instruct-0630-tooluse-eval-aligned-r32), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
456
Steps covered
21 (step 0–448)
With recovered skill
456
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-4b-0630-tooluse-eval-aligned-r32-spare-games-envs.TACTBench-Samples
TACTBench Demonstration Samples
This repository contains five full-context demonstration examples from
TACTBench. It does not contain the TACT training set or the remaining hidden
TACTBench evaluation set. The samples use the same full-history representation
as the benchmark evaluation and illustrate direct correction, error
explanation, guided revision, clarification checking, affective feedback, and
retry elicitation.
Data
data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.socialjax-harvest-frame-aligned-512
SocialJax Harvest frame-aligned dynamics dataset
Compact tokenizer-code dataset for training action-conditioned SocialJax Harvest dynamics models.
Repository: ParoleLM/socialjax-harvest-frame-aligned-512
Format: frame_aligned_socialjax_dynamics_v3
Tokenizer codes per frame: 88
Codebook size: 512
Maximum agents: 7
Total size: 4.385 GB
Train: 9,720 rollouts, 19,916,280 frames
Validation: 540 rollouts, 1,106,460 frames
Test: 540 rollouts, 1,106,460 frames
Files… See the full description on the dataset page: https://huggingface.co/datasets/ParoleLM/socialjax-harvest-frame-aligned-512.Bilingual-SFT-2.0-Pashto-English-Aligned
Bilingual SFT 2.0 — Pashto English Aligned 🇦🇫🇬🇧
Bilingual-SFT-2.0-Pashto-English-Aligned is a bilingual supervised fine-tuning dataset designed to improve Large Language Models (LLMs) in Pashto ↔ English understanding, instruction following, conversation, and bilingual generation.
The dataset uses a conversational messages format and is intended for modern instruction-tuning pipelines, including Hugging Face Transformers, TRL, Unsloth, Axolotl, and other SFT frameworks.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-2.0-Pashto-English-Aligned.new-twi-tts-aligned-ref-prepped
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
New Twi Tts Aligned Ref Prepped
tulu_delta-learning_Qwen2.5-3B-1.5B_reward-alignedtulu_delta-learning_3B-1.5B_reward-alignedculturally_aligned_arabic_stories_subset_a
📚 Culturally Aligned Arabic Stories Dataset (Subset A)
A curated 110-example subset of the Crafting Culturally Aligned Narratives dataset, designed for the development and evaluation of Arabic children’s story generation models aligned with Islamic and cultural values.
✨ Overview
Language: Modern Standard Arabic (MSA)
Samples: 110 prompt–response pairs
Format: JSONL (id, language, prompt, response, source, license)
Moral domains: honesty, courage, generosity… See the full description on the dataset page: https://huggingface.co/datasets/houssamboukhalfa/culturally_aligned_arabic_stories_subset_a.pusht_norm4_stopreq_plain_aligned100k
PushT Plain Stopreq Aligned To CoT 100k
Plain records are selected from /data/home/jiaxin/unified_world_model/data/pusht_96_norm4_visual_nomarker_data/data by the CoT source_record manifest.
Images and actions are unchanged; only the full prompt receives the stop-required line.
verl_humanual_book_aligned_samplesdelta-Qwen2.5-3B-vs-1.5B-reward_alignedpusht_norm4_stopreq_cot_aligned100k
PushT CoT Stopreq Candidate Shuffle Aligned 100k
Built from /data/home/raychai/hf_datasets/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot_stopreq_candidate_shuffle_20260604_004542 using manifest pusht_stopreq_plain_cot_aligned100k_v1.
Training rows are the first 100000 rows of the rewritten CoT source, resharded into 8 files for BAGEL_JSONL_STREAMING parity with the plain run.
legacy_pusht_norm4_allstep_cot_stopreq_aligned100k_ordered
PushT CoT Stopreq Candidate Shuffle Aligned 100k
Built from /data/home/raychai/hf_datasets/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot_stopreq_candidate_shuffle_20260604_004542 using manifest pusht_stopreq_plain_cot_aligned100k_v1.
Training rows are the first 100000 rows of the rewritten CoT source, resharded into 8 files for BAGEL_JSONL_STREAMING parity with the plain run.
self-aligned-instruction-datasetsalamandra40b-aligned_Results_ca_prompt1_testsalamandra40b-aligned_Results_es_prompt1_testdextr-aligned-bboxesself-aligned-instruction-dataset-assignmenttransformers-en-ko-aligned-docs
Transformers EN-KO Aligned Docs
This dataset package contains English-Korean aligned text pairs derived from the docs/source/en and docs/source/ko trees in huggingface/transformers.
Repository layout
data/: published dataset splits only
metadata/: filtering, blacklist, and build-status artifacts
docs/: agent harness and dataset construction notes
AGENTS.md: short Codex entry point for this dataset repo
Contents
data/train.jsonl: final training split with… See the full description on the dataset page: https://huggingface.co/datasets/jmj-minju/transformers-en-ko-aligned-docs.
