CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01duke-trust-lab /human-aligned-similarity-benchmark Human Aligned Similarity Benchmark You are welcome to go to alignedmachine.com to contribute. Overview This dataset contains human-aligned similarity judgments for embedding text and multimodal AI model evaluation. The benchmark is designed to assess how well AI models align with human cognitive preferences in similarity perception across text and image modalities. Dataset Structure Concept Files This dataset contains human preference judgments for… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/human-aligned-similarity-benchmark.textn<1K0 likes179 downloads9mo agoHugging Face02kyutai /hifitts2-aligned HiFiTTS-2 word alignments Word-level forced alignments for the HiFiTTS-2 corpus (44 kHz subset, resampled to 24 kHz), as used to train pocket-tts models. Like HiFiTTS-2 itself, this dataset contains no audio — only pointers and annotations. The audio is downloaded from LibriVox and cut locally. Contents train/train_aligned-*.jsonl.gz — the full aligned training manifest eval_aligned.jsonl.gz — a 1000-utterance held-out split scripts/download_audio.py — fetches… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/hifitts2-aligned.tabulartext-to-speech10M<n<100M2 likes154 downloads1mo agoHugging Face03abhik1505040 /doc-aligned-crossSum-subsettext100K<n<1M0 likes146 downloads2y agoHugging Face04realzL /unified-alignedtext1M<n<10M0 likes80 downloads3mo agoHugging Face05chenjigaram /eu-financial-regulation-aligned EU Financial Regulation, Aligned Across 24 Languages Eight EU financial regulations, split to the paragraph, in every official EU language, with the alignment verified rather than assumed. 56,838 rows — 2,442 provisions × up to 24 languages. Why the alignment is exact Most multilingual legal corpora are aligned by matching sentences, which is approximate and fails on exactly the long provisions people care about. This one does not do that. EUR-Lex assigns… See the full description on the dataset page: https://huggingface.co/datasets/chenjigaram/eu-financial-regulation-aligned.texttext-retrieval10K<n<100K0 likes42 downloads1mo agoHugging Face06emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-en PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.texttext-generation100K<n<1M0 likes40 downloads3mo agoHugging Face07PTTREP /asynchow-code-aligned-minutes AsynChow Code-Aligned Minutes This dataset is a unit-normalized variant of the AsynChow data released with fangru-lin/procedure_generalization_llm, pinned to source commit d9bf3485cd41c1050d33471d922c826f474efec1. It contains three aligned representations of each weighted DAG scheduling problem: natural: natural-language steps and precedence constraints; graph: adjacency-list and duration-dictionary representation; python: executable-style Python representation from the… See the full description on the dataset page: https://huggingface.co/datasets/PTTREP/asynchow-code-aligned-minutes.tabularquestion-answering1K<n<10K0 likes26 downloads2d agoHugging Face08emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-fr PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.texttext-generation10K<n<100K0 likes24 downloads3mo agoHugging Face09msr-spare-1 /qwen3-4b-0630-tooluse-eval-aligned-r32-spare-games-envs qwen3-4B-Instruct-0630-tooluse-eval-aligned-r32 — generated environments Environments generated by the SPARE proposer during training run 050mlekj (qwen3-4B-Instruct-0630-tooluse-eval-aligned-r32), recovered from the spare-viz durable cache. The run's scratch directory no longer exists; this dataset is the surviving copy. Games 456 Steps covered 21 (step 0–448) With recovered skill 456 With hint 0 Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-4b-0630-tooluse-eval-aligned-r32-spare-games-envs.textn<1K0 likes23 downloads1mo agoHugging Face10Taxonomy-Aligned-Conversational-Tutor /TACTBench-Samples TACTBench Demonstration Samples This repository contains five full-context demonstration examples from TACTBench. It does not contain the TACT training set or the remaining hidden TACTBench evaluation set. The samples use the same full-history representation as the benchmark evaluation and illustrate direct correction, error explanation, guided revision, clarification checking, affective feedback, and retry elicitation. Data data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.tabulartext-generationn<1K0 likes23 downloads1d agoHugging Face11ParoleLM /socialjax-harvest-frame-aligned-512 SocialJax Harvest frame-aligned dynamics dataset Compact tokenizer-code dataset for training action-conditioned SocialJax Harvest dynamics models. Repository: ParoleLM/socialjax-harvest-frame-aligned-512 Format: frame_aligned_socialjax_dynamics_v3 Tokenizer codes per frame: 88 Codebook size: 512 Maximum agents: 7 Total size: 4.385 GB Train: 9,720 rollouts, 19,916,280 frames Validation: 540 rollouts, 1,106,460 frames Test: 540 rollouts, 1,106,460 frames Files… See the full description on the dataset page: https://huggingface.co/datasets/ParoleLM/socialjax-harvest-frame-aligned-512.tabular10K<n<100K0 likes16 downloads4mo agoHugging Face12nassimjp /Bilingual-SFT-2.0-Pashto-English-Aligned Bilingual SFT 2.0 — Pashto English Aligned 🇦🇫🇬🇧 Bilingual-SFT-2.0-Pashto-English-Aligned is a bilingual supervised fine-tuning dataset designed to improve Large Language Models (LLMs) in Pashto ↔ English understanding, instruction following, conversation, and bilingual generation. The dataset uses a conversational messages format and is intended for modern instruction-tuning pipelines, including Hugging Face Transformers, TRL, Unsloth, Axolotl, and other SFT frameworks.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-2.0-Pashto-English-Aligned.texttext-generation100K<n<1M0 likes16 downloads1mo agoHugging Face13ghananlpcommunity /new-twi-tts-aligned-ref-prepped This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. New Twi Tts Aligned Ref Prepped text10K<n<100K0 likes14 downloads3mo agoHugging Face14cpsu04 /tulu_delta-learning_Qwen2.5-3B-1.5B_reward-alignedtabular100K<n<1M0 likes13 downloads7mo agoHugging Face15cpsu04 /tulu_delta-learning_3B-1.5B_reward-alignedtabular100K<n<1M0 likes11 downloads7mo agoHugging Face16houssamboukhalfa /culturally_aligned_arabic_stories_subset_a 📚 Culturally Aligned Arabic Stories Dataset (Subset A) A curated 110-example subset of the Crafting Culturally Aligned Narratives dataset, designed for the development and evaluation of Arabic children’s story generation models aligned with Islamic and cultural values. ✨ Overview Language: Modern Standard Arabic (MSA) Samples: 110 prompt–response pairs Format: JSONL (id, language, prompt, response, source, license) Moral domains: honesty, courage, generosity… See the full description on the dataset page: https://huggingface.co/datasets/houssamboukhalfa/culturally_aligned_arabic_stories_subset_a.texttext-generationn<1K0 likes8 downloads11mo agoHugging Face17novastar114 /pusht_norm4_stopreq_plain_aligned100k PushT Plain Stopreq Aligned To CoT 100k Plain records are selected from /data/home/jiaxin/unified_world_model/data/pusht_96_norm4_visual_nomarker_data/data by the CoT source_record manifest. Images and actions are unchanged; only the full prompt receives the stop-required line. tabular100K<n<1M0 likes8 downloads4mo agoHugging Face18asingh15 /verl_humanual_book_aligned_samplestextn<1K0 likes7 downloads2mo agoHugging Face19cpsu04 /delta-Qwen2.5-3B-vs-1.5B-reward_alignedtabular100K<n<1M0 likes6 downloads7mo agoHugging Face20novastar114 /pusht_norm4_stopreq_cot_aligned100k PushT CoT Stopreq Candidate Shuffle Aligned 100k Built from /data/home/raychai/hf_datasets/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot_stopreq_candidate_shuffle_20260604_004542 using manifest pusht_stopreq_plain_cot_aligned100k_v1. Training rows are the first 100000 rows of the rewritten CoT source, resharded into 8 files for BAGEL_JSONL_STREAMING parity with the plain run. tabular100K<n<1M0 likes6 downloads4mo agoHugging Face21novastar114 /legacy_pusht_norm4_allstep_cot_stopreq_aligned100k_ordered PushT CoT Stopreq Candidate Shuffle Aligned 100k Built from /data/home/raychai/hf_datasets/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot_stopreq_candidate_shuffle_20260604_004542 using manifest pusht_stopreq_plain_cot_aligned100k_v1. Training rows are the first 100000 rows of the rewritten CoT source, resharded into 8 files for BAGEL_JSONL_STREAMING parity with the plain run. tabular100K<n<1M0 likes5 downloads4mo agoHugging Face22Nikichoksi /self-aligned-instruction-datasettextn<1K0 likes4 downloads1y agoHugging Face23rsepulvedat /salamandra40b-aligned_Results_ca_prompt1_testtextn<1K0 likes4 downloads1y agoHugging Face24rsepulvedat /salamandra40b-aligned_Results_es_prompt1_testtextn<1K0 likes3 downloads1y agoHugging Face25bluecopa /dextr-aligned-bboxestabular1K<n<10K0 likes2 downloads9mo agoHugging Face26Nikichoksi /self-aligned-instruction-dataset-assignmenttextn<1K0 likes1 downloads1y agoHugging Face27jmj-minju /transformers-en-ko-aligned-docsgated Transformers EN-KO Aligned Docs This dataset package contains English-Korean aligned text pairs derived from the docs/source/en and docs/source/ko trees in huggingface/transformers. Repository layout data/: published dataset splits only metadata/: filtering, blacklist, and build-status artifacts docs/: agent harness and dataset construction notes AGENTS.md: short Codex entry point for this dataset repo Contents data/train.jsonl: final training split with… See the full description on the dataset page: https://huggingface.co/datasets/jmj-minju/transformers-en-ko-aligned-docs.tabulartranslation1K<n<10K1 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.