CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dsfsi /govza-sa-cabinet-statements-sentence-aligned Gov-ZA Multilingual Cabinet Statements (Sentence-Aligned) Dataset Description This dataset contains sentence-aligned parallel text from South African government cabinet statements in 11 official languages. The data is sourced from the Government Communication and Information System (GCIS) and scraped from www.gov.za/cabinet-statements. Key Features: 📊 55 language pair combinations covering 11 South African languages 🔗 Sentence-level alignment using LASER embeddings 📈… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/govza-sa-cabinet-statements-sentence-aligned.texttranslation100K<n<1M1 likes6.3k downloads9mo agoHugging Face02ghanaopenai /new-twi-tts-aligned-ipa new-twi-tts-aligned + IPA phonemes ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme transcription for every clip, produced with ghananlpcommunity/ghana-speech-phoneme-asr. Audio included — this is self-contained, no join with the source dataset needed. Contents split clips hours phoneme units mean units/clip test 16,140 17.24 663,140 41.1 train 145,258 155.21 5,945,389 40.9 Columns column type meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.audioautomatic-speech-recognition100K<n<1M0 likes1.5k downloads2mo agoHugging Face03nwdxlgzs /sentence-undl_zh2en_alignedraw:bot-yaya/undl_zh2en_aligned work:split text10M<n<100M0 likes950 downloads1y agoHugging Face04bcv-commons /aligned-mwe aligned_mwe — multi-word target expressions per lexeme Where lexeme-alignments is one row per surface token, this is one row per lexeme rendered by a contiguous multi-word phrase (חֶסֶד → "kasih setia", בֵּית → "tempat pengirikan"). Mined from the aligner's per-verse t_idx positions: only spans whose target token positions are contiguous (max−min+1 == len) qualify — scattered tokens that merely all linked to a lexeme are dropped (and counted in the manifest as scattered_dropped… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/aligned-mwe.tabulartranslation1M<n<10M0 likes868 downloads11d agoHugging Face05ghanaopenai /new-twi-tts-aligned This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi TTS Dataset A speech dataset of Twi (Akan) extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Text-To-Speech (TTS) models. 📂 Dataset Structure Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned.audio100K<n<1M0 likes703 downloads3mo agoHugging Face06VibeCuisine /jetson1-060926-subtask-place_alignedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060926-subtask-place_aligned.tabularrobotics1K<n<10K0 likes690 downloads3mo agoHugging Face07VibeCuisine /jetson1-060826-subtask-grab2_alignedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060826-subtask-grab2_aligned.tabularrobotics1K<n<10K0 likes668 downloads3mo agoHugging Face08AdoCleanCode /korea_speech_mfa_aligned_validationaudio100K<n<1M0 likes515 downloads8mo agoHugging Face09maliced /librispeech_mfcc_alignedtextautomatic-speech-recognition100K<n<1M0 likes449 downloads1y agoHugging Face10ghananlpcommunity /new-twi-tts-aligned-ipa new-twi-tts-aligned + IPA phonemes ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme transcription for every clip, produced with ghananlpcommunity/ghana-speech-phoneme-asr. Audio included — this is self-contained, no join with the source dataset needed. Contents split clips hours phoneme units mean units/clip test 16,140 17.24 663,140 41.1 train 145,258 155.21 5,945,389 40.9 Columns column type meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/new-twi-tts-aligned-ipa.audioautomatic-speech-recognition100K<n<1M0 likes412 downloads2mo agoHugging Face11bot-yaya /undl_ru2en_aligned Dataset Card for "undl_ru2en_aligned" More Information needed tabulartranslation10M<n<100M0 likes342 downloads1y agoHugging Face12ntt123 /VietBibleVox-aligned VietBibleVox Dataset The VietBibleVox Dataset is based on the data extracted from open.bible specifically for the Vietnamese language. As the original data is provided under the cc-by-sa-4.0 license, this derived dataset is also licensed under cc-by-sa-4.0. The dataset comprises 29,185 pairs of (verse, audio clip), with each verse from the Bible read in Vietnamese by a male voice. The verses are the original texts and may not be directly usable for training text-to-speech models.… See the full description on the dataset page: https://huggingface.co/datasets/ntt123/VietBibleVox-aligned.audiotext-to-speech10K<n<100K2 likes304 downloads1y agoHugging Face13benhadjermed /warsh_quran_alignedaudio10K<n<100K1 likes273 downloads6mo agoHugging Face14manassehzw /shona-bible-bdsc-aligned Shona Bible Speech Alignment Dataset Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC source audio made available by Biblica, Inc. through Open.Bible. This release contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech segments covering approximately 75.55 hours. Dataset summary Language: Shona (sna) Speaker: narrator 1 Speaker sex: male Books: 66 Clips: 31,284 Audio: approximately 75.55 hours Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.audioautomatic-speech-recognition10K<n<100K3 likes267 downloads17d agoHugging Face15AdoCleanCode /portuguese_tedx_alignedaudio10K<n<100K0 likes235 downloads7mo agoHugging Face16AdoCleanCode /spanish_voxpopuli_alignedaudio10K<n<100K0 likes230 downloads8mo agoHugging Face17duplexio /emilia-yodas-en-aligned Emilia-YODAS EN Word-Aligned Word-level forced-alignment timestamps for the English subset of amphion/Emilia-Dataset (Emilia-YODAS split), produced with Qwen/Qwen3-ForcedAligner-0.6B. No audio is redistributed — this dataset contains only metadata (IDs, transcripts already present in Emilia-YODAS, and per-word [start, end] timestamps). To use it, join on id with the original Emilia-YODAS audio. Stats Metric Value Utterances 4,516,833 Total audio 11,572.7… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-aligned.textautomatic-speech-recognition1M<n<10M0 likes210 downloads5mo agoHugging Face18Shiki42 /robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150 ABORTED research artifact — retained for reproducibility, not deleted. This artifact is retained as historical evidence only. Its cached segment-start main-camera observations are not eligible for dynamic-main-view claims. See ABORTED.yaml for the machine-readable archival record. Archival registry mapping: experiment_id: E003 robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150 This dataset was created using LeRobot. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150.tabularrobotics10K<n<100K0 likes209 downloads2mo agoHugging Face19OpenSakura /OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve OpenSakura Eve LN Aligned Dataset OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve is a Japanese-to-Chinese light-novel translation dataset in OpenSakura ALIGNED format. Stats below are computed from the actual generated parquet files. Dataset Summary Metric Value Dataset ID OpenSakura/OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve Total rows 631,009 Total parquet files 213 (train: 148, validation: 22, test: 21) Total size 7,531,530,066 bytes (~7.53 GB, ~7.01 GiB)… See the full description on the dataset page: https://huggingface.co/datasets/OpenSakura/OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve.tabulartranslation100K<n<1M2 likes199 downloads7mo agoHugging Face20hyeonseong-kim-k /pick-mustard-wholebody-50hz-half-alignedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "Unitree_G1_Dex3_GR00T_N16", "total_episodes": 101, "total_frames": 23497, "total_tasks": 1, "total_videos": 101, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:101" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyeonseong-kim-k/pick-mustard-wholebody-50hz-half-aligned.tabularrobotics10K<n<100K0 likes189 downloads9mo agoHugging Face21bot-yaya /undl_ar2en_aligned Dataset Card for "undl_ar2en_aligned" More Information needed tabulartranslation10M<n<100M0 likes177 downloads1y agoHugging Face22AdoCleanCode /correct_transcription_alignedaudio100K<n<1M0 likes154 downloads8mo agoHugging Face23DNA-LLM /aligned_seqstabular1K<n<10K0 likes150 downloads2y agoHugging Face24AdoCleanCode /free_st_chinese_mandarin_corpus_mfa_alignedaudio10K<n<100K0 likes131 downloads8mo agoHugging Face25arpandeepk /swe-zero-aligned-v9text10K<n<100K0 likes125 downloads4mo agoHugging Face26GGLabYale /MTBench_finance_aligned_pairs_short MTBench: A Multimodal Time Series Benchmark MTBench (Huggingface, Github, Arxiv) is a suite of multimodal datasets for evaluating large language models (LLMs) in temporal and cross-modal reasoning tasks across finance and weather domains. Each benchmark instance aligns high-resolution time series (e.g., stock prices, weather data) with textual context (e.g., news articles, QA prompts), enabling research into temporally grounded and multimodal understanding. 🏦 Stock… See the full description on the dataset page: https://huggingface.co/datasets/GGLabYale/MTBench_finance_aligned_pairs_short.textn<1K0 likes124 downloads1y agoHugging Face27GGLabYale /MTBench_finance_aligned_pairs_long MTBench: A Multimodal Time Series Benchmark MTBench (Huggingface, Github, Arxiv) is a suite of multimodal datasets for evaluating large language models (LLMs) in temporal and cross-modal reasoning tasks across finance and weather domains. Each benchmark instance aligns high-resolution time series (e.g., stock prices, weather data) with textual context (e.g., news articles, QA prompts), enabling research into temporally grounded and multimodal understanding. 🏦 Stock… See the full description on the dataset page: https://huggingface.co/datasets/GGLabYale/MTBench_finance_aligned_pairs_long.textn<1K1 likes114 downloads1y agoHugging Face28bot-yaya /undl_zh2en_aligned 联合国数字图书馆的段落级中-英对齐平行语料 用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。 bleu score 这里贴一份吧,懒得转格式了,我不太懂看,可能很差( Language & Paragraph Count & Avg Tokens & bleu1 & bleu2 & bleu3 & bleu4 \\ \midrule ar & 59754 & 52.71873 & 0.73799 & 0.58027 & 0.48118 & 0.40782 \\ de & 187 & 69.58824 & 0.62058 & 0.38837 & 0.26155 & 0.18271 \\ es & 66537 & 50.70776 & 0.74566 & 0.58545 & 0.48445 & 0.41073 \\ fr & 68765 & 52.13133 & 0.67895 & 0.49830 &… See the full description on the dataset page: https://huggingface.co/datasets/bot-yaya/undl_zh2en_aligned.tabulartranslation10M<n<100M2 likes107 downloads1y agoHugging Face29arpandeepk /swe-zero-aligned-v10text1K<n<10K0 likes101 downloads4mo agoHugging Face30ghananlpcommunity /new-twi-tts-aligned_normalised This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. audio100K<n<1M0 likes95 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.