datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
govza-sa-cabinet-statements-sentence-aligned
Gov-ZA Multilingual Cabinet Statements (Sentence-Aligned)
Dataset Description
This dataset contains sentence-aligned parallel text from South African government cabinet statements in 11 official languages. The data is sourced from the Government Communication and Information System (GCIS) and scraped from www.gov.za/cabinet-statements.
Key Features:
📊 55 language pair combinations covering 11 South African languages
🔗 Sentence-level alignment using LASER embeddings
📈… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/govza-sa-cabinet-statements-sentence-aligned.new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.sentence-undl_zh2en_alignedraw:bot-yaya/undl_zh2en_aligned
work:split
aligned-mwe
aligned_mwe — multi-word target expressions per lexeme
Where lexeme-alignments is one row per surface
token, this is one row per lexeme rendered by a contiguous multi-word phrase (חֶסֶד → "kasih setia",
בֵּית → "tempat pengirikan"). Mined from the aligner's per-verse t_idx positions: only spans whose
target token positions are contiguous (max−min+1 == len) qualify — scattered tokens that merely all
linked to a lexeme are dropped (and counted in the manifest as scattered_dropped… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/aligned-mwe.new-twi-tts-aligned
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi TTS Dataset
A speech dataset of Twi (Akan) extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Text-To-Speech (TTS) models.
📂 Dataset Structure
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned.jetson1-060926-subtask-place_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060926-subtask-place_aligned.jetson1-060826-subtask-grab2_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060826-subtask-grab2_aligned.korea_speech_mfa_aligned_validationlibrispeech_mfcc_alignednew-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/new-twi-tts-aligned-ipa.undl_ru2en_aligned
Dataset Card for "undl_ru2en_aligned"
More Information needed
VietBibleVox-aligned
VietBibleVox Dataset
The VietBibleVox Dataset is based on the data extracted from open.bible specifically for the Vietnamese language. As the original data is provided under the cc-by-sa-4.0 license, this derived dataset is also licensed under cc-by-sa-4.0.
The dataset comprises 29,185 pairs of (verse, audio clip), with each verse from the Bible read in Vietnamese by a male voice.
The verses are the original texts and may not be directly usable for training text-to-speech models.… See the full description on the dataset page: https://huggingface.co/datasets/ntt123/VietBibleVox-aligned.warsh_quran_alignedshona-bible-bdsc-aligned
Shona Bible Speech Alignment Dataset
Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC
source audio made available by Biblica, Inc. through Open.Bible. This release
contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech
segments covering approximately 75.55 hours.
Dataset summary
Language: Shona (sna)
Speaker: narrator 1
Speaker sex: male
Books: 66
Clips: 31,284
Audio: approximately 75.55 hours
Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.portuguese_tedx_alignedspanish_voxpopuli_alignedemilia-yodas-en-aligned
Emilia-YODAS EN Word-Aligned
Word-level forced-alignment timestamps for the English subset of
amphion/Emilia-Dataset
(Emilia-YODAS split), produced with
Qwen/Qwen3-ForcedAligner-0.6B.
No audio is redistributed — this dataset contains only metadata (IDs,
transcripts already present in Emilia-YODAS, and per-word [start, end]
timestamps). To use it, join on id with the original Emilia-YODAS audio.
Stats
Metric
Value
Utterances
4,516,833
Total audio
11,572.7… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-aligned.robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150
ABORTED research artifact — retained for reproducibility, not deleted.
This artifact is retained as historical evidence only. Its cached segment-start main-camera observations are not eligible for dynamic-main-view claims. See ABORTED.yaml for the machine-readable archival record.
Archival registry mapping:
experiment_id: E003
robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150
This dataset was created using LeRobot.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150.OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve
OpenSakura Eve LN Aligned Dataset
OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve is a Japanese-to-Chinese light-novel translation dataset in OpenSakura ALIGNED format.
Stats below are computed from the actual generated parquet files.
Dataset Summary
Metric
Value
Dataset ID
OpenSakura/OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve
Total rows
631,009
Total parquet files
213 (train: 148, validation: 22, test: 21)
Total size
7,531,530,066 bytes (~7.53 GB, ~7.01 GiB)… See the full description on the dataset page: https://huggingface.co/datasets/OpenSakura/OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve.pick-mustard-wholebody-50hz-half-alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Dex3_GR00T_N16",
"total_episodes": 101,
"total_frames": 23497,
"total_tasks": 1,
"total_videos": 101,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyeonseong-kim-k/pick-mustard-wholebody-50hz-half-aligned.undl_ar2en_aligned
Dataset Card for "undl_ar2en_aligned"
More Information needed
correct_transcription_alignedaligned_seqsfree_st_chinese_mandarin_corpus_mfa_alignedswe-zero-aligned-v9MTBench_finance_aligned_pairs_short
MTBench: A Multimodal Time Series Benchmark
MTBench (Huggingface, Github, Arxiv) is a suite of multimodal datasets for evaluating large language models (LLMs) in temporal and cross-modal reasoning tasks across finance and weather domains.
Each benchmark instance aligns high-resolution time series (e.g., stock prices, weather data) with textual context (e.g., news articles, QA prompts), enabling research into temporally grounded and multimodal understanding.
🏦 Stock… See the full description on the dataset page: https://huggingface.co/datasets/GGLabYale/MTBench_finance_aligned_pairs_short.MTBench_finance_aligned_pairs_long
MTBench: A Multimodal Time Series Benchmark
MTBench (Huggingface, Github, Arxiv) is a suite of multimodal datasets for evaluating large language models (LLMs) in temporal and cross-modal reasoning tasks across finance and weather domains.
Each benchmark instance aligns high-resolution time series (e.g., stock prices, weather data) with textual context (e.g., news articles, QA prompts), enabling research into temporally grounded and multimodal understanding.
🏦 Stock… See the full description on the dataset page: https://huggingface.co/datasets/GGLabYale/MTBench_finance_aligned_pairs_long.undl_zh2en_aligned
联合国数字图书馆的段落级中-英对齐平行语料
用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。
bleu score 这里贴一份吧,懒得转格式了,我不太懂看,可能很差(
Language & Paragraph Count & Avg Tokens & bleu1 & bleu2 & bleu3 & bleu4 \\
\midrule
ar & 59754 & 52.71873 & 0.73799 & 0.58027 & 0.48118 & 0.40782 \\
de & 187 & 69.58824 & 0.62058 & 0.38837 & 0.26155 & 0.18271 \\
es & 66537 & 50.70776 & 0.74566 & 0.58545 & 0.48445 & 0.41073 \\
fr & 68765 & 52.13133 & 0.67895 & 0.49830 &… See the full description on the dataset page: https://huggingface.co/datasets/bot-yaya/undl_zh2en_aligned.swe-zero-aligned-v10new-twi-tts-aligned_normalised
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
