datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vukuzenzele-sentence-aligned
The Vuk'uzenzele South African Multilingual Corpus
Github: https://github.com/dsfsi/vukuzenzele-nlp/
Zenodo:
Arxiv Preprint:
Give Feedback 📑: DSFSI Resource Feedback Form
About
The dataset was obtained from the South African government magazine Vuk'uzenzele, created by the Government Communication and Information System (GCIS).
The original raw PDFS were obtatined from the Vuk'uzenzele website.
The datasets contain government magazine editions in 11 languages… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/vukuzenzele-sentence-aligned.aligned-mwe
aligned_mwe — multi-word target expressions per lexeme
Where lexeme-alignments is one row per surface
token, this is one row per lexeme rendered by a contiguous multi-word phrase (חֶסֶד → "kasih setia",
בֵּית → "tempat pengirikan"). Mined from the aligner's per-verse t_idx positions: only spans whose
target token positions are contiguous (max−min+1 == len) qualify — scattered tokens that merely all
linked to a lexeme are dropped (and counted in the manifest as scattered_dropped… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/aligned-mwe.jetson1-060926-subtask-place_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060926-subtask-place_aligned.jetson1-060826-subtask-grab2_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060826-subtask-grab2_aligned.undl_ru2en_aligned
Dataset Card for "undl_ru2en_aligned"
More Information needed
robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150
ABORTED research artifact — retained for reproducibility, not deleted.
This artifact is retained as historical evidence only. Its cached segment-start main-camera observations are not eligible for dynamic-main-view claims. See ABORTED.yaml for the machine-readable archival record.
Archival registry mapping:
experiment_id: E003
robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150
This dataset was created using LeRobot.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150.OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve
OpenSakura Eve LN Aligned Dataset
OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve is a Japanese-to-Chinese light-novel translation dataset in OpenSakura ALIGNED format.
Stats below are computed from the actual generated parquet files.
Dataset Summary
Metric
Value
Dataset ID
OpenSakura/OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve
Total rows
631,009
Total parquet files
213 (train: 148, validation: 22, test: 21)
Total size
7,531,530,066 bytes (~7.53 GB, ~7.01 GiB)… See the full description on the dataset page: https://huggingface.co/datasets/OpenSakura/OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve.pick-mustard-wholebody-50hz-half-alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Dex3_GR00T_N16",
"total_episodes": 101,
"total_frames": 23497,
"total_tasks": 1,
"total_videos": 101,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyeonseong-kim-k/pick-mustard-wholebody-50hz-half-aligned.undl_ar2en_aligned
Dataset Card for "undl_ar2en_aligned"
More Information needed
hifitts2-aligned
HiFiTTS-2 word alignments
Word-level forced alignments for the HiFiTTS-2
corpus (44 kHz subset, resampled to 24 kHz), as used to train
pocket-tts models.
Like HiFiTTS-2 itself, this dataset contains no audio — only pointers and
annotations. The audio is downloaded from LibriVox and cut locally.
Contents
train/train_aligned-*.jsonl.gz — the full aligned training manifest
eval_aligned.jsonl.gz — a 1000-utterance held-out split
scripts/download_audio.py — fetches… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/hifitts2-aligned.aligned_seqsundl_zh2en_aligned
联合国数字图书馆的段落级中-英对齐平行语料
用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。
bleu score 这里贴一份吧,懒得转格式了,我不太懂看,可能很差(
Language & Paragraph Count & Avg Tokens & bleu1 & bleu2 & bleu3 & bleu4 \\
\midrule
ar & 59754 & 52.71873 & 0.73799 & 0.58027 & 0.48118 & 0.40782 \\
de & 187 & 69.58824 & 0.62058 & 0.38837 & 0.26155 & 0.18271 \\
es & 66537 & 50.70776 & 0.74566 & 0.58545 & 0.48445 & 0.41073 \\
fr & 68765 & 52.13133 & 0.67895 & 0.49830 &… See the full description on the dataset page: https://huggingface.co/datasets/bot-yaya/undl_zh2en_aligned.genus_aligned_v2setimes-en-tr-aligned-corpus
SETimes EN-TR — Sentence-Aligned, LLM-Cleaned
A cleaned and re-aligned version of the SETimes English-Turkish parallel corpus. The original SETimes data is paragraph-style — each "pair" can contain a headline, a dateline, several body sentences, and a source citation, all glued together on one line. This version splits everything into proper sentence pairs so each row is one English sentence next to its Turkish translation.
144,064 sentence pairs, split into train (142,064)… See the full description on the dataset page: https://huggingface.co/datasets/atahanuz/setimes-en-tr-aligned-corpus.umi_cam_aligned_ee_pose_headThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose_head.mobile-stretch3-lift-place-150-hybrid28-alignedumi_cam_aligned_ee_poseThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose.umi_cam_aligned_ee_pose_curThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose_cur.SO101-DualArm-teleop_fold_towel_100epi_10fps_action_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 100,
"total_frames": 21649,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/SO101-DualArm-teleop_fold_towel_100epi_10fps_action_aligned.umi_cam_aligned_ee_pose_head_curThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose_head_cur.umi_cam_aligned_ee_pose_cur_postThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose_cur_post.viral_aligned_genus_levelglm52-aligned-rubric-traces
GLM-5.2 Aligned Rubric-Writing Traces
23777 teacher traces from GLM-5.2 on the aligned rubric-writing task, collected to
distill / warmstart a smaller rubric-writer. For each (user, book) example the teacher is shown a
persona-conditioned prompt (a user's past book reviews) and asked to (1) predict what that user
would likely write about a new book and (2) produce a <rubric> of numbered criteria for scoring
candidate reviews on coverage of that prediction. The full generation —… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/glm52-aligned-rubric-traces.undl_es2en_aligned
Dataset Card for "undl_es2en_aligned"
More Information needed
3_0_hil_data_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 122,
"total_frames": 268044,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:122"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/3_0_hil_data_aligned.lane_lift_id_20_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 20,
"total_frames": 592,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YSL2683/lane_lift_id_20_aligned.m9-verifier-38k-aligned
M9 Verifier 38K Aligned
This dataset contains 38,564 prompts with verifier-compatible gold answers for
an M9 RLVR-GRPO experiment in a unified post-training study with Qwen3-1.7B.
It is an independent research artifact, not an official release from the model
or paper authors.
The bank was reconstructed from the frozen
YangyiH/openreasoning_mixed_100k
prompt mixture. Every recovered row was matched to the frozen base row by
domain, source shard, and prompt SHA-256 before verifier… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/m9-verifier-38k-aligned.asynchow-code-aligned-minutes
AsynChow Code-Aligned Minutes
This dataset is a unit-normalized variant of the AsynChow data released with
fangru-lin/procedure_generalization_llm,
pinned to source commit d9bf3485cd41c1050d33471d922c826f474efec1.
It contains three aligned representations of each weighted DAG scheduling
problem:
natural: natural-language steps and precedence constraints;
graph: adjacency-list and duration-dictionary representation;
python: executable-style Python representation from the… See the full description on the dataset page: https://huggingface.co/datasets/PTTREP/asynchow-code-aligned-minutes.yam-episode-0109-aligned-cropThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "yam_right_linear_4310",
"total_episodes": 1,
"total_frames": 226,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nuffnuff/yam-episode-0109-aligned-crop.pac-bench-100pct-aligned_rephrased
