datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aligned-mwe
aligned_mwe — multi-word target expressions per lexeme
Where lexeme-alignments is one row per surface
token, this is one row per lexeme rendered by a contiguous multi-word phrase (חֶסֶד → "kasih setia",
בֵּית → "tempat pengirikan"). Mined from the aligner's per-verse t_idx positions: only spans whose
target token positions are contiguous (max−min+1 == len) qualify — scattered tokens that merely all
linked to a lexeme are dropped (and counted in the manifest as scattered_dropped… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/aligned-mwe.jetson1-060926-subtask-place_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060926-subtask-place_aligned.jetson1-060826-subtask-grab2_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060826-subtask-grab2_aligned.undl_ru2en_aligned
Dataset Card for "undl_ru2en_aligned"
More Information needed
robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150
ABORTED research artifact — retained for reproducibility, not deleted.
This artifact is retained as historical evidence only. Its cached segment-start main-camera observations are not eligible for dynamic-main-view claims. See ABORTED.yaml for the machine-readable archival record.
Archival registry mapping:
experiment_id: E003
robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150
This dataset was created using LeRobot.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_put_obj_cabinet_start-aligned50_finish-aligned50_hybrid_150.OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve
OpenSakura Eve LN Aligned Dataset
OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve is a Japanese-to-Chinese light-novel translation dataset in OpenSakura ALIGNED format.
Stats below are computed from the actual generated parquet files.
Dataset Summary
Metric
Value
Dataset ID
OpenSakura/OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve
Total rows
631,009
Total parquet files
213 (train: 148, validation: 22, test: 21)
Total size
7,531,530,066 bytes (~7.53 GB, ~7.01 GiB)… See the full description on the dataset page: https://huggingface.co/datasets/OpenSakura/OpenSakura-DS-260220-LN-ja-zh-ALIGNED-Eve.pick-mustard-wholebody-50hz-half-alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Dex3_GR00T_N16",
"total_episodes": 101,
"total_frames": 23497,
"total_tasks": 1,
"total_videos": 101,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyeonseong-kim-k/pick-mustard-wholebody-50hz-half-aligned.undl_ar2en_aligned
Dataset Card for "undl_ar2en_aligned"
More Information needed
aligned_seqsundl_zh2en_aligned
联合国数字图书馆的段落级中-英对齐平行语料
用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。
bleu score 这里贴一份吧,懒得转格式了,我不太懂看,可能很差(
Language & Paragraph Count & Avg Tokens & bleu1 & bleu2 & bleu3 & bleu4 \\
\midrule
ar & 59754 & 52.71873 & 0.73799 & 0.58027 & 0.48118 & 0.40782 \\
de & 187 & 69.58824 & 0.62058 & 0.38837 & 0.26155 & 0.18271 \\
es & 66537 & 50.70776 & 0.74566 & 0.58545 & 0.48445 & 0.41073 \\
fr & 68765 & 52.13133 & 0.67895 & 0.49830 &… See the full description on the dataset page: https://huggingface.co/datasets/bot-yaya/undl_zh2en_aligned.genus_aligned_v2umi_cam_aligned_ee_pose_headThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose_head.mobile-stretch3-lift-place-150-hybrid28-alignedSO101-DualArm-teleop_fold_towel_100epi_10fps_action_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 100,
"total_frames": 21649,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/SO101-DualArm-teleop_fold_towel_100epi_10fps_action_aligned.umi_cam_aligned_ee_poseThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose.umi_cam_aligned_ee_pose_curThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose_cur.umi_cam_aligned_ee_pose_head_curThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose_head_cur.umi_cam_aligned_ee_pose_cur_postThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
23
],
"names": [
"left_ee.x",
"left_ee.y",
"left_ee.z",
"left_ee.qx",
"left_ee.qy",
"left_ee.qz"… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/umi_cam_aligned_ee_pose_cur_post.viral_aligned_genus_levelundl_es2en_aligned
Dataset Card for "undl_es2en_aligned"
More Information needed
3_0_hil_data_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 122,
"total_frames": 268044,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:122"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/3_0_hil_data_aligned.glm52-aligned-rubric-traces
GLM-5.2 Aligned Rubric-Writing Traces
23777 teacher traces from GLM-5.2 on the aligned rubric-writing task, collected to
distill / warmstart a smaller rubric-writer. For each (user, book) example the teacher is shown a
persona-conditioned prompt (a user's past book reviews) and asked to (1) predict what that user
would likely write about a new book and (2) produce a <rubric> of numbered criteria for scoring
candidate reviews on coverage of that prediction. The full generation —… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/glm52-aligned-rubric-traces.lane_lift_id_20_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 20,
"total_frames": 592,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YSL2683/lane_lift_id_20_aligned.m9-verifier-38k-aligned
M9 Verifier 38K Aligned
This dataset contains 38,564 prompts with verifier-compatible gold answers for
an M9 RLVR-GRPO experiment in a unified post-training study with Qwen3-1.7B.
It is an independent research artifact, not an official release from the model
or paper authors.
The bank was reconstructed from the frozen
YangyiH/openreasoning_mixed_100k
prompt mixture. Every recovered row was matched to the frozen base row by
domain, source shard, and prompt SHA-256 before verifier… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/m9-verifier-38k-aligned.yam-episode-0109-aligned-cropThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "yam_right_linear_4310",
"total_episodes": 1,
"total_frames": 226,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nuffnuff/yam-episode-0109-aligned-crop.pac-bench-100pct-aligned_rephrasedaprm-snorkelai_agent_finance_reasoning-aligned-v2pac-bench-100pct-alignedbible-ptbr-gun-gub-alignedobjects_collector_lastonly_test5_alignedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Inspire",
"total_episodes": 5,
"total_frames": 2389,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shivubind/objects_collector_lastonly_test5_aligned.
