datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fr3_pickplace_extended_new_cmd_SYNC_part_correctedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "fr3",
"total_episodes": 301,
"total_frames": 276017,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:301"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yio-ye2004/fr3_pickplace_extended_new_cmd_SYNC_part_corrected.fr3_pickplace_extended_new_cmd_SYNC_part_corrected_with_prompt_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "fr3",
"total_episodes": 301,
"total_frames": 276017,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:301"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yio-ye2004/fr3_pickplace_extended_new_cmd_SYNC_part_corrected_with_prompt_test.piper_pick_and_place_correctedpreference_data_llama_factory_corrected_format_text_onlyschaeffer_thesis_correctedThe SCHAEFFER dataset (Spectro-morphogical Corpus of Human-annotated Audio with Electroacoustic Features for Experimental Research), is a compilation of 788 raw audio data accompanied by human annotations and morphological acoustic features.
The audio files adhere to the concept of Sound Objects introduced by Pierre Scaheffer, a framework for the analysis and creation of sound that focuses on its typological and morphological characteristics.
Inside the dataset, the annotation are provided in… See the full description on the dataset page: https://huggingface.co/datasets/dbschaeffer/schaeffer_thesis_corrected.ag_news_corrected_labelspreference_data_llama_factory_corrected_formatcorrected_labels_ag_newsur5e_bt_lemon_apple_orange_pink_blue_plate_2cam_150ep_gripper_correctedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5e",
"total_episodes": 150,
"total_frames": 35511,
"total_tasks": 6,
"total_videos": 300,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:150"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gribes02/ur5e_bt_lemon_apple_orange_pink_blue_plate_2cam_150ep_gripper_corrected.Megawika_correctedso101_smolvla_corrected66_dagger15This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.wrist": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/msKim100/so101_smolvla_corrected66_dagger15.MMSoc_Memotion_corrected
Memotion for MTEB
This repository is a byte-preserving cleanup of
mteb/MMSoc_Memotion at revision
f77e225ae55c1987b0b8cbf6badd1c10296f5f34 for
MTEB issue #5158.
The pinned source and the
official Kaggle release
both contain the same truncated train image, image_5119.png (row
4578, SHA-256 63175a3560ace1e74d4e7913206c94b6293473e09539df0e5be791df62e6a2a9). The PNG is missing part of its IDAT
payload. Permissive decoding fabricates 35 black rows, so this dataset excludes
that one… See the full description on the dataset page: https://huggingface.co/datasets/mteb/MMSoc_Memotion_corrected.slimorca-deduped-cleaned-correctedCREDIT: https://huggingface.co/cgato
there was some minor formatting errors in, corrected and pushed to the Open-Orca org*
What is this dataset?
Half of the Slim Orca Deduped dataset, but further cleaned by removing instances of soft prompting.
I removed a ton prompt prefixes which did not add any information or were redundant. Ex. "Question:", "Q:", "Write the Answer:", "Read this:", "Instructions:"
I also removed a ton of prompt suffixes which were simply there to lead the model to answer as… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/slimorca-deduped-cleaned-corrected.AutoMultiTurnByCalm3-22B-Corrected-reformattedso101_smolvla_corrected_66This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.wrist": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/msKim100/so101_smolvla_corrected_66.long_cable_insertion_2_corrected_raw_plus_dagger2xThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "arx",
"total_episodes": 701,
"total_frames": 96716,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 0,
"video_files_size_in_mb": 0,
"fps": 50,
"splits": {
"train": "0:701"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/long_cable_insertion_2_corrected_raw_plus_dagger2x.corrected-mt-bench-ja
Corrected MT-Bench-ja
Inflection AIによるCorrected MT-Benchの日本語訳です。
一部の設問はStability AIによるJapanese MT-Benchを使用しています。
SQAC-CorrectedRaDeR_MATH_allquerytypes_reward1_correctedghanaian-english-words-corrected-transcriptions
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghanaian English Transcript Corrections
A dataset of mistranscribed words and phrases from Ghanaian news media YouTube videos, corrected using Llama 3.1 405B.
Source
Extracted from YouTube transcripts of Ghanaian… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghanaian-english-words-corrected-transcriptions.comics_dataset_512_inv_manga_correctedslimorca-corrected-chatmlaustralian-insurance-pii-dataset-correctedpick-place-three-blocks-full-emg-correctedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 54,
"total_frames": 15183,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:54"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jasontchan/pick-place-three-blocks-full-emg-corrected.final_corrected_hausa_dataset2slimorca-deduped-cleaned-corrected-textcube_in_cup_corrected_for_right_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Peter-zieg-uga/cube_in_cup_corrected_for_right_2.final_corrected_hausa_datasetamharic-asr-llm-correctedrobot-dataset-corrected
corrected-mapping
LeRobot v2.1 format dataset for robot manipulation.
Dataset Structure
Episodes: 1 episodes of robot manipulation
Total Frames: 327 frames
Cameras: 3 camera views per episode
observation.images.base_camera_sensor_image_raw
observation.images.arm1_camera_sensor_image_raw
observation.images.arm2_camera_sensor_image_raw
Robot: Bimanual manipulator with 34 joints
Format: LeRobot v2.1
Usage with LeRobot
from… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/robot-dataset-corrected.
