datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xarm_lift_mediumThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 800,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/xarm_lift_medium.xarm_push_mediumThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 800,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/xarm_push_medium.xarm_lift_medium_replayThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 800,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/xarm_lift_medium_replay.datacomp_medium
DataComp Medium Pool
This repository contains metadata files for the medium pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_medium.xarm_push_medium_replayThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 800,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/xarm_push_medium_replay.datacomp-medium-pool-translatedpa-warm-start-sft-medium-5b-mix
geodesic-research/pa-warm-start-sft-medium-5b-mix
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/pa-warm-start-sft-medium-5b-mix", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-medium-5b-mix.brainformer-e-mediumPGLearn-Medium-NewYork2030LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumprompt-swap-medium12-e2-mxfp4-mergedPGLearn-Medium-2869_pegase-nminus1PGLearn-Medium-NewYork2030-nminus1halfcheetah-medium-replay-v0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "open-door",
"total_episodes": 101,
"total_frames": 100899,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/halfcheetah-medium-replay-v0.gsm_infinite_medium_32kmedium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.gsm_infinite_medium_0PGLearn-Medium-1354_pegase-nminus1shadow_fma_medium_1K_240427PGLearn-Medium-1888_rtefactor_medium_32kgsm_infinite_medium_128kmedium-web-pentesting
Medium Web Pentesting Articles
Dataset Description
A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body.
This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.robolab-mgh-mustard-mediumThis dataset was created using LeRobot.
Dataset Description
Simulated manipulation demonstrations generated by our MimicGen reimplementation on the RoboLab
(Isaac Lab) benchmark. One cell of a generator x task x difficulty grid (MimicGen) / generator x source-count grid (PGDG).
What this is
Generator
MimicGen (our reimplementation, not the authors' code)
Task
mustard
Initial-pose randomization
medium — position 50%, yaw ±25°
Episodes
3045… See the full description on the dataset page: https://huggingface.co/datasets/DAVIAN-Robotics/robolab-mgh-mustard-medium.gsm_infinite_medium_8kpickcube_medium_nocamperturb_500This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 501,
"total_frames": 71968,
"total_tasks": 1,
"total_videos": 1002,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:501"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dureduck/pickcube_medium_nocamperturb_500.factor_medium_zerocontextPGLearn-Medium-1354_pegasePGLearn-Medium-2869_pegasemetra-alpha-ma-medium-a05-n3000This dataset was created using LeRobot.
Dataset Description
Successful BananaInBowl demonstrations collected from an RFCL-trained SAC policy in
Isaac Lab (RoboLab), for behaviour-cloning research on strategy diversity.
Lane: metra-alpha — a sweep of the METRA intrinsic-reward scale alpha, with every
other axis held fixed (banana task, 50 demos, sf=0.5, z_dim=3, phi_space=full, z_unit=true).
Each dataset is one (difficulty level, alpha) cell. Difficulty here: medium.
Collection:… See the full description on the dataset page: https://huggingface.co/datasets/DAVIAN-Robotics/metra-alpha-ma-medium-a05-n3000.
