datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bridge_orig_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "widowx",
"total_episodes": 53192,
"total_frames": 1893026,
"total_tasks": 19974,
"total_videos": 212768,
"total_chunks": 54,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:53192"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/bridge_orig_lerobot.generic_data_v2mmBERT-pretrain-p2-fineweb2-remaining
mmBERT Pre-training Data P2
Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite.
NOTE: this is only P2 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p2-fineweb2-remaining.X-Atlas-Orion
X-Atlas/Orion
X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human
protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular
identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.ProLong-TextFullreddit_mds_incrementalcc_en_middle_mds_incrementalX-Atlas-Orion
X-Atlas Orion Dataset (SLAF Format)
Attribution
This is a re-release of data originally generated by Xaira Therapeutics.
Original Dataset: Xaira-Therapeutics/X-Atlas-Orion
Original Format: Parquet files
This Release: Same data in SLAF (Sparse Lazy Array Format)
License: CC-BY-NC-SA-4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0)
Original Citation:
@article{huang2025xatlasorion,
title={X-Atlas/Orion: Genome-wide Perturb-seq Datasets via a Scalable… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/X-Atlas-Orion.ProLongTextscenetok_originchunk-extmmBERT-pretrain-p3-others
mmBERT Pre-training Data P3
Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite.
NOTE: this is only P3 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p3-others.bridge_orig_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "widowx",
"total_episodes": 53192,
"total_frames": 1893026,
"total_tasks": 19974,
"total_videos": 212768,
"total_chunks": 54,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:53192"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ryanqian1994/bridge_orig_lerobot.old-data-store-realoxford_spires_datasetWe present the Oxford Spires Dataset, captured in and around well-known landmarks in Oxford using a custom-built multi-sensor perception unit as well as a millimetre-accurate map from a terrestrial LiDAR scanner (TLS). The perception unit includes three global shutter colour cameras, an automotive 3D LiDAR scanner, and an inertial sensor — all precisely calibrated.
Project page
Paper
Arxiv
Video
Code
Sample Usage
Download the Dataset
You can download the dataset from… See the full description on the dataset page: https://huggingface.co/datasets/ori-drs/oxford_spires_dataset.mmBERT-pretrain-p1-fineweb2-langs
mmBERT Pre-training Data P1
Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite.
NOTE: this is only P1 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p1-fineweb2-langs.Robotic_Origami_Challenge
Robotic Origami Challenge: Fold Plane Demonstrations
Real-world LeRobot demonstrations for dexterous paper-airplane folding.
Overview
Robotic Origami Challenge: Fold Plane Demonstrations is a real-world teleoperation dataset for folding a paper airplane with a bimanual dexterous robot system. It is released by Sharpa in a LeRobot-compatible format for the Robotic Origami Challenge community.
Origami is a demanding benchmark for embodied AI:… See the full description on the dataset page: https://huggingface.co/datasets/SharpaIT/Robotic_Origami_Challenge.brand-assetsmmBERT-pretraining-data-chunk1
mmBERT Training Data (Ready-to-Use)
Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training.
This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk1.tulu_flan_mds_incremental-tokensbridge_origAlgoTune
Website |
Paper |
Code
How good are language models at coming up with new algorithms? To try to answer this, we built a benchmark, AlgoTune, comprised of 154 widely used math, physics, and computer science functions. For each function, the goal is to write code that produces the same outputs as the original function, while being faster. In addition to the benchmark, we also provide an agent, AlgoTuner, which allows language models to easily optimize code.… See the full description on the dataset page: https://huggingface.co/datasets/oripress/AlgoTune.thinking_bridge_orig_lerobot_output_qwen3vlD2E-Original
D2E-Original
Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation
This is the dataset for D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI. 273.4 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open-world, sandbox, and more), for training vision-action models and game agents.
What's included:
Video + Audio: H.264 encoded at FHD/QHD 60fps with game audio.
Input events: Keyboard… See the full description on the dataset page: https://huggingface.co/datasets/open-world-agents/D2E-Original.tulu_flan_mds_incrementalrefinedweb_mds_incrementalogb-full-original
OGB full original archives
Public, byte-for-byte mirror of 17 official Open Graph Benchmark (OGB) and OGB Large-Scale Challenge archive downloads used by the Emulated-Inc graph benchmark environment.
Original ZIP archives are stored under archives//. Each dataset directory includes metadata.json with the authoritative source URL, exact byte size, SHA-256 digest, and repository archive path. Archives were transferred directly from the official SNAP/DGL hosts through ephemeral… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/ogb-full-original.cc_news_mds_incremental-tokenscc_en_head_mds_incrementalS4ASen4AgriNet is a Sentinel-2 based time series multi country benchmark dataset, tailored for
agricultural monitoring applications with Machine and Deep Learning. It is annotated from
farmer declarations collected via the Land Parcel Identification System (LPIS) for harmonizing
country wide labels. These declarations have only recently been made available as open data,
allowing for the first time the labelling of satellite imagery from ground truth data.
We proceed to propose and standardise a new crop type taxonomy across Europe that address
Common Agriculture Policy (CAP) needs, based on the Food and Agriculture Organization (FAO)
Indicative Crop Classification scheme. Sen4AgriNet is the only multi-country, multi-year dataset
that includes all spectral information. It is constructed to cover the period 2016-2020 for
Catalonia and France, while it can be extended to include additional countries.
