datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rampnet-datasetRampNet is a two-stage pipeline that addresses the scarcity of curb ramp detection datasets by using government location data to automatically generate over 210,000 annotated Google Street View panoramas. This new dataset is then used to train a state-of-the-art curb ramp detection model that significantly outperforms previous efforts. In this repo, we provide our generated curb ramp dataset that we use to train the model.
Each parquet row contains a panoramic image and… See the full description on the dataset page: https://huggingface.co/datasets/projectsidewalk/rampnet-dataset.SynthCheX-75K-v2
SynthCheX-75K
SynthCheX-75K is released as a part of the CheXGenBench paper. It is a synthetic dataset generated using Sana (0.6B) [1] fine-tuned on chest radiographs. Sana (0.6B) establishes the SoTA performance on the CheXGenBench benchmark.
The dataset contains 75,649 high-quality image-text samples along with the pathological annotations.
Filtration Process for SynthCheX-75K
Generative models can lead to both high and low-fidelity generations on different subsets… See the full description on the dataset page: https://huggingface.co/datasets/raman07/SynthCheX-75K-v2.ramen_benchmark_jp_beirThis is a copy of https://huggingface.co/datasets/jinaai/ramen_benchmark_jp reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/ramen_benchmark_jp_beir.sim_ram_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 150,
"total_frames": 60000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:150"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hongdaaaaaaaa/sim_ram_scripted.sim_ram_scripted_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 100,
"total_frames": 40000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hongdaaaaaaaa/sim_ram_scripted_2.plant-diseases-100kDiabetic_Retinopathy_Preprocessed_Dataset_256x256This is dataset comes from this Kaggle Dataset
from the user Sachin Kumar.
The goal of the dataset is for the Varun AIM Projects to easily start running and download the dataset on their local computer in the HF libraries as the directory I strongly recommedn to use.
indic_wikisourceDataset contains page image urls and the corresponding annotations from wikisource.
Also has information whether the page has been validated/proofread.
I expect this dataset to be useful for creating OCR models for printed text in Indic languages.
Languages Covered:
as - Assamese
bn - Bengali
gu - Gujarati
hi - Hindi
kn - Kannada
ml - Malayalam
mr - Marathi
or - Odiya
pa - Punjabi
sa - Sanskrit
ta - Tamil
te - Telugu
sim_ram_scripted_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 100,
"total_frames": 40000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hongdaaaaaaaa/sim_ram_scripted_3.rampnet-benchmark
RampNet Benchmark Imagery
⚠️ This benchmark is not part of the RampNet paper
It did not exist when RampNet was published. The paper's tag,
v1.0-iccv2025 (August 2025),
contains no benchmark/ directory at all — its evaluation was a 1,000-panorama manually
labeled gold set (manual_labels/, imagery in
rampnet-dataset), drawn from
the same three training cities.
These 9 city splits were built eleven months later, between 2026-07-22 and 2026-08-01, as
post-publication… See the full description on the dataset page: https://huggingface.co/datasets/projectsidewalk/rampnet-benchmark.rampnet-crop-model-dataset-round1
RampNet Crop-Model Dataset — Round 1 (Project Sidewalk crops)
The training data behind round 1 of the RampNet Stage 1 crop model — 27,704 crops,
13.37 GB — from RampNet: A Two-Stage Pipeline for Bootstrapping Curb Ramp Detection in
Streetscape Images from Open Government Metadata (O'Meara et al., ICCV'25 CV4A11y workshop,
arXiv:2508.09415).
The crop model is what turns a government curb ramp GPS coordinate into a pixel keypoint on a
panorama; every label in
rampnet-dataset was… See the full description on the dataset page: https://huggingface.co/datasets/projectsidewalk/rampnet-crop-model-dataset-round1.iphone_stairs_ramps
iphone_stairs_ramps
Description
Processed entire iphone_stairs_ramps with filter_every_nth=1, 100% of data, and num_subsampled_points=5
Processing Parameters
mateoguaman/iphone_chin:
exclude_outliers_pct: 0
filter_by_curvature: false
filter_every_nth: 1
horizon:
1000: 1.0
num_subsampled_points: 5
mateoguaman/iphone_hip:
exclude_outliers_pct: 0
filter_by_curvature: false
filter_every_nth: 1
horizon:1000: 1.0
num_subsampled_points: 5… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/iphone_stairs_ramps.smallnorb
Dataset Card for "smallnorb"
Dataset Description
NOTE: This dataset is an unofficial port of small NORB based on a repo from Andrea Palazzi using this script. For complete and accurate information, we highly recommend visiting the dataset's original homepage.
Homepage: https://cs.nyu.edu/~ylclab/data/norb-v1.0-small/
Paper: https://ieeexplore.ieee.org/document/1315150
Dataset Summary
From the dataset's homepage:
This database is intended for experiments in… See the full description on the dataset page: https://huggingface.co/datasets/Ramos-Ramos/smallnorb.audiofeaturesalbumcovers
Dataset Card for "audiofeaturesalbumcovers"
More Information needed
JovianVortexHuntermusicdiffuser
Dataset Card for "musicdiffuser"
More Information needed
docvqa-tables-listsFiltered the table/list question type from the HuggingFaceM4/DocumentVQA dataset.
Original Dataset is not mine and licencing driven by licencing of original dataset. Posted this as it may be of use to others.
libero-pick_up_the_black_bowl_next_to_the_ramekin_and_place_it_on_the_plateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 45,
"total_frames": 5940,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:45"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/libero-pick_up_the_black_bowl_next_to_the_ramekin_and_place_it_on_the_plate.fullalbumcovers
Dataset Card for "fullalbumcovers"
More Information needed
libero-pick_up_the_black_bowl_on_the_ramekin_and_place_it_on_the_plateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 39,
"total_frames": 4472,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:39"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/libero-pick_up_the_black_bowl_on_the_ramekin_and_place_it_on_the_plate.audiofeatures2albumcovers
Dataset Card for "audiofeatures2albumcovers"
More Information needed
albumcoversongtitle
Dataset Card for "albumcoversongtitle"
More Information needed
ramen_benchmark_jp
Japanese Ramen Retrieval
Marketing document from Ramen restaurants.
Questions: 29
Language: Japanese
Example:
{
'query': '新潟市がラーメンを愛する街であることを示す具体的な数字として、市民は月にどのくらいの頻度でラーメンを食べていますか?', # As a specific figure showing that Niigata City is a place that loves ramen, how often do citizens eat ramen in a month?
'image_filename': 'page_8.jpg',
'image': <PIL.PngImagePlugin.PngImageFile image mode=RGB size=2867x2024 at 0x7D73469EC820>}
}
Disclaimer
This dataset may… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/ramen_benchmark_jp.sketch_to_hedsketch_to_next_sketchfurnishka_training_dataramen_benchmark_jp_deprecated
Japanese Ramen Retrieval
Marketing document from Ramen restaurants.
Questions: 29
Language: Japanese
Example:
{
'query': '新潟市がラーメンを愛する街であることを示す具体的な数字として、市民は月にどのくらいの頻度でラーメンを食べていますか?', # As a specific figure showing that Niigata City is a place that loves ramen, how often do citizens eat ramen in a month?
'image_filename': 'page_8.jpg',
'image': <PIL.PngImagePlugin.PngImageFile image mode=RGB size=2867x2024 at 0x7D73469EC820>}
}
Disclaimer
This dataset may… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/ramen_benchmark_jp_deprecated.dataset_chromosomejiro-style-ramenMore Information needed
ramzes_pink_cat
