datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NavOL
🧭 NavOL artifacts
This dataset repository contains the model checkpoints, Dingo robot asset,
processed 50-scene training asset, and benchmark data used by NavOL. Source
code and executable data tools are maintained at
https://github.com/WAboutMe/NavOL.
🧠 Checkpoints
File
Intended use
models/navdp-cross-modal.ckpt
NavDP initialization checkpoint required to start NavOL training
models/checkpoints/navol-mpc-iter1000.pt
Default NavOL inference and… See the full description on the dataset page: https://huggingface.co/datasets/WAboutme/NavOL.kvqa
KVQA: Knowledge-aware Visual Question Answering
KVQA is a visual question-answering dataset created by Sanket Shah, Anand Mishra, Naganand Yadati, and Partha Pratim Talukdar at IIIT Hyderabad / IISc. It is designed for questions about named entities in images that require world knowledge and reasoning over a knowledge graph.
The original release contains 24,602 images, approximately 183,000 manually verified question-answer pairs, more than 18,000 named entities, and a… See the full description on the dataset page: https://huggingface.co/datasets/wabdalm/kvqa.WABAD
WABAD: A World Annotated Bird Acoustic Dataset for Passive Acoustic Monitoring
Original dataset available and full authors list: https://zenodo.org/records/17293588
Dataset corresponding author: Cristian Pérez Granados (cristian.perez@ctfc.cat), Esther Sebastián-González (esther.sebastian@ua.es)
Additional dataset curation: Ben McEwen (b.j.mcewen@tilburguniversity.edu), Corentin Bernard (corentin.bernard@lis-lab.fr)
Dataset Curation
The original dataset (available on… See the full description on the dataset page: https://huggingface.co/datasets/DBD-research-group/WABAD.w_abcyuth_324burgerbotv2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur10e",
"total_episodes": 21,
"total_frames": 12455,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/martinfg-wabo/burgerbotv2.burgerbotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur10e",
"total_episodes": 1,
"total_frames": 650,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/martinfg-wabo/burgerbot.wxf_taskPanda-70Mwikilarge
WikiLarge
HuggingFace implementation of the WikiLarge corpus for sentence simplification gathered by Zhang, Xingxing and Lapata, Mirella.
/!\ I am not one of the creators of the dataset, I just needed a HF version of this dataset and uploaded it. I encourage you to read the paper introducing the dataset: Sentence Simplification with Deep Reinforcement Learning (Zhang & Lapata, EMNLP 2017)
Uses
This dataset can be used to train sentence simplification… See the full description on the dataset page: https://huggingface.co/datasets/waboucay/wikilarge.pusht-20episodesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 20,
"total_frames": 2822,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/martinfg-wabo/pusht-20episodes.turk_corpusCorpus of sentences gathered from Wikipedia and simplifications proposed by Amazon MTurk workers.
Data gathered by Wei Xu, Courtney Napoles, Ellie Pavlick, Quanze Chen and Chris Callison-Burch.imgbedvirtual_teleop_pickplace_30fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "nova",
"total_episodes": 1,
"total_frames": 422,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/martinfg-wabo/virtual_teleop_pickplace_30fps.wa-block-aef-statsred-cube-export-testtrimmning-testtest123Morefund-split-revised1Morefund-splitMorefund-split-revisedMorefund-split-revised2Morefund-split-verifiedMorefund-split-csvMorefund-split-csv1wabnds_wab9gps5mswabgh
