datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
egoscaler-v2
EgoScalerV2 Dataset
This dataset accompanies our work on Developing Vision-Language-Action Model from Egocentric Videos. It provides 6DoF object trajectories paired with egocentric visual observations and natural-language action descriptions, formatted in the LeRobot v2.0 schema so it can be consumed directly by LeRobot-compatible pipelines.
🌐 Project page: https://biscue5.github.io/egovla-project-page/
📄 Paper: Developing Vision-Language-Action Model from Egocentric Videos… See the full description on the dataset page: https://huggingface.co/datasets/Biscue5/egoscaler-v2.Motion-o-MCoT-PLM-motion-keyframes
Motion-o-MCoT (PLM + motion keyframes)
Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/.
Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats).
Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.bishkek-transport
Bishkek Public Transport
Open, continuously-growing data on the public-transport network of Bishkek,
the capital of the Kyrgyz Republic. The data originates from the Bishkek
mayoralty's public-transport monitoring system (the same feed behind the city's
official live transit map and its "My City" mobile service). Only publicly
visible transit information is included: stop locations and the live positions of
buses, trolleybuses/electric buses, and marshrutkas (shared minibuses).… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/bishkek-transport.ASVspoof_2019_LAZINC_22PubChemBiST
BiST
💻 Github Repo
English | 简体中文
Introduction
BiST is a large-scale bilingual translation dataset, with "BiST" standing for Bilingual Synthetic Translation dataset. Currently, the dataset contains approximately 60M entries and will continue to expand in the future.
BiST consists of two subsets, namely en-zh and zh-en, where the former represents the source language, collected from public data as real-world content; the latter represents the target language… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/BiST.USPTO_50Kencoded_LA_2021
ASVspoof 2021 LA subset
sorted original dataset from ASVspoof 2021 LA subset
credit: ASVspoof 2021 challenge released under The databases are available under an Open Data Commons Attribution Licence and can be downloaded from the Zenodo repository.
bi-so101-fruits-classificationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so101_follower",
"total_episodes": 2,
"total_frames": 2910,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jy13/bi-so101-fruits-classification.ASVspoof_2021_DFChEMBLbishkek-real-estate
🏠 Bishkek Real Estate Dataset
The largest open dataset of apartment listings from Kyrgyzstan
Dataset Files • Quick Start • Features • Benchmarks • Use Cases
📊 Overview
Comprehensive dataset of 8,821 apartment listings from Bishkek, Kyrgyzstan with:
💰 Prices in USD and per m²
📍 GPS coordinates for 99.6% of listings
🏗️ Building data: year built, house type, floors
🏢 Residential complex info (name, class, developer)
📝 Descriptions in Russian
Why… See the full description on the dataset page: https://huggingface.co/datasets/raimbekovm/bishkek-real-estate.wl-bisindo-10classMOSESiqra_eval_all_refship-detection-sliced-bis
Dataset Card for "ship-detection-sliced-bis"
More Information needed
gsplatIqra_train_processed_whisperBISOARM-100_PICK_AND_PLACE_ESSALUDIqra_train_processed_whisper-phonemebisac_expanded_finalbisonprocessed_Iqra_train_with_syllablesIqra_train_with_arabic_BWsm-cv-ar-entrade-idp-datasetiqraEval_whisper_processed_arabic_BWself_defined_biso101This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "biso101",
"total_episodes": 100,
"total_frames": 66737,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RyanPan27/self_defined_biso101.bi_so_101_clothesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 250,
"total_frames": 104750,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 25,
"splits": {
"train": "0:250"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RebelDot/bi_so_101_clothes.
