datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data_gouv_datasets_catalog-full-documents
🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée
Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data.
Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes :
titre et description,
organisation productrice,
licence,
couverture spatiale et temporelle,
fréquence de mise à jour,
formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.house_kg_full_dataset
house.kg — Kyrgyzstan Real Estate (multimodal)
A complete snapshot of house.kg, the largest real-estate
board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller
identities, agency ratings, reviews — and 227,294 photographs.
Field names are English; values are kept in the original language (Russian/Kyrgyz),
exactly as the site renders them.
💻 Scraper source code on GitHub →
The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.beamit-annotated-full-texts-dataset
Dataset Card for "beamit-annotated-full-texts-dataset"
More Information needed
kernel-vuln-dataset-full
Linux Kernel Vulnerability-Introducing Commits Dataset
Dataset Description
A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed.
Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix.
How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/pebblebed/kernel-vuln-dataset-full.kernel-vuln-dataset-full
Linux Kernel Vulnerability-Introducing Commits Dataset
Dataset Description
A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed.
Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix.
How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/quguanni/kernel-vuln-dataset-full.youtube_processed_full_dataset_finalfull_dataset_grasping-tagged
full_dataset_grasping
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
full_dataset_grasping
full_dataset_grasping
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
Full-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.nanonla-qwen3-8b-L24-data-full
Qwen3-8B NLA — FULL parquets (activation_vector regenerated)
The slim NLA splits with the activation_vector column recomputed (raw layer-24
residual at the final token of detokenized_text_truncated). Three configs:
av_sft / ar_sft (warm-start SFT) and rl (RL + held-out eval). Each has a
different prompt schema, hence separate configs.
prolog-dataset-fullDataset with Prolog code / query pairs and execution results.
full-modality-data
Full Modality Dataset Statistics
Video Statistics
Total Videos: 28,472
Total Duration: 1422.33 hours
Average Duration: 179.84 seconds
Median Duration: 160.08 seconds
Duration Range: 10.04s - 1780.03s
QA Statistics
Total Questions: 1,444,526
Average Questions per Video: 50.7
Questions per Video Range: 14 - 450
Question Type Distribution
OE: 1,444,526 (100.0%)
Question Category Distribution
temporal: 96,873 (6.7%)
causal: 96,873… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/full-modality-data.new_full_dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.cam_zed_left": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/H2Ozone/new_full_data.house_kg_full_dataset_frames
house.kg — Kyrgyzstan Real Estate, over time
Sale and rental listings scraped from house.kg, the largest
real-estate board in Kyrgyzstan, re-measured on a schedule. Field names are
English; values are kept in the original language (Russian), exactly as the site
renders them.
Coverage: 2026-09-08. This is the baseline snapshot; later runs append new partitions.
Subsets
subset
rows
description
listings
25,264
one row per advertisement — current state plus… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset_frames.SepsisPrediction_Dataset_Full
Dataset Card for "SepsisPrediction_Dataset_Full"
More Information needed
lekiwi-dataset-cross-fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 60,
"total_frames": 18000,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/PRFitz/lekiwi-dataset-cross-full.eval_so101_30_fulldata120This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 1947,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Rorschach4153/eval_so101_30_fulldata120.lekiwi-full-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_follower",
"total_episodes": 150,
"total_frames": 64992,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:150"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CRPlab/lekiwi-full-dataset.Planning-Data-Math-Full-Soln-thinkupdated_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v2full_eval_data_imdbdataset_pick_oranges_full_taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 33,
"total_frames": 32767,
"total_tasks": 1,
"total_videos": 66,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:33"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/weblucas/dataset_pick_oranges_full_task.dataset_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212_2_global_step_70full_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 4,
"total_frames": 1356,
"total_tasks": 1,
"total_videos": 8,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:4"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/danielsanjosepro/full_dataset.updated_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v1lekiwi-dataset-circle-fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 70,
"total_frames": 83980,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:70"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/PRFitz/lekiwi-dataset-circle-full.swe_bench_dep_data_fullmy_dataset_fullfull_tdc_datasetsdataset_v5_charuko_right_full
dataset_v5_charuko_right_full
Synthetic LeRobot dataset built by concatenating each dataset_v5_charuko_approach episode with each dataset_v5_charuko_right_1 episode (all-pairs cross product), with a short linear-interpolated state+action bridge between them.
Episodes: 10 (5 approaches × 2 placements)
Total frames: 4723
FPS: 30
Task string: Place the coke on <placeholder> — sentinel meant to be substituted at training time.
Generated by tools/eval3_v5_concat_pairs.py in the… See the full description on the dataset page: https://huggingface.co/datasets/RobotLearningVLA/dataset_v5_charuko_right_full.
