datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
full_tmdb_movies_datasetdata_gouv_datasets_catalog-full-documents
🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée
Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data.
Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes :
titre et description,
organisation productrice,
licence,
couverture spatiale et temporelle,
fréquence de mise à jour,
formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.house_kg_full_dataset
house.kg — Kyrgyzstan Real Estate (multimodal)
A complete snapshot of house.kg, the largest real-estate
board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller
identities, agency ratings, reviews — and 227,294 photographs.
Field names are English; values are kept in the original language (Russian/Kyrgyz),
exactly as the site renders them.
💻 Scraper source code on GitHub →
The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.beamit-annotated-full-texts-dataset
Dataset Card for "beamit-annotated-full-texts-dataset"
More Information needed
kernel-vuln-dataset-full
Linux Kernel Vulnerability-Introducing Commits Dataset
Dataset Description
A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed.
Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix.
How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/pebblebed/kernel-vuln-dataset-full.kernel-vuln-dataset-full
Linux Kernel Vulnerability-Introducing Commits Dataset
Dataset Description
A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed.
Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix.
How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/quguanni/kernel-vuln-dataset-full.youtube_processed_full_dataset_finalfull_dataset_grasping-tagged
full_dataset_grasping
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.full_dataset_grasping
full_dataset_grasping
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
Full_Length_Motion_Capture_Dataset
Apple Arts Studios Full-Length Motion Capture Dataset
Dataset Overview
The Apple Arts Studios Full-Length Motion Capture Dataset is a professionally captured, full-body human-motion dataset containing 199 hours and 30 minutes of continuous motion capture data.
Unlike segmented motion datasets, this repository preserves the complete capture sequences without separating individual actions into short clips.
The recordings retain their continuous capture… See the full description on the dataset page: https://huggingface.co/datasets/Appleartsstudios/Full_Length_Motion_Capture_Dataset.Full-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.nanonla-qwen3-8b-L24-data-full
Qwen3-8B NLA — FULL parquets (activation_vector regenerated)
The slim NLA splits with the activation_vector column recomputed (raw layer-24
residual at the final token of detokenized_text_truncated). Three configs:
av_sft / ar_sft (warm-start SFT) and rl (RL + held-out eval). Each has a
different prompt schema, hence separate configs.
prolog-dataset-fullDataset with Prolog code / query pairs and execution results.
full-modality-data
Full Modality Dataset Statistics
Video Statistics
Total Videos: 28,472
Total Duration: 1422.33 hours
Average Duration: 179.84 seconds
Median Duration: 160.08 seconds
Duration Range: 10.04s - 1780.03s
QA Statistics
Total Questions: 1,444,526
Average Questions per Video: 50.7
Questions per Video Range: 14 - 450
Question Type Distribution
OE: 1,444,526 (100.0%)
Question Category Distribution
temporal: 96,873 (6.7%)
causal: 96,873… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/full-modality-data.new_full_dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.cam_zed_left": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/H2Ozone/new_full_data.house_kg_full_dataset_frames
house.kg — Kyrgyzstan Real Estate, over time
Sale and rental listings scraped from house.kg, the largest
real-estate board in Kyrgyzstan, re-measured on a schedule. Field names are
English; values are kept in the original language (Russian), exactly as the site
renders them.
Coverage: 2026-09-08. This is the baseline snapshot; later runs append new partitions.
Subsets
subset
rows
description
listings
25,264
one row per advertisement — current state plus… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset_frames.SepsisPrediction_Dataset_Full
Dataset Card for "SepsisPrediction_Dataset_Full"
More Information needed
lekiwi-dataset-cross-fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 60,
"total_frames": 18000,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/PRFitz/lekiwi-dataset-cross-full.new_helium_adapter_checkpoint_librispeech_full_dataseteval_so101_30_fulldata120This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 1947,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Rorschach4153/eval_so101_30_fulldata120.lekiwi-full-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_follower",
"total_episodes": 150,
"total_frames": 64992,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:150"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CRPlab/lekiwi-full-dataset.full-ev-infrastrucutre
EV Charging Stations in Finland — Merged Dataset
File: ev_stations_merged.csv
Rows: 5,410 (one row per charging location) · Columns: 25
Companion file: merge_review.csv (3,158 rows - the matched pairs only, for manual review)
This dataset combines two independently collected lists of public EV charging locations in Finland into a single table. It is the base table for the DSP project: recommending 50 new fast-charging locations.
1. Source datasets
Google… See the full description on the dataset page: https://huggingface.co/datasets/data-sci-project/full-ev-infrastrucutre.Planning-Data-Math-Full-Soln-thinknew_helium_adapter_checkpoint_librispeech_full_datasetupdated_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v2full_eval_data_imdbopenfacades-dataset-fulldataset_pick_oranges_full_taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 33,
"total_frames": 32767,
"total_tasks": 1,
"total_videos": 66,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:33"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/weblucas/dataset_pick_oranges_full_task.dataset_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212_2_global_step_70
