datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beamit-full-texts-dataset
Dataset Card for "beamit-full-texts-dataset"
More Information needed
tadabur-lora-data-fulldata_gouv_datasets_catalog-full-documents
🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée
Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data.
Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes :
titre et description,
organisation productrice,
licence,
couverture spatiale et temporelle,
fréquence de mise à jour,
formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.gpt2small_full_training_datayt_full_image_dataset
Dataset Card for "yt_full_image_dataset"
More Information needed
ONS-Capacity-Factor-Dataset-Full
Dataset Card for ONS-Capacity-Factor-Dataset-Full
Dataset Summary
The ONS-Capacity-Factor-Dataset-Full provides hourly data for wind and solar power plants in Brazil. These values are published by the Operador Nacional do Sistema Elétrico (ONS) — the Brazilian National Electric System Operator — which is responsible for coordinating and controlling electricity generation and transmission in the National Interconnected System (SIN).
The dataset includes data from 2009 to… See the full description on the dataset page: https://huggingface.co/datasets/SamuelM0422/ONS-Capacity-Factor-Dataset-Full.lucidprots_full_data
Dataset Card for "lucidprots_full_data"
More Information needed
BenglaAI_Full_Datasetsoda-vec-data-full_pmc_title_abstract
SODA-VEC Clean Dataset
This is a cleaned and filtered version of the SODA-VEC dataset, containing high-quality biomedical title-abstract pairs from PubMed Central (PMC) articles.
Dataset Overview
Total examples: 26,573,900
Training set: 26,473,900 examples (99.6%)
Validation set: 50,000 examples (0.2%)
Test set: 50,000 examples (0.2%)
Quality Filtering Applied
This dataset has been processed with the following quality filters:
Abstract Length… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract.house_kg_full_dataset
house.kg — Kyrgyzstan Real Estate (multimodal)
A complete snapshot of house.kg, the largest real-estate
board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller
identities, agency ratings, reviews — and 227,294 photographs.
Field names are English; values are kept in the original language (Russian/Kyrgyz),
exactly as the site renders them.
💻 Scraper source code on GitHub →
The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.anv-data-ke-somali-fullanv-data-ke-somali-fullisbndb-full-database
Dataset Card for "isbndb-annas"
More Information needed
beamit-annotated-full-texts-dataset
Dataset Card for "beamit-annotated-full-texts-dataset"
More Information needed
kernel-vuln-dataset-full
Linux Kernel Vulnerability-Introducing Commits Dataset
Dataset Description
A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed.
Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix.
How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/pebblebed/kernel-vuln-dataset-full.fullDatasetWithSplitkernel-vuln-dataset-full
Linux Kernel Vulnerability-Introducing Commits Dataset
Dataset Description
A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed.
Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix.
How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/quguanni/kernel-vuln-dataset-full.youtube_processed_full_dataset_finalfull_dataset_grasping-tagged
full_dataset_grasping
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
R3-full-datasetBurn_To_Win_Full-datasetChinese-Braille-Dataset-Full-Tone
Chinese Braille Sentence Corpus (Full Tone)
📃 [Paper] •
💻 [Code] •
📖 [Passage corpus] •
🎬 [Demo]
The sentence-level half of the Braille–Chinese parallel corpus used in
"Vision-Braille: A Curriculum Learning Toolkit and Braille–Chinese Corpus for Braille
Translation" (EMNLP 2026 Main Conference).
Every Braille sequence here retains all tone markers (retention rate r = 100). This is
the source corpus: the tone-omission variants used for curriculum training are… See the full description on the dataset page: https://huggingface.co/datasets/Violet-yo/Chinese-Braille-Dataset-Full-Tone.soda-vec-data-full_pmc_title_abstract_paired
SODA-VEC Paired Dataset for Negative Sampling
This is a paired version of the SODA-VEC dataset, specifically formatted for negative sampling training with MultipleNegativesRankingLoss.
Dataset Overview
Total examples: 26,573,900
Format: Paired (anchor-positive) for contrastive learning
Source: EMBO/soda-vec-data-full_pmc_title_abstract
Purpose: Training sentence transformers with negative sampling
Data Format
Each example contains:
anchor (string): The title… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract_paired.full_dataset_grasping
full_dataset_grasping
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
Full-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.full_pose_retrain_datasetvqasynth_opencv3d_dataset_3600_wotem_v2_fulldataset-preferences-llm-course-full-dataset
Dataset Card for dataset-preferences-llm-course-full-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davanstrien/dataset-preferences-llm-course-full-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dataset-preferences-llm-course-full-dataset.species-dataset-full-oakR3-full-dataset-no-glue
