datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
science-datalake
Science Data Lake
A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline.
Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below.
What's Unique
This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/J0nasW/science-datalake.my-cloud-data-lake
🌊 Zero-Cost Cloud Data Lake on Hugging Face
A complete serverless data pipeline that converts 12GB+ of CSV data to optimized Parquet format and serves it via a FastAPI on Hugging Face Spaces.
🏗️ Architecture Overview
📁 Local CSV Data (12GB+)
↓
🔄 Conversion Script (DuckDB + SNAPPY)
↓
📦 Parquet Files (4.4GB - 67% compression)
↓
☁️ Hugging Face Dataset
↓
🚀 Serverless FastAPI
↓
🌐 Public API Endpoint
📊 Dataset Information
Source: 793… See the full description on the dataset page: https://huggingface.co/datasets/OMCHOKSI108/my-cloud-data-lake.carbon-emission-datalake
Carbon Emission Data Lake
Dataset ini berisi data historis & prediksi emisi karbon serta suhu untuk wilayah
Texas dan Jakarta, dihasilkan secara otomatis oleh pipeline Prefect + dbt.
Kolom: generated_date, generated_at, region, live_temperature,
base_emission_mt, carbon_emission_forecast_mt, prophet_upper, prophet_lower,
is_anomaly (0=Normal, 1=Anomaly, NULL=model belum cukup data), recommended_carbon_cap_mt.
Modul ML: Prophet (forecast 30 hari), IsolationForest (deteksi… See the full description on the dataset page: https://huggingface.co/datasets/sigit48/carbon-emission-datalake.biomni-data-lakeQuant_DataLake
Quant_DataLake 数据湖存储层详解
文档版本: v3.0 (Doc-Code 对齐: Content-Addressable 因子注册表 + 12 因子 + _factor_states 路径修正 + 数据完备性标注)生成日期: 2026-07-24定位: 仅描述数据湖仓物理存储层 (Parquet+JSON 文件布局、Schema 契约、分区规范)。ETL 管道代码设计见 data_pipeline/README.md。语言: 中文 (Technical Chinese)
一、 数据湖层级解剖 (DataLake Anatomy)
1.1 三层架构总览 (Bronze → Silver → Gold)
数据湖采用 Medallion Architecture (奖章架构) 三层递进设计,所有数据以 Hive 风格 year=YYYY/ 分区 存储于 Parquet 文件中。
┌──────────────────────────────────────┐… See the full description on the dataset page: https://huggingface.co/datasets/zzlmgg/Quant_DataLake.science-datalake
Science Data Lake
A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline.
Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below.
What's Unique
This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/GodotCN/science-datalake.auditeur-datalakemodis-lake-powell-toy-dataset
MODIS Water Lake Powell Toy Dataset
Dataset Summary
Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water)
Dataset Structure
Data Fields
water: Label, water or not-water (binary)
sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000)
sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000)
sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000)
sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-toy-dataset.secop-datalake
SECOP II Colombia — Contratos Públicos · VeedurIA Datalake
Dataset de contratación pública colombiana normalizado y anotado con índices de riesgo fiscal por VeedurIA · operado por SWAL.
Fuentes de datos
SECOP II – Sistema Electrónico de Contratación Pública (Colombia)
API: datos.gov.co/resource/jbjy-vk9h
Licencia: Datos Abiertos Colombia · CC-BY 4.0
Schema de cada fila
Campo
Tipo
Descripción
id
string
Referencia única del contrato… See the full description on the dataset page: https://huggingface.co/datasets/iberi22/secop-datalake.data-lakebiomni-datalakexg-freeze-frame-data
xG Freeze-Frame Data — StatsBomb 360
~15.58M freeze-frame rows from 323 StatsBomb 360 matches, capturing player positions at the moment of each shot. Each row represents one visible player in one shot event.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
Quick Start
from datasets import load_dataset
ds = load_dataset("luxury-lakehouse/xg-freeze-frame-data")
df = ds["train"].to_pandas()
# Average number of visible players per shot… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/xg-freeze-frame-data.modis-lake-powell-raster-dataset
MODIS Water Lake Powell Raster Dataset
Dataset Summary
Raster dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water)
Dataset Structure
Data Fields
water: Label, water or not-water (binary)
sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000)
sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000)
sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000)
sur_refl_b04_1:… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-raster-dataset.scoutgpt-training-data
ScoutGPT Training Data — Player Action Sequences
Per-player match-level action sequences unified across StatsBomb Open Data and Wyscout open data. Each row is one player-match's ordered SPADL action sequence with contextual tokens (competition, season, score state, half, opponent), serialized as the token stream that the ScoutGPT transformer consumes during training and at inference.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/scoutgpt-training-data.xg-shot-data
xG Shot Data — StatsBomb + Wyscout
~131K professional soccer shots from StatsBomb Open Data (95K) and Wyscout (43K), with geometric features, categorical context, and goal labels. Partitioned by data_source for selective loading.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
⚠️ Schema change (cut-over 2026-07-22)
This dataset emits both legacy and canonical Kimball key columns side-by-side. The legacy match_id column will be removed on… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/xg-shot-data.xg-shot-data-v3
Pre-Shot xG v3 — Shot Data (Tabular Corpus)
The tabular half of the training corpus for xg_model_v3, the canonical-SPADL-native pre-shot expected goals model from the luxury-lakehouse analytics platform. One row per shot, all providers — the provider is the data_source column, not a separate file format. Sourced from the gold fct_action_values fact.
Each shot is joinable to its freeze-frame player set (dataset xg-shot-freeze-frames) and to the full action-level corpus… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/xg-shot-data-v3.football2vec-training-data
Football2Vec Training Data — SPADL Action Sequences
Tokenized SPADL action sequences for training the Football2Vec v2 transformer encoder. One row per player-match, covering ~87,000 sequences across ~3,000 professional soccer matches from StatsBomb Open Data and Wyscout.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
Quick Start
from datasets import load_dataset
ds = load_dataset("luxury-lakehouse/football2vec-training-data")
df =… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/football2vec-training-data.e-commerce-olist-data-lakepining-for-the-data
Pining for the Data
Soccer tracking data in SkillCorner V3 format (match JSON + tracking JSONL at 10fps),
redistributed as-is from SkillCorner open data under the MIT license.
"It's not pinin', it's passed on! This parrot is no more!"
— Monty Python's Flying Circus, Dead Parrot sketch
What This Is
Tracking data from SkillCorner open data (MIT license). See NOTICE.
Format: SkillCorner V3 (match JSON + tracking JSONL)
Frame rate: 10 fps
Data: Redistributed as-is, no… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/pining-for-the-data.football2vec-360-training-data
Football2Vec 360 Training Data — SPADL Sequences with Freeze-Frame Context
Tokenized SPADL action sequences with StatsBomb 360 freeze-frame context for training the Football2Vec 360-Enriched model. One row per player-match, covering ~2M actions across 323 professional soccer matches from StatsBomb 360 Open Data.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
Quick Start
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/football2vec-360-training-data.rebot-lerobot-dataafrica-lake-chad-basin-fts-appeal-data
Lake Chad Basin FTS Appeal Data | Africa (original)
Size category: n<1K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-lake-chad-basin-fts-appeal-data.attica-traffic-datalake-oldattica-traffic-datalakeDSAIT4000-data-lakefinancial-datalake-buckupautotrain-data-cancer-lakera
AutoTrain Dataset for project: cancer-lakera
Dataset Description
This dataset has been automatically processed by AutoTrain for project cancer-lakera.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<600x450 RGB PIL image>",
"feat_image_id": "ISIC_0024329",
"feat_lesion_id": "HAM_0002954",
"target": 0… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/autotrain-data-cancer-lakera.data-lake-processeddatalake-part2
datalake-part2
TilQazyna деректер архивінің жалғасы · Продолжение архива данных TilQazyna · Companion TilQazyna data archive
Қазақша · Русский · English
Қазақша
datalake-part2 — TilQazyna деректер архивінің көлемі 2.04 ГБ бөлігі. Репозиторий негізгі datalake-пен байланысты дайындық сатысындағы материалдарды бөлек сақтайды.
Құрамы
Файлдар staged бумасына орналасқан. Материалдардың схемасы, пішімі және өңдеу қадамдары жарияланбаған, сондықтан олар… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/datalake-part2.kinyalm-data-lake
KinyaLM Data Lake
This public-gated dataset repository stores KinyaLM data artifacts for team
review. Visitors can see the dataset page, but file access requires a
logged-in Hugging Face account and gated access acceptance.
Current draft batches (1,000 rows total):
sft-drafts-2026-07-13-batch-001: 286 draft rows
sft-drafts-2026-07-13-batch-002: 714 draft rows
Each batch includes:
draft SFT JSONL,
a master review TSV and any review shards,
a review package zip,
manifests with… See the full description on the dataset page: https://huggingface.co/datasets/kinyalm/kinyalm-data-lake.
