datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.credit-card-clients
Default of Credit Card Clients Dataset
The following was retrieved from UCI machine learning repository.
Dataset Information
This dataset contains information on default payments, demographic factors, credit data, history of payment, and bill statements of credit card clients in Taiwan from April 2005 to September 2005.
Content
There are 25 variables:
ID: ID of each client
LIMIT_BAL: Amount of given credit in NT dollars (includes individual and family/supplementary credit
SEX:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/credit-card-clients.MoleculeNet_ClinTox
MoleculeNet ClinTox
Load and return the ClinTox dataset, part of MoleculeNet [1] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict drug approval viability, by predicting clinical trial toxicity and final FDA approval status. Both tasks are binary.
Characteristic
Description
Tasks
2
Task type
multitask classification
Total samples
1477
Recommended split
scaffold
Recommended metric
AUROC
References
[1]… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ClinTox.in1k_clip_qwen25vl_3b_224res_64tokens_new_ptin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptclinvar_variant_summarysource data from https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz
climate-fever-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever-qrels.Climate-Change-Indicators-For-African-Countries
Climate Change Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Climate-Change-Indicators-For-African-Countries.clinical-trials-xml-2018-2024real_clinical_cases_of_Famous_Old_TCM_Doctors
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors 数据集简介
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors是一个包含了当代著名老中医临床病例的数据集。这些病例数据来源于《当代名老中医典型医案集》(Contemporary Famous Old Chinese Medicine Doctors' Typical Cases Collection)一书。该数据集收录了多位德高望重的老中医大家的真实门诊病历,涵盖了多种常见病和疑难杂症。每个病例都包括病情描述、辨证论治思路、具体治疗方药等宝贵的一手临床资料。这些医案凝聚了老一辈名医的智慧和经验,对于中医的传承发展和临床应用研究,都有重要价值。通过对这些案例的挖掘分析,能够总结老中医诊疗思维、理法方药的特点,为现代中医临床实践提供有益借鉴。
Introduction to TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors… See the full description on the dataset page: https://huggingface.co/datasets/TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors.ClimateX
ClimateX: Expert Confidence in Climate Statements
What do LLMs know about climate? Let's find out!
ClimateX Dataset
We introduce the Expert Confidence in Climate Statements (ClimateX) dataset, a novel, curated, expert-labeled, natural language dataset of 8094 statements extracted or paraphrased from the IPCC Assessment Report 6: Working Group I report, Working Group II report, and Working Group III report, respectively.
Each statement is labeled with the corresponding… See the full description on the dataset page: https://huggingface.co/datasets/rlacombe/ClimateX.clipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.Climate_and_Health
Iganga Mayuge Health and Demographic Surveillance Site (IMHDSS)
This document provides an overview of a verbal autopsy dataset linked with climate data collected at the Iganga Mayuge Health and Demographic Surveillance Site (IMHDSS). The dataset covers the period from January 1, 2007 to December 31, 2022 and includes information on deaths among site residents. Causes of death were classified using the Inter-VA classification algorithm.
Overview
The dataset is a random… See the full description on the dataset page: https://huggingface.co/datasets/APHRC/Climate_and_Health.LDS-retrain-bank-adamw-N16k-bs256-clip1.0detector-clickbait-br-datasets
Detector Clickbait BR - Datasets
Este repositório contém os datasets utilizados para o treinamento do modelo detector-clickbait-br-model, um classificador de textos em português brasileiro capaz de identificar títulos clickbait.
📚 Descrição dos Datasets
1. detector-clickbait-br-raw.csv
Dataset original contendo os dados iniciais sem processamento.
Características:
Dados brutos coletados originalmente
Pode conter duplicatas
Pode conter valores nulos
Formato:… See the full description on the dataset page: https://huggingface.co/datasets/rodrigoaraujorosa/detector-clickbait-br-datasets.International_Classification_Diseases_Clinical_Modification_icd10cm_order_April_2024showdown-clicks
showdown-clicks
General Agents
🤗 Dataset | GitHub
showdown is a suite of offline and online benchmarks for computer-use agents.
showdown-clicks is a collection of 5,679 left clicks of humans performing various tasks in a macOS desktop environment. It is intended to evaluate instruction-following and low-level control capabilities of computer-use agents.
As of March 2025, we are releasing a subset of the full set, showdown-clicks-dev, containing 557 clicks. All examples are… See the full description on the dataset page: https://huggingface.co/datasets/generalagents/showdown-clicks.WIT-es_jina-clip-v2_sampleclinical_evidence
OpenTargets Clinical Evidence Dataset
This OpenTargets clinical evidence dataset represents a comprehensive collection of clinical trial data linking genetic targets to diseases, containing 32 structured columns that capture the complete lifecycle of clinical studies.
The dataset is annotated with clinical trials' predicted stop reason and genetic evidence when available and was referenced in the Why Clinical Trials Stop: The Role of Genetics.
The study analyzed 28,842 stopped… See the full description on the dataset page: https://huggingface.co/datasets/opentargets/clinical_evidence.clinical-trial-outcomes-predictions
Clinical Trial Outcomes Prediction Dataset
A dataset of 1,366 binary forecasting questions about clinical trial outcomes, automatically generated and labeled using Lightning Rod Labs' Future-as-Label methodology.
Dataset Description
This dataset contains questions about pharmaceutical clinical trials from 2023-2024, paired with verified outcomes (success/failure). Each question asks whether a specific trial will meet its endpoints, receive FDA approval, or complete by a… See the full description on the dataset page: https://huggingface.co/datasets/3rdSon/clinical-trial-outcomes-predictions.power-plant-climate-exposure-index
Global Power-Plant Climate Exposure Screening Index (CESI)
Per-plant outdoor-environment severity for the world's power fleet: a 0-100 index, a CES1-CESX screening class and the dominant stressor, derived from each plant's own coordinates (WorldClim normals, Koppen-Geiger class, distance to coast). The full model ships as severity_model.py.
Canonical record: doi.org/10.5281/zenodo.22172589 · Publisher: Inzonex
Load
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Inzinion/power-plant-climate-exposure-index.africa-synth-climate-carbon-emissions-africa-all
Africa Synth Climate Carbon Emissions Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-climate-carbon-emissions-africa-all.china-myeloma-clinical-trials
China Multiple Myeloma Clinical Trials — Open Dataset
Multiple myeloma clinical trials registered in China, curated from official NMPA / CDE filings by the China Myeloma Digital Network (CMDN), an independent non-profit patient advocacy organisation.
This is a mirror. The citable version of record lives at doi.org/10.5281/zenodo.22690814; the source repository is chinamyeloma/china-myeloma-clinical-trials; the documentation is at chinamyeloma.org.
Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/chinamyeloma/china-myeloma-clinical-trials.ClimateEvalambient-o-clip-iqa-patches-imagenet
Ambient Diffusion Omni (Ambient-o): Training Good Models with Bad Data
Dataset Description
Ambient Diffusion Omni (Ambient-o) is a framework for using low-quality, synthetic, and out-of-distribution images to improve the quality of diffusion models. Unlike traditional approaches that rely on highly curated datasets, Ambient-o extracts valuable signal from all available images during training, including data typically discarded as "low-quality."
This dataset card is for… See the full description on the dataset page: https://huggingface.co/datasets/adrianrm/ambient-o-clip-iqa-patches-imagenet.africa-synth-climate-deforestation-land-use-change-all
Africa Synth Climate Deforestation Land Use Change All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-climate-deforestation-land-use-change-all.clinical-quad-unblinding-sae-cluster-media-leak-trial-halt-decision-v0.1Clinical Quad Unblinding SAE Cluster Media Leak Trial Halt Decision v0.1
Each row is a site weekly snapshot.
Core quad
Emergency unblindingSAE clusterMedia leak riskTrial halt decision risk
Target
label_trial_halt_risk_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6
What this repo does
This repository contains a Clarus v0.6 intervention pathway dataset focused on respiratory collapse dynamics.
The dataset evaluates whether a model can determine if a proposed intervention meaningfully stabilizes a deteriorating respiratory system.
The task requires reasoning from:
system state
trajectory toward instability
boundary geometry
recovery geometry
intervention vector
projected trajectory consequence
The model cannot read the answer directly.
It must… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6.ClimateCause
ClimateCause: Complex and Implicit Causal Structures in Climate Reports
Introduction
ClimateCause is a manually expert-annotated dataset of higher-order causal structures from science-for-policy climate reports, including implicit and nested causality. Cause-effect expressions are normalized and disentangled into individual causal relations to facilitate graph construction, with unique annotations for cause-effect correlation, relation type, and spatiotemporal… See the full description on the dataset page: https://huggingface.co/datasets/laallein/ClimateCause.clinical-trial-outcomes-2020plus
Clinical Trial Outcomes (2020+) with Normalized Endpoints
124,790 normalized endpoints across 14,170 clinical studies that started on or after
2020-01-01 and have posted results on ClinicalTrials.gov.
Snapshot: 2026-09-08. Source: ClinicalTrials.gov API v2 (U.S. National Library of Medicine).
Built with ctgov — the same normalizer, released as an
MIT-licensed package with zero dependencies. So this snapshot is not a dead artifact: you can
re-run it against the live registry, or… See the full description on the dataset page: https://huggingface.co/datasets/GooseWithStories/clinical-trial-outcomes-2020plus.
