datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.MoleculeNet_ClinTox
MoleculeNet ClinTox
Load and return the ClinTox dataset, part of MoleculeNet [1] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict drug approval viability, by predicting clinical trial toxicity and final FDA approval status. Both tasks are binary.
Characteristic
Description
Tasks
2
Task type
multitask classification
Total samples
1477
Recommended split
scaffold
Recommended metric
AUROC
References
[1]… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ClinTox.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.clinc150MTS_Dialogue-Clinical_Note
MTS Dialogue (Clinical Note Summarisation)
Main Dataset
The MTS-Dialog dataset is a new collection of 1.7k short doctor-patient conversations and corresponding summaries (section headers and contents).
The training set consists of 1,201 pairs of conversations and associated summaries.
The validation set consists of 100 pairs of conversations and their summaries.
The "dialogue" column contain Doctor-Patient conversation. The "section_text" column contains the Clinical Note of the… See the full description on the dataset page: https://huggingface.co/datasets/har1/MTS_Dialogue-Clinical_Note.in1k_clip_qwen25vl_3b_224res_64tokens_new_ptin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptclinvar_variant_summarysource data from https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz
ExpansionRx_OpenADMET_RLM_CLint
ExpansionRx-OpenADMET RLM CLint
RLM CLint (rat liver microsomal intrinsic clearance) dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict the rat liver microsomal intrinsic clearance (RLM CLint) of molecules.
Note that this dataset was not part of the original challenge. It was provided by the organizers afterward as an additional endpoint.
Characteristic
Description
Tasks
1… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_RLM_CLint.climate-fever-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever-qrels.ClimaQA
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025)
Check the paper's webpage and GitHub for more info!
The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.Climate-Change-Indicators-For-African-Countries
Climate Change Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Climate-Change-Indicators-For-African-Countries.clinical-trials-xml-2018-2024real_clinical_cases_of_Famous_Old_TCM_Doctors
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors 数据集简介
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors是一个包含了当代著名老中医临床病例的数据集。这些病例数据来源于《当代名老中医典型医案集》(Contemporary Famous Old Chinese Medicine Doctors' Typical Cases Collection)一书。该数据集收录了多位德高望重的老中医大家的真实门诊病历,涵盖了多种常见病和疑难杂症。每个病例都包括病情描述、辨证论治思路、具体治疗方药等宝贵的一手临床资料。这些医案凝聚了老一辈名医的智慧和经验,对于中医的传承发展和临床应用研究,都有重要价值。通过对这些案例的挖掘分析,能够总结老中医诊疗思维、理法方药的特点,为现代中医临床实践提供有益借鉴。
Introduction to TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors… See the full description on the dataset page: https://huggingface.co/datasets/TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors.ClimateX
ClimateX: Expert Confidence in Climate Statements
What do LLMs know about climate? Let's find out!
ClimateX Dataset
We introduce the Expert Confidence in Climate Statements (ClimateX) dataset, a novel, curated, expert-labeled, natural language dataset of 8094 statements extracted or paraphrased from the IPCC Assessment Report 6: Working Group I report, Working Group II report, and Working Group III report, respectively.
Each statement is labeled with the corresponding… See the full description on the dataset page: https://huggingface.co/datasets/rlacombe/ClimateX.clipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.Climate_and_Health
Iganga Mayuge Health and Demographic Surveillance Site (IMHDSS)
This document provides an overview of a verbal autopsy dataset linked with climate data collected at the Iganga Mayuge Health and Demographic Surveillance Site (IMHDSS). The dataset covers the period from January 1, 2007 to December 31, 2022 and includes information on deaths among site residents. Causes of death were classified using the Inter-VA classification algorithm.
Overview
The dataset is a random… See the full description on the dataset page: https://huggingface.co/datasets/APHRC/Climate_and_Health.climate-news-articles
🌍 Jeu de données d'articles de presse française labellisés comme traitant ou non des sujets liés au climat
🇬🇧 / 🇺🇸 : as this data set is based only on French data, all explanations are written in French in this repository. The goal of the dataset is to train a model to classify titles of French newspapers in two categories : if it's about climate or not.
🗺️ Le contexte
Ce jeu de données de classification de titres d'article de presse française a été réalisé pour… See the full description on the dataset page: https://huggingface.co/datasets/pierre-loic/climate-news-articles.climate-stance-detectiondetector-clickbait-br-datasets
Detector Clickbait BR - Datasets
Este repositório contém os datasets utilizados para o treinamento do modelo detector-clickbait-br-model, um classificador de textos em português brasileiro capaz de identificar títulos clickbait.
📚 Descrição dos Datasets
1. detector-clickbait-br-raw.csv
Dataset original contendo os dados iniciais sem processamento.
Características:
Dados brutos coletados originalmente
Pode conter duplicatas
Pode conter valores nulos
Formato:… See the full description on the dataset page: https://huggingface.co/datasets/rodrigoaraujorosa/detector-clickbait-br-datasets.International_Classification_Diseases_Clinical_Modification_icd10cm_order_April_2024showdown-clicks
showdown-clicks
General Agents
🤗 Dataset | GitHub
showdown is a suite of offline and online benchmarks for computer-use agents.
showdown-clicks is a collection of 5,679 left clicks of humans performing various tasks in a macOS desktop environment. It is intended to evaluate instruction-following and low-level control capabilities of computer-use agents.
As of March 2025, we are releasing a subset of the full set, showdown-clicks-dev, containing 557 clicks. All examples are… See the full description on the dataset page: https://huggingface.co/datasets/generalagents/showdown-clicks.CLIP-ViT-L-14-336-L20-features
OpenAI/CLIP-ViT-L/14@336 Layer 20 features, CLIP+BLIP labels
Feature activation max visualization of the 4096 Features @ L20
CLIP+BLIP labels (may or may not describe what a neuron truly encodes!)
⚠️ May contain sensitive images, albeit abstract. Use responsibly!
Examples:
WIT-es_jina-clip-v2_sampleClinicalTrial-gov_QABiogen_ADME_HLM_CLint
Biogen ADME HLM CLint
Biogen_ADME_HLM_CLint dataset from the Biogen ADME benchmark [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict log10 of human liver microsomal intrinsic clearance (HLM CLint, in mL/min/kg) of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3087
Recommended split
time
Recommended metricMAE
References
[1]
Fang, Cheng, et al.
"Prospective Validation of… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/Biogen_ADME_HLM_CLint.csa-clinical-stage-asset-intelligence-sample
CSA — Clinical-Stage Asset Intelligence · Free Sample
Clinical trials, FDA, and SEC — linked to the drug asset and the listed sponsor, with a
forward catalyst calendar. This is a free 150-row sample of the nearest-term
catalysts; the full snapshot carries 2,221 forward catalysts (955 linked to
124 listed sponsors) and 1,890 resolved assets.
Data, not investment advice. CSA is information, not a recommendation to buy, sell,
or hold any security. Estimated catalyst dates (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/csa-clinical-stage-asset-intelligence-sample.ICD10-Clinical-TerminologyICD10-Clinical-Terminology
pyarrow fast search demonstration for context AI MMoE
clinical_evidence
OpenTargets Clinical Evidence Dataset
This OpenTargets clinical evidence dataset represents a comprehensive collection of clinical trial data linking genetic targets to diseases, containing 32 structured columns that capture the complete lifecycle of clinical studies.
The dataset is annotated with clinical trials' predicted stop reason and genetic evidence when available and was referenced in the Why Clinical Trials Stop: The Role of Genetics.
The study analyzed 28,842 stopped… See the full description on the dataset page: https://huggingface.co/datasets/opentargets/clinical_evidence.clickbait_title_classificationDataset introduced in Stop Clickbait: Detecting and Preventing Clickbaits in Online News Mediaby Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, Niloy Ganguly
Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, and Niloy Ganguly. "Stop Clickbait: Detecting and Preventing Clickbaits in Online News Media”. In Proceedings of the 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), San Fransisco, US, August 2016.
Cite:… See the full description on the dataset page: https://huggingface.co/datasets/marksverdhei/clickbait_title_classification.
