datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChemistryQAChemistryQA is a complex QA task which cannot be solved by end-to-end neural networks. To answer chemical questions, machines need to understand questions, apply chemistry and math knowledge, and do calculation and reasoning. ChemistryQA contains about 4500 questions covering around 200 chemistry topics, which are collected from https://socratic.org/chemistry.
All credits go to chemistry-qa project by Microsoft (https://github.com/microsoft/chemistry-qa)
Trademarks
This project may contain… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/ChemistryQA.GreenHyperSpectra
🌱 GreenHyperSpectra: A multi-source hyperspectral dataset for global vegetation trait prediction 🌱
GreenHySpectra is a collection of hyperspectral reflectance data of vegetation from different sources. It is intended for Regression machine learning task for plant trait prediction with self and semi-supervised learning.
Spatial coverage
📁 Configurations
1. GreenHyperSpectra: Unlabeled set
Files: all CSVs under unlabeled/… See the full description on the dataset page: https://huggingface.co/datasets/Avatarr05/GreenHyperSpectra.Avazu_x1
Avazu_x1
Dataset description:
This dataset contains about 10 days of labeled click-through data on mobile advertisements. It has 22 feature fields including user features and advertisement attributes. The preprocessed data are randomly split into 7:1:2* as the training set, validation set, and test set, respectively.
The dataset statistics are summarized as follows:
Dataset
Total
#Train
#Validation
#Test
Avazu_x1
40,428,967
28,300,276
4,042,897
8,085,794
Source:… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Avazu_x1.Avazu_x4
Avazu_x4
Dataset description:
This dataset contains about 10 days of labeled click-through data on mobile advertisements. It has 22 feature fields including user features and advertisement attributes. Following the same setting with the AutoInt work, we split the data randomly into 8:1:1 as the training set, validation set, and test set, respectively.
The dataset statistics are summarized as follows:
Dataset
Total
#Train
#Validation
#Test
Avazu_x4
40,428,967
32,343,172… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Avazu_x4.Avazu_x2
Avazu_x2
Dataset description:
This dataset contains about 10 days of labeled click-through data on mobile advertisements. It has 22 feature fields including user features and advertisement attributes. Following the same setting in the AutoGroup work, we randomly split 80% of the data for training and validation, and the remaining 20% for testing, respectively. For all categorical fields, we filter infrequent features by setting the threshold min_category_count=20 and replace them… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Avazu_x2.AVATARavatar-functionsThere is no difference between 'train' and 'test', these are just used thus the csv file can be detected by huggingface.
max_java_exp_len=1784
max_python_exp_len=1469
avaDist_Uncertainty
Dist_Uncertainty – Hyperspectral Case Studies for Plant Trait Uncertainty Assessment
Dataset Description
This dataset contains hyperspectral remote sensing imagery and associated land cover labels used as out-of-domain (OOD) test cases for evaluating uncertainty estimation methods in deep learning-based plant trait retrievals. It accompanies the paper by Cherif et al. (2025, Biogeosciences) and supports the evaluation of a distance-based uncertainty method (Dis_UN)… See the full description on the dataset page: https://huggingface.co/datasets/Avatarr05/Dist_Uncertainty.clinical-therapeutic-protocol-avatar-state-space-construction-v0.1What this dataset tests
Whether a model can construct physiologically plausible patient avatarsas state-spaces suitable for long-horizon protocol stress-testing.
Required outputs
avatar_profile
state_vector_schema
baseline_coherence_score_0_100
fragility_points
Avatar must include
comorbidity stack
baseline state vector
adherence profile
fragility points
Typical failures
avatars that ignore physiology links
schemas missing key streams
fragility points not tied to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-therapeutic-protocol-avatar-state-space-construction-v0.1.radiotherapy-availability-who-africa
Radiotherapy Availability - WHO African Region | Africa (World Health Organization)
Size category: 10K<n<100K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/radiotherapy-availability-who-africa.clinical-bed-available-transfer-execution-coherence-risk-v0.1What this repo is for
Detect when
a bed is confirmed
but the patient does not move
Common breaks
portering delay
cleaning delay
bed reallocated to higher priority
escort delay for ICU
Examples you can use
ICU stepdown bed confirmed at 14:00 but transfer at 20:00
ED bed assigned but cleaning not done
bed confirmed then removed
You use it to flag
boarding risk
capacity loss risk
healthcare-equipment-availability-maintenance-coherence-risk-v0.1What this repo is for
Detect equipment readiness breakdown before delays and safety incidents.
Tracks alignment between equipment need, availability, maintenance status, location visibility, and backup coverage.
Helps hospitals prevent hidden shortages, unsafe workarounds, and procedure delays.
available_world_outlets
Available World Outlets
A curated list of 832 working news/media outlets worldwide, verified
to be online and publishing recent articles (latest article within 90 days as
of the scrape date).
Source
Compiled from a starting list of 1,451 outlets across all countries. Each URL
was fetched and inspected for article-publication metadata (Open Graph / JSON-LD /
<time> tags). Only outlets reachable, returning HTML, and showing a recent
publication date are included.… See the full description on the dataset page: https://huggingface.co/datasets/ahmdjlt/available_world_outlets.clinical-transfer-escalation-bed-availability-coherence-risk-v0.1What this repo is for
Detect when escalation need
and bed or transfer capacity
fall out of alignment
before
ward deterioration
and preventable ICU delays.
streaming-availability
Streaming Availability by Service and Country
1,938 aggregated records from 685,896 streaming availability records across 20+ services in 50+ countries.
Source
DropThe.org — Data platform tracking streaming availability globally.
Links
DropThe.org
Streaming Statistics
Methodology
vulneracode-logscarpark_availabilitymaritime-pilot-tug-availability-queue-coherence-risk-v0.1What this repo is for
Detect when ships can’t move even though berths exist.
You use it to flag:
pilot shortages
tug shortages
weather amplification
anchorage buildup before ports declare congestion
shakespearean-and-modern-english-conversational-dataset
Dataset Card for Shakespearean and Modern English Conversational Dataset
Dataset Summary
This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details.
price_and_availability_dataclinical-result-available-clinician-review-coherence-risk-v0.1What this repo is for
Detect when
results exist in the record
but no clinician reviews them
or review comes late
or action is wrong
Examples you can use
critical potassium not reviewed
critical bleed on CT not reviewed
outpatient critical INR inbox miss
You use it to flag
silent failure risk
engine-predictive-maintenanceReal_estatepoisoned-alpacapoisoned_alpaca_ERROR_MESSAGEWineQTcoinbase-app-reviewsbinance-app-reviews
Dataset Card for Dataset Name
Binance app reviews collected from Google Play Store (28-09-2020 to 27-09-2025)
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/Avaneeshkarthik/binance-app-reviews.avan
