datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ts-satfire
Dataset Card for TS-SatFire
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The TS-SatFire dataset is a comprehensive multi-temporal remote sensing dataset designed to cover the entire life cycle of wildfires. It provides a unified framework to support three critical and interconnected wildfire monitoring tasks: active fire detection, daily burned area… See the full description on the dataset page: https://huggingface.co/datasets/SamuelWu318/ts-satfire.sat_multiple_choice_math_may_23This is the set of math SAT questions from the May 2023 SAT, taken from here: https://www.mcelroytutoring.com/lower.php?url=44-official-sat-pdfs-and-82-official-act-pdf-practice-tests-free.
Questions that included images were not included but all other math questions, including those that have tables were included.
e-commerceairline-satisfaction-analysis
Airline Passenger Satisfaction – EDA Report
This project analyzes the Airline Passenger Satisfaction Dataset, containing 103,904 rows and 25 columns describing passenger demographics, flight information, and service ratings.The goal is to understand which factors influence satisfaction, identify important service features,and compare satisfaction between different traveler types and flight classes.
Dataset Overview
The dataset includes:
Passenger demographics (age… See the full description on the dataset page: https://huggingface.co/datasets/drukeroni/airline-satisfaction-analysis.satclip
Dataset Card for S2-100K
The S2-100K dataset is a dataset of 100,000 multi-spectral satellite images sampled from Sentinel-2 via the Microsoft Planetary Computer. Copernicus Sentinel data is captured between Jan 1, 2021 and May 17, 2023. The dataset is sampled approximately uniformly over landmass and only includes images without cloud coverage. The dataset is available for research purposes only. If you use the dataset, please cite our paper. More information on the dataset can… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/satclip.sat-questions-and-answers-for-llm
SAT History Questions and Answers 🏛️ - Text Classification Dataset
This dataset contains a collection of questions and answers for the SAT Subject Test in World History and US History. Each question is accompanied by a corresponding answers and the correct response.
The dataset includes questions from various topics, time periods, and regions on both World History and US History.
💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/sat-questions-and-answers-for-llm.SatBirdxsPlotOpenThis folder includes data files for experiments in the across-datasets setup between SatBird and sPlotOpen.
Laws_and_Constitution_of_Indiasatellite-ground-stations-bgp-dataChameleon-Radiology-Reportsairline_satisfactionsat-listado-69b
SAT Contribuyentes 69-B / 69-B Bis
Listados oficiales del SAT (Servicio de Administración Tributaria) de contribuyentes publicados bajo los Artículos 69-B y 69-B Bis del Código Fiscal de la Federación.
Fuente: http://omawww.sat.gob.mx/cifras_sat/Paginas/DatosAbiertos/contribuyentes_publicados.html
Última actualización
Artículo 69-B: 2026-03-19T00:10:07Z (información al 2026-08-31)
Artículo 69-B Bis: 2026-03-12T18:44:12Z (información al 2026-08-19)… See the full description on the dataset page: https://huggingface.co/datasets/mayrop/sat-listado-69b.trec-cast-2019
TREC Conversational Assistance Track (CAsT)
There are currently few datasets appropriate for training and evaluating models for Conversational Information Seeking (CIS). The main aim of TREC CAsT is to advance research on conversational search systems. The goal of the track is to create a reusable benchmark for open-domain information centric conversational dialogues.
Year 1 (TREC 2019)
Read the TREC 2019 Overview paper.
2019 Data
Topics… See the full description on the dataset page: https://huggingface.co/datasets/satyanshu404/trec-cast-2019.MentalChat16K
🗣️ Synthetic Counseling Conversations Dataset
📝 Description
Synthetic Data 10K
This dataset consists of 9,775 synthetic conversations between a counselor and a client, covering 33 mental health topics such as 💑 Relationships, 😟 Anxiety, 😔 Depression, 🤗 Intimacy, and 👨👩👧👦 Family Conflict. The conversations were generated using the OpenAI GPT-3.5 Turbo model and a customized adaptation of the Airoboros self-generation framework.
The… See the full description on the dataset page: https://huggingface.co/datasets/sathyam123/MentalChat16K.educational-satisfaction-eafc7f
educational-satisfaction-eafc7f
Synthetic products test data: 42 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number… See the full description on the dataset page: https://huggingface.co/datasets/GraniteEvan/educational-satisfaction-eafc7f.Customer_Satisfaction_Surveyafrisenti
Dataset Summary
AfriSenti is the largest sentiment analysis dataset for under-represented African languages, covering 110,000+ annotated tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yoruba).
The datasets are used in the first Afrocentric SemEval shared task, SemEval 2023 Task 12: Sentiment analysis for African languages (AfriSenti-SemEval).… See the full description on the dataset page: https://huggingface.co/datasets/sathvikchandra77/afrisenti.low-television-8444d5
low-television-8444d5
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/satomigoto6/low-television-8444d5.cold-ambition-e6935d
cold-ambition-e6935d
Synthetic weather test data: 54 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/satotakuma/cold-ambition-e6935d.salesdatacontent-eye-2a4aa2
content-eye-2a4aa2
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/satomikako2/content-eye-2a4aa2.satireMS-Marco-Prompt-generationen-sat-latin-splitalpaca_data_cleaned_bhojpuri
Dataset Card for Dataset Name
This repository contains a translated version of the Alpaca-Cleaned dataset, originally provided by Yahma on Hugging Face. The dataset has been translated into Bhojpuri, a language spoken in the northern-eastern part of India and the Terai region of Nepal.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The Alpaca-Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/SatyamDev/alpaca_data_cleaned_bhojpuri.geo-data-satclinical-oxygen-prescription-delivery-saturation-coherence-risk-v0.1What this repo is for
Detect when
oxygen is prescribed with a target
but delivery
and achieved saturations
do not match
Examples you can use
oxygen prescribed but not given
COPD patient over-oxygenated
hypoxia persists with no titration
You use it to flag
respiratory harm risk
explainable-user-profilesExplained user profiles for 50 users based on IMDB Dataset created by GPT-4o.
agi-eval-sat-math-judgmentsbodh_llm_alpha_dataset
