datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pettahdambulla_vegforecastindustrial-sensor-anomaly-data
Industrial Equipment Sensor Anomaly Data
Overview
Synthetic multivariate sensor data from a simulated manufacturing plant with 5 equipment units (EQ-001 through EQ-005). Each unit generates 10,000 one-minute-interval readings across 11 sensor channels, 2 metadata fields, 3 derived features, and equipment operating mode labels.
The dataset is designed for anomaly detection benchmarking. It embeds 4 distinct anomaly types at approximately 4.5% prevalence:
Thermal runaway —… See the full description on the dataset page: https://huggingface.co/datasets/Petsteb/industrial-sensor-anomaly-data.jobcannon-psychometric-responses
JobCannon Psychometric Response Dataset
v3 — 54,431 item-level responses across nine instruments and 25 languages.
Anonymized, item-level responses to nine open-domain psychometric instruments,
collected from real test-takers on JobCannon. Each row is
one completed assessment: the raw per-item answers, the computed dimensional
scores, and the dominant result type.
This is a first-party dataset — our own users' responses, not a
re-publication of someone else's data.… See the full description on the dataset page: https://huggingface.co/datasets/PeterKol/jobcannon-psychometric-responses.flight-mh370-revisited-data
Flight MH370 Revisited — oversized source files
Companion data for the research repository
gmkf7vfyfb-web/flight-mh370-revisited.
That repository holds the complete project tree — model code, source data,
reports, figures, posterior outputs, the 148-file source-library audit and the
handoff dossier — from the consolidated snapshot of 14 August 2026. Seven files
were too large to keep in Git (one is 336 MB, above GitHub's hard 100 MB
per-file limit), so they live here instead.… See the full description on the dataset page: https://huggingface.co/datasets/peteabiome/flight-mh370-revisited-data.PETA_TEM_Sol
PETA_TEM_Sol Dataset
Description: Solubility mutation dataset.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
Github
PETA: evaluating the impact of protein transfer learning with sub-word tokenization on downstream applications
https://github.com/ginnm/ProteinPretraining
Citation
Please cite our work if you use our dataset.
@article{tan2024peta,
title={PETA: evaluating the impact of protein transfer learning with… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/PETA_TEM_Sol.wet-vs-dry-pet-food-prices-raw-dataset-2026
44,583 prices: 4 categories, 12 U.S. ZIPs, 29 days
Wet vs Dry Pet Food Prices Raw Dataset (2026)
How do listed and package-standardized prices compare across wet and dry dog and cat food, 12 selected U.S. ZIP markets, and 29 days?
This fixed research snapshot contains 44,583 unaggregated, quality-filtered price observations across 4 categories, 12 U.S. ZIP markets, and 29 consecutive dates from July 21 through August 18, 2026. The analysis-ready CSV preserves product titles… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/wet-vs-dry-pet-food-prices-raw-dataset-2026.PETA_LGK_Sol
PETA_LGK_Sol Dataset
Description: Solubility mutation dataset.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
Github
PETA: evaluating the impact of protein transfer learning with sub-word tokenization on downstream applications
https://github.com/ginnm/ProteinPretraining
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/PETA_LGK_Sol.PETA_CHS_Sol
PETA_CHS_Sol Dataset
Description: Solubility mutation dataset.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
Github
PETA: evaluating the impact of protein transfer learning with sub-word tokenization on downstream applications
https://github.com/ginnm/ProteinPretraining
Citation
Please cite our work if you use our dataset.
@article{tan2024peta,
title={PETA: evaluating the impact of protein transfer learning with… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/PETA_CHS_Sol.pet-airlines
Pet Air Travel Dataset (Japan Outbound)
A structured, machine-readable dataset of pet travel policies for major international airlines, focused on flights departing from Japan. Designed to be directly consumable by AI agents and downstream applications without scraping.
Overview
The web is full of human-readable pet travel guides; almost none are usable by an AI agent or a typed application without per-airline scraping. pet_airlines provides the missing structured… See the full description on the dataset page: https://huggingface.co/datasets/Aulvem/pet-airlines.pet-food-recall-risk
Pet Food Recall Risk Classification
A small supervised multi-label text classification dataset for categorising
pet food recall and safety-alert records into risk categories.
Built as an academic assignment for an Information Retrieval course.
All source records come from official public recall and safety-alert portals.
Task
Supervised multi-label text classification.
Given a structured text constructed from brand name, product description, and
recall reason, predict one… See the full description on the dataset page: https://huggingface.co/datasets/ShurongSR/pet-food-recall-risk.119k-prices-petflation-2026
119,316 prices: 11 categories, 12 U.S. ZIPs, 29 days
119K Prices: Petflation 2026
How do pet food, treats, litter, grooming, cleanup supplies, and training-pad prices compare across U.S. ZIP markets when package quantities are standardized within each category?
This fixed research snapshot contains 119,316 unaggregated, quality-filtered retail price observations across 11 pet-care categories, 12 U.S. ZIP markets, and 29 consecutive dates from July 21 through August 18, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/119k-prices-petflation-2026.jobcannon-entertainment-responses
JobCannon Entertainment Quiz Response Dataset
v1. 77,284 item-level responses to five for-fun quizzes, in 24 languages.
Anonymized, item-level answers to five entertainment quizzes taken by real
visitors on JobCannon between March and August 2026.
One row is one completed quiz: the raw per-item answers, the category totals the
site computed from them, and the result the taker was shown.
We collected all of it on our own traffic. It is not a repackaging of somebody
else's file.… See the full description on the dataset page: https://huggingface.co/datasets/PeterKol/jobcannon-entertainment-responses.peta-indonesia-ikpturkish_pets_balanced_datasetAugmented_CIC-IDS2017tech-ready-restaurants-in-the-tampa-st-petersburg-clearwater-metro-area-fl-us-176638
Tech-Ready Restaurants in the Tampa-St. Petersburg-Clearwater Metro Area, FL, US
Free sample dataset from BeamStation
The dataset "Tech-Ready Restaurants in the Tampa-St. Petersburg-Clearwater Metro Area, FL, US" provides a weekly‑updated list of 89 dining establishments that meet specific criteria for technology adoption. Each record represents a restaurant with a Beam Score above 70, indicating strong financial health and market presence, and a sentiment score from the last 30… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/tech-ready-restaurants-in-the-tampa-st-petersburg-clearwater-metro-area-fl-us-176638.petra-sadouski-moi-shybolet-autabiiagrafichnyia-arabeski-petra-sadouski
Мой шыболет. Аўтабіяграфічныя арабэскі
Metadata
Author: Пётра Садоўскі
Title: Мой шыболет. Аўтабіяграфічныя арабэскі
Narrator: Пётра Садоўскі
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/petra-sadouski-moi-shybolet-autabiiagrafichnyia-arabeski-petra-sadouski.evalo_faqpetro9recipes
