datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pettahdambulla_vegforecastglobal-top-Index-exploring-trends-in-stock-Market
Global Top Index: Exploring Trends in Stock Markets
About the Dataset
The Global Top Index dataset offers a detailed view of daily trading activities from several of the world's leading stock market indices. This dataset is ideal for conducting comprehensive analyses to uncover insights and predictive trends in the international stock markets.
Dataset Contents
The dataset encompasses the following key data points for each trading session across multiple dates… See the full description on the dataset page: https://huggingface.co/datasets/pettah/global-top-Index-exploring-trends-in-stock-Market.RealEdit
RealEdit Dataset
RealEdit is a large-scale, authentic dataset of image edits collected from Reddit's r/PhotoshopRequest and r/estoration. It is divided into two splits: train and test.
Note: This dataset contains image URLs that may become inactive over time. If you are a researcher and would like access to archived images, please fill out this Google Form.
We recommend using gallery-dl to download the images, though you are free to use any method.
Test Split… See the full description on the dataset page: https://huggingface.co/datasets/peter-sushko/RealEdit.Ads_Creative_Ad_Copy_Programmatic
Dataset Summary
The Programmatic Ad Creatives dataset contains 7097 samples of online programmatic ad creatives along with their ad sizes. The dataset includes 8 unique ad sizes, such as (300, 250), (728, 90), (970, 250), (300, 600), (160, 600), (970, 90), (336, 280), and (320, 50). The dataset is in a tabular format and represents a random sample from Project300x250.com's complete creative data set. It is primarily used for training and evaluating natural language processing models… See the full description on the dataset page: https://huggingface.co/datasets/PeterBrendan/Ads_Creative_Ad_Copy_Programmatic.industrial-sensor-anomaly-data
Industrial Equipment Sensor Anomaly Data
Overview
Synthetic multivariate sensor data from a simulated manufacturing plant with 5 equipment units (EQ-001 through EQ-005). Each unit generates 10,000 one-minute-interval readings across 11 sensor channels, 2 metadata fields, 3 derived features, and equipment operating mode labels.
The dataset is designed for anomaly detection benchmarking. It embeds 4 distinct anomaly types at approximately 4.5% prevalence:
Thermal runaway —… See the full description on the dataset page: https://huggingface.co/datasets/Petsteb/industrial-sensor-anomaly-data.jobcannon-psychometric-responses
JobCannon Psychometric Response Dataset
v3 — 54,431 item-level responses across nine instruments and 25 languages.
Anonymized, item-level responses to nine open-domain psychometric instruments,
collected from real test-takers on JobCannon. Each row is
one completed assessment: the raw per-item answers, the computed dimensional
scores, and the dominant result type.
This is a first-party dataset — our own users' responses, not a
re-publication of someone else's data.… See the full description on the dataset page: https://huggingface.co/datasets/PeterKol/jobcannon-psychometric-responses.PETraPETra
Description:
We introduce PETra, the first multilingual corpus and detection framework for pragmatic explicitation. The corpus consists of 3,000 sentence pairs from TED-Multi and Europarl, covers twelve language pairs, and includes additions such as entity descriptions, measurement conversions, and translator remarks. We identify candidate explicitation cases through null alignments and refined using active learning with human annotation.
Citation:
If you use this dataset, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/Doosme/PETra.Ads_Creative_Text_Programmatic
Dataset Summary
The Programmatic Ad Creatives dataset contains 1000 samples of online programmatic ad creatives along with their ad sizes. The dataset includes 8 unique ad sizes, such as (300, 250), (728, 90), (970, 250), (300, 600), (160, 600), (970, 90), (336, 280), and (320, 50). The dataset is in a tabular format and represents a random sample from Project300x250.com's complete creative data set. It is primarily used for training and evaluating natural language processing models… See the full description on the dataset page: https://huggingface.co/datasets/PeterBrendan/Ads_Creative_Text_Programmatic.flight-mh370-revisited-data
Flight MH370 Revisited — oversized source files
Companion data for the research repository
gmkf7vfyfb-web/flight-mh370-revisited.
That repository holds the complete project tree — model code, source data,
reports, figures, posterior outputs, the 148-file source-library audit and the
handoff dossier — from the consolidated snapshot of 14 August 2026. Seven files
were too large to keep in Git (one is 336 MB, above GitHub's hard 100 MB
per-file limit), so they live here instead.… See the full description on the dataset page: https://huggingface.co/datasets/peteabiome/flight-mh370-revisited-data.PETA_TEM_Sol
PETA_TEM_Sol Dataset
Description: Solubility mutation dataset.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
Github
PETA: evaluating the impact of protein transfer learning with sub-word tokenization on downstream applications
https://github.com/ginnm/ProteinPretraining
Citation
Please cite our work if you use our dataset.
@article{tan2024peta,
title={PETA: evaluating the impact of protein transfer learning with… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/PETA_TEM_Sol.crpe_vlmevalkitwet-vs-dry-pet-food-prices-raw-dataset-2026
44,583 prices: 4 categories, 12 U.S. ZIPs, 29 days
Wet vs Dry Pet Food Prices Raw Dataset (2026)
How do listed and package-standardized prices compare across wet and dry dog and cat food, 12 selected U.S. ZIP markets, and 29 days?
This fixed research snapshot contains 44,583 unaggregated, quality-filtered price observations across 4 categories, 12 U.S. ZIP markets, and 29 consecutive dates from July 21 through August 18, 2026. The analysis-ready CSV preserves product titles… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/wet-vs-dry-pet-food-prices-raw-dataset-2026.PETA_LGK_Sol
PETA_LGK_Sol Dataset
Description: Solubility mutation dataset.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
Github
PETA: evaluating the impact of protein transfer learning with sub-word tokenization on downstream applications
https://github.com/ginnm/ProteinPretraining
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/PETA_LGK_Sol.PETA_CHS_Sol
PETA_CHS_Sol Dataset
Description: Solubility mutation dataset.
Number of labels: 1
Problem Type: regression
Columns:
aa_seq: protein amino acid sequence
Github
PETA: evaluating the impact of protein transfer learning with sub-word tokenization on downstream applications
https://github.com/ginnm/ProteinPretraining
Citation
Please cite our work if you use our dataset.
@article{tan2024peta,
title={PETA: evaluating the impact of protein transfer learning with… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/PETA_CHS_Sol.pet-airlines
Pet Air Travel Dataset (Japan Outbound)
A structured, machine-readable dataset of pet travel policies for major international airlines, focused on flights departing from Japan. Designed to be directly consumable by AI agents and downstream applications without scraping.
Overview
The web is full of human-readable pet travel guides; almost none are usable by an AI agent or a typed application without per-airline scraping. pet_airlines provides the missing structured… See the full description on the dataset page: https://huggingface.co/datasets/Aulvem/pet-airlines.pet-food-recall-risk
Pet Food Recall Risk Classification
A small supervised multi-label text classification dataset for categorising
pet food recall and safety-alert records into risk categories.
Built as an academic assignment for an Information Retrieval course.
All source records come from official public recall and safety-alert portals.
Task
Supervised multi-label text classification.
Given a structured text constructed from brand name, product description, and
recall reason, predict one… See the full description on the dataset page: https://huggingface.co/datasets/ShurongSR/pet-food-recall-risk.119k-prices-petflation-2026
119,316 prices: 11 categories, 12 U.S. ZIPs, 29 days
119K Prices: Petflation 2026
How do pet food, treats, litter, grooming, cleanup supplies, and training-pad prices compare across U.S. ZIP markets when package quantities are standardized within each category?
This fixed research snapshot contains 119,316 unaggregated, quality-filtered retail price observations across 11 pet-care categories, 12 U.S. ZIP markets, and 29 consecutive dates from July 21 through August 18, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/119k-prices-petflation-2026.ro-political-ai
Dataset Card for "RO-Political-Texts"
Dataset Description
Dataset Summary
RO-Political-AI is a collection of datasets for RoNLP, focused on detecting the difference between political texts written by humans, during the presidential elections from Romania (2024-2025), and synthetic texts generated with LLM models, as well as on studying linguistic characteristics such as slang, idioms, figurative expressions and "monkey business" behaviors in political discourse… See the full description on the dataset page: https://huggingface.co/datasets/petrematei/ro-political-ai.Jordan-Peterson-Conversation-for-NLPThis dataset contains dialogues from Jordan Peterson through either quora answers or interview transcripts. The dataset was manually created to imitate conversation.
ro-political
Dataset Card for "RO-Political-Texts"
Dataset Description
Dataset Summary
RO-Political Corpus is a collection of datasets for RoNLP, focused on detecting the difference between political texts written by humans, during the presidential elections from Romania (2024-2025), and synthetic texts generated with LLM models, as well as on studying linguistic characteristics such as slang, idioms, figurative expressions and "monkey business" behaviors in political… See the full description on the dataset page: https://huggingface.co/datasets/petrematei/ro-political.petcare_sampleCOFOG-feedbackstackoverflow-kubernetes-questionscovert from https://huggingface.co/datasets/mcipriano/stackoverflow-kubernetes-questions/blob/main/README.md
format from parquet to csv
coverting code as below
import pandas as pd
from pandas import read_parquet
data = read_parquet("~/Downloads/kubernetes_dump.parquet")
#print(data.count())
#data.head()
data.to_csv('/tmp/out.csv', index=False)
jobcannon-entertainment-responses
JobCannon Entertainment Quiz Response Dataset
v1. 77,284 item-level responses to five for-fun quizzes, in 24 languages.
Anonymized, item-level answers to five entertainment quizzes taken by real
visitors on JobCannon between March and August 2026.
One row is one completed quiz: the raw per-item answers, the category totals the
site computed from them, and the result the taker was shown.
We collected all of it on our own traffic. It is not a repackaging of somebody
else's file.… See the full description on the dataset page: https://huggingface.co/datasets/PeterKol/jobcannon-entertainment-responses.3D_Synthetic_Petroleum_Derived_GEMS
3D Synthetic Petroleum-Derived GEMS (MVP Release)
📌 Dataset Overview
This dataset contains 197 elite, highly complex 3D molecular structures derived from petroleum fractions. Designed specifically for petrochemicals, materials science, organic semiconductors, and specialized additives, these compounds represent a curated "Golden Fund" of stable, complex hydrocarbons.
Unlike drug-like molecules, this dataset focuses on polycyclic architectures, rigid 3-ring… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/3D_Synthetic_Petroleum_Derived_GEMS.osworld_tasks_filespeta-indonesia-ikpturkish_pets_balanced_datasetCodeNet_Python_118Augmented_CIC-IDS2017
