datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipa-childes-split
IPA-CHILDES split
This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular,
the following changes have been implemented:
column processed_gloss dropped as it duplicates information of gloss up to punctuation
column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+)
column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package
columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.nyc-2025-marathon-splitssmart-logistics-delay-split-v0
Smart Logistics delay split v0 (frozen)
The 1,000-row Kaggle Smart Logistics Supply Chain Dataset (user ziya07, listed as CC0 on
https://www.kaggle.com/datasets/ziya07/smart-logistics-supply-chain-dataset when read on
2026-09-03) together with the exact 800/200 ShareGPT-style JSONL split on which
Yuchiwang02/Llama-3.2-1B-DelaySentinel
was fine-tuned in September 2025.
What this dataset makes possible
A complete, worked example of label leakage, small enough to… See the full description on the dataset page: https://huggingface.co/datasets/Yuchiwang02/smart-logistics-delay-split-v0.brainteaser_splitedi2p-adversarial-split
I2P - Adversarial Samples
We here provide a subset of the inappropriate image prompts (I2P) benchmark that are solid candidates for adversarial testing.
Specifically, all prompts in this dataset provided here are reasonably likely to produce inappropriate images and bypass the MidJourney prompt filter.
More details are provided in our AACL workshop paper: "Distilling Adversarial Prompts from Safety Benchmarks:
Report for the Adversarial Nibbler Challenge"
cebquad_splitCreated for the Cebuano Question Answering System. Articles were scraped from the SunStar Superbalita website and were pseudonymized. The dataset for the articles are found here.
Researcher:Jhoanna Rica Lagumbaylagumbay.jhoanna@gmail.comUniversity of the Philippines CebuDepartment of the Computer Science
TAU-2018-Real-splitsmini-split-ac-prices-raw-dataset-2026
25,761 raw U.S. mini-split AC price observations across 12 ZIP markets and 29 days.
Mini-Split AC Prices Raw Dataset (2026)
Analyze 25,761 unaggregated product-level listed retail prices for complete mini-split air-conditioning systems across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field.
What “raw” means here:… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/mini-split-ac-prices-raw-dataset-2026.eval2_180_split_inspection_posebasedalzheimerdataset_splitengine-condition-splits1_test_split_dataset
Mock Product Reviews Dataset
Dataset Description
A synthetic product review dataset for text classification and sentiment analysis tasks. The dataset contains user reviews across multiple product categories with ratings, sentiment labels, and metadata.
Dataset Summary
Total samples: 300
Train split: 210 samples (70.0%)
Validation split: 45 samples (15.0%)
Test split: 45 samples (15.0%)
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/ldmLDM77/1_test_split_dataset.stocks-splitscountdigit_max10000_sample50000_split0.8qwen35-9b-test-impl-split-coop
What this is
Cooperative two-agent coding dataset: 48 task pairs across 15 repos (random-50 subset), generated
with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a test-impl-split prompt variant —
one agent is assigned the role of writing tests; the other writes the implementation. The two
patches cover non-overlapping files, which eliminates merge conflicts entirely.
Key finding: This variant achieves a 100% clean merge rate (0 conflicts across all 48 pairs),
but a 0%… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-test-impl-split-coop.temporal_splittitanic-train-splittitanic-test-splitPhishFuzzer-split
Dataset Card
PhishingFuzzer splitted for DSR301m
MeAJOR-split
Dataset Card
Dataset Details
MeAJOR data splitted
Aya_dataset_potentially_misplaced_splitdataset_splits_v1Aya_dataset_potentially_mislabeled_splitarm_o0_phi4_multi_full_14_all_splitsdescription_to_claims_splitcountdigit_max1000000_sample50000_split0.8tourism-dataset-splitstourism-dataset-splits-12042026imdb_clustering_train_test_splitECGData_clean_splits
