datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ICPC_Data
ICPC World Finals — a discriminative subset, with model traces
24 ICPC World Finals problems (2021–2025), together with the full transcripts of an
LLM attempting each of them three times under simulated contest rules.
Selection
The model
Every run in this dataset comes from:
nvidia/Nemotron-Cascade-2-30B-A3B
The partitions
Every one of the 53 problems was run 3 times (seeds 1, 2, 3). Each problem was then
placed by its pass rate and… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.ICLR2024-paper-statsICLR_2025_Accepted_PapersPaper Decision Results for ICLR 2025
ICLR 2025 Accepted Paper List
https://openreview.net/group?id=ICLR.cc/2025/Conference#tab-accept
multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.appliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.whiskydb-fine-spirits-sample
🥃 WhiskyDB — Fine Spirits & Whisky Dataset (Free Sample)
Full dataset: whiskydb.dataengineered.io · $49 one-time (or $49 / month with the monthly refresh) → Buy once · Subscribe · the same sample on Kaggle
A free sample of WhiskyDB: a structured, relational dataset of whiskies and fine spirits built entirely from open, legally accessible public sources — government label registries (US TTB COLA), corporate registries (UK Companies House), the EU eAmbrosia GI register, Open… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/whiskydb-fine-spirits-sample.recalldb-product-recalls-sample
RecallDB — U.S. Product Recall Database (Sample)
Full dataset: recalldb.dataengineered.io · $49 one-time snapshot → Buy on Stripe · the same sample on Kaggle
127,783 official recalls · 292,790 recalled products · CPSC · FDA · FSIS · NHTSA · USCG · 100% source-linked
RecallDB is a normalized, provenance-tracked dataset of official U.S. federal product recalls. It joins five official source families into one relational model: CPSC consumer products, NHTSA vehicles, FDA/openFDA… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/recalldb-product-recalls-sample.winedb-fine-wines-and-vintages
🍷 WineDB — Fine Wine & Vintages Sample Dataset
Full dataset: winedb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A curated free sample of the WineDB dataset: highly normalized relational tables tracking fine wine producers, cuvees, exact vintage varietal blend percentages (SUM <= 100.001, enforced by SQLite triggers), alcohol content (ABV %), organoleptic tasting descriptors, and secondary market valuation indices. Prefer SQLite? The same… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/winedb-fine-wines-and-vintages.floradb-houseplants-care-sample
🌿 FloraDB — Houseplant Care & Pet-Toxicity Dataset (Free Sample)
Full dataset: floradb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A free sample of FloraDB: a structured dataset that turns subjective houseplant care advice — "bright indirect light", "water when dry" — into quantitative engineering metrics (Lux thresholds, watering-day intervals, temperature and humidity ranges), joined to ASPCA dog/cat toxicity and grounded on the GBIF… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/floradb-houseplants-care-sample.International_Classification_Diseases_Clinical_Modification_icd10cm_order_April_2024vn-ict-provincial
Vietnam ICT Index provincial panel (2006-2020. gaps 2008, 2021)
Vietnam ICT Index provincial panel from official annual reports. Years 2006-2007 are full_64 (incl. Hà Tây). 2009-2020 are full_63. No standalone cover-year surveys for 2008 or 2021 (2008 Bang18 column = Index 2007 published Dec 2008. 2021 skipped - next cover year is 2022. do not confuse with DTI 2021). From 2016 the overall table uses 3 pillars (HTKT, HTNL, ƯD). earlier years keep five pillars including SXKD and… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-ict-provincial.ResourceEstimation_HLOGenCNN
HLO Feature Dataset for Deep Learning Resource Estimation
Dataset Summary
The HLO Feature Dataset is a collection of compiler-level graph features (HLO graphs) extracted from deep learning training workloads. Alongside detailed metadata (model configs, GPU stats), this dataset enables machine learning approaches for:
⏱️ Training Time Prediction
📉 Resource Consumption Estimation
⚡ HPC and GPU Scheduling Optimization
🧩 Graph-based Neural Architecture Analysis
This… See the full description on the dataset page: https://huggingface.co/datasets/ICICLE-AI/ResourceEstimation_HLOGenCNN.iclr2026-lm-logprobs
LM Log-Probabilities for Value Bias Analysis
Next-token log-probability distributions from 12 language models across 54 prompts, used in the paper:
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A.F. Thompson, Elle, Vincent Adam, Hannah Rose Kirk, Christopher Summerfield, Tsvetomira Dumbalska (ICLR 2026)
Part of the Oxford-HIPlab collection for this paper.
Dataset description
Each CSV contains the full next-token log-probability… See the full description on the dataset page: https://huggingface.co/datasets/Oxford-HIPlab/iclr2026-lm-logprobs.isaac-gr00t-ikea-training-report
Isaac GR00T IKEA training report
Portable export of the W&B run unitree_g1_ikea_batch32_20260730. The run stopped after a clean host
shutdown; the last W&B metric is step 31,750 and the
last complete checkpoint is 31,000.
Summary
First logged training loss: 1.4667
Last logged training loss: 0.1256
Lowest 1,000-step rolling loss: 0.1249 at step 31,750
Mean GPU compute utilization: 51.8%
Mean allocated GPU memory: 29.0%
Provisional checkpoint choice: 30,000… See the full description on the dataset page: https://huggingface.co/datasets/ICRA-Competitions/isaac-gr00t-ikea-training-report.roasterdb-specialty-coffee-sample
☕ RoasterDB — Specialty Coffee Dataset (Free Sample)
Full dataset: roasterdb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A free sample of RoasterDB: a structured dataset of specialty-coffee products scraped from the direct storefronts of curated artisan roasters worldwide, with tasting notes normalized to the Specialty Coffee Association (SCA) Flavor Wheel and a source URL on every record so any fact can be re-verified.
This sample contains 100… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/roasterdb-specialty-coffee-sample.permutation-pools
Permutation Pools
Permutation-augmentation pools: demonstration-order permutations of the label-transfer and evidence-tracking sets, used to enlarge filtered test sets to 1000 rows with balanced golds.
Part of the icl-heads collection: the datasets behind a study of which
Llama-3.1-8B-Instruct attention heads mediate in-context evidence accumulation
and in-context label mapping, discovered with Differentiable Circuit Masking
(a learned per-head mask over K/V activations patched… See the full description on the dataset page: https://huggingface.co/datasets/icl-heads/permutation-pools.label-transfer
Label Transfer
Label-transfer datasets: the same questions in the prompt and counterfactual contexts, with the answer vocabulary replaced (Yes->Ba / No->Ga) or swapped (Yes->No / No->Yes), at 4 and 8 demonstrations.
Part of the icl-heads collection: the datasets behind a study of which
Llama-3.1-8B-Instruct attention heads mediate in-context evidence accumulation
and in-context label mapping, discovered with Differentiable Circuit Masking
(a learned per-head mask over K/V… See the full description on the dataset page: https://huggingface.co/datasets/icl-heads/label-transfer.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/harryCJ/multimodal-ICS-provenance.surface-audit
TruthfulQA-476 — a surface-form-cleaned binary-choice TruthfulQA
TruthfulQA-476 is the recommended drop-in replacement for the binary-choice TruthfulQA
evaluation set. It keeps 476 of the 790 original question pairs, in the original schema, chosen so
that a classifier restricted to six surface features of the answer text (negation, hedging, length,
token statistics) can no longer separate correct from incorrect answers above chance, while the
ranking of models on the subset… See the full description on the dataset page: https://huggingface.co/datasets/iclr2027-surface-audit/surface-audit.wireless-vendor-identifiers
Wireless Vendor Identifier Metadata
This dataset contains public wireless and Bluetooth metadata relevant to BLE
research. It is designed for defensive analysis, asset classification,
literature support, and reproducible metadata lookup. The dataset is split into
multiple tables so it is not just an OUI mirror.
It deliberately excludes complete BLE advertisement payloads, packet templates,
radio timing recipes, device captures, effectiveness labels, and any instructions
for… See the full description on the dataset page: https://huggingface.co/datasets/ictrun/wireless-vendor-identifiers.icml26-repro-nonlinear-autoencoder-pca
Nonlinear autoencoder / PCA reproduction
This bundle reproduces the paper with the authors' pinned release
(SPOC-group/advantage_nonlinearity@4378017) plus an independent direct
population-gradient-flow audit.
run_official_amp_campaign.py: 96 finite-dimensional runs of the released
two-spike AMP at d=256/512 and eight sample ratios.
run_official_ae_campaign.py: 48 full-batch Adam runs of the released tied,
one-neuron ReLU/ELU autoencoder at d=512/1024, with exact PCA and 20,000… See the full description on the dataset page: https://huggingface.co/datasets/SabaPivot/icml26-repro-nonlinear-autoencoder-pca.datasets_for_icdeBTC-Data-1Hour-2018-2023
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/IceMasterT/BTC-Data-1Hour-2018-2023.deciban
deciban — a vote-level corpus for studying diversity of thought in LLM ensembles
Ten open-weight language models (20–31B parameters), each answering every question of
three benchmarks 24 times at temperature 1.0, with every individual vote retained:
767,520 multiple-choice inferences and 23,520 graded free-text forensic trials, plus
parse/termination reasons and abstentions as first-class outcomes. The corpus exposes the
joint answer distribution between models — which questions… See the full description on the dataset page: https://huggingface.co/datasets/IcyApril/deciban.iced-coffee-and-cold-brew-prices-raw-dataset-2026
60,942 raw U.S. iced coffee and cold brew price observations across 12 ZIP markets and 29 days.
Iced Coffee & Cold Brew Prices Raw Dataset (2026)
Analyze 60,942 unaggregated product-level listed retail prices for iced coffee, cold brew, and liquid coffee concentrates across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field.… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/iced-coffee-and-cold-brew-prices-raw-dataset-2026.DPO_ID-Wiki_10kTesting
HOW TO WRANGLING THIS DATASET TO DPO & CHATML FORMAT
def return_prompt_and_responses(samples) -> dict[str, str, str]:
return {
"prompt": [
"<|im_start|>user\n" + i + "<|im_end|>\n"
for i in samples["PROMPT"]
],
"chosen": [
"<|im_start|>assistant\n" + j + "<|im_end|>"
for j in samples["CHOSEN"]
],
"rejected": [
"<|im_start|>assistant\n" + k + "<|im_end|>"
for k in… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/DPO_ID-Wiki_10kTesting.scaphandre_ram_usage
Scaphandre RAM Usage Dataset
Dataset Description
This dataset contains RAM usage monitoring data collected using Scaphandre for the ICOS Federated Learning infrastructure.
Overview
Source: Scaphandre energy monitoring tool
Collection Method: Live system monitoring (continuous fetching)
Purpose: Training data for ICOS FL
Update Frequency: Real-time collection with 3s intervals
Data Schema
Column
Type
Description
timestamp
float
Unix… See the full description on the dataset page: https://huggingface.co/datasets/ICOS-AI/scaphandre_ram_usage.BTC-Data-Daily-2014-2023
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/IceMasterT/BTC-Data-Daily-2014-2023.clinical-quad-stress-buffer-lag-coupling-icu-collapse-v0.4
What this repo does
This repository contains a Clarus v0.4 cascade boundary discovery dataset modeling ICU collapse.
Earlier Clarus datasets focused on detecting cascade states or forecasting collapse trajectories.
Version v0.4 extends the framework to a harder task:
detecting whether a system lies on the instability boundary itself.
The dataset models ICU deterioration as a coupled physiological system in which rising system stress, declining physiological buffer, delayed… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-stress-buffer-lag-coupling-icu-collapse-v0.4.clinical-quad-stress-buffer-lag-coupling-icu-collapse-v0.6
What this repo does
This repository contains a Clarus v0.6 intervention pathway dataset focused on ICU collapse dynamics.
The dataset evaluates whether a model can determine if a proposed intervention meaningfully stabilizes a deteriorating critical care system.
The task requires reasoning from:
system state
trajectory toward instability
boundary geometry
recovery geometry
intervention vector
projected trajectory consequence
The model cannot read the answer directly.
It must infer… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-stress-buffer-lag-coupling-icu-collapse-v0.6.
