datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
everyday-manipulation-3d-raw
Everyday Manipulation 3D (raw RGB-D)
1,513 clips · 10.28 hours · 279 GiB · 4 participants · 10 manipulation tasks · 42 recording sittings
Chest-mounted iPhone Pro capture of everyday two-handed manipulation by
CaryX AI. Clips were recorded with
Record3D, an iOS app that captures the
iPhone's LiDAR RGB-D stream. Each clip is the app's .r3d recording with the
audio track removed; the sensor streams are unmodified: synchronised RGB,
metric LiDAR depth, per-frame ARKit 6-DoF camera… See the full description on the dataset page: https://huggingface.co/datasets/CaryxAI/everyday-manipulation-3d-raw.credit-card-clients
Default of Credit Card Clients Dataset
The following was retrieved from UCI machine learning repository.
Dataset Information
This dataset contains information on default payments, demographic factors, credit data, history of payment, and bill statements of credit card clients in Taiwan from April 2005 to September 2005.
Content
There are 25 variables:
ID: ID of each client
LIMIT_BAL: Amount of given credit in NT dollars (includes individual and family/supplementary credit
SEX:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/credit-card-clients.uci-credit-card-defaultriftbound-cards
Riftbound TCG Card Database
Machine-readable snapshot of every card in
Riftbound: The League of Legends TCG, auto-scraped
from the official Card Gallery and errata pages.
Snapshot date: 2026-09-23
Cards: 1188
Source repo (scraper + pipeline): https://github.com/LouisCourrian/riftbound-cards
Every GitHub release publishes the same three formats as attached assets and
mirrors them here.
Files
File
What
cards.csv
Full corpus, one row per card. Array fields… See the full description on the dataset page: https://huggingface.co/datasets/Wysme/riftbound-cards.NBA-Player-Career-Stats
Dataset Description
This dataset contains a single CSV file with lifetime statistics for NBA players. The data includes various box score stats and personal information for each player's career.
Data Fields
The CSV file contains the following columns:
FULL_NAME: The player's full name
AST: Total career assists
BLK: Total career blocks
DREB: Total career defensive rebounds
FG3A: Total 3-point field goal attempts
FG3M: Total 3-point field goals made
FG3_PCT: 3-point field… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/NBA-Player-Career-Stats.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.naive-physics-ironing-v0.2
nAIve physics — Ironing Pilot v0.2 + Interaction Analysis v0.3
Visual Preview
Original RGB demonstration — IRON_009
▶ Watch IRON_009 original RGB demonstration
v0.3 interaction analysis — IRON_009
▶ Watch IRON_009 analysed interaction video
Raw → analysed: the first video is the original RGB demonstration; the second shows the v0.3 garment semantics, tool tracking, and temporal interaction analysis derived from the same episode.
A… See the full description on the dataset page: https://huggingface.co/datasets/CaramelCoffee19/naive-physics-ironing-v0.2.north-carolina-layoffs-warn-act-notices-daily
North Carolina WARN Act layoff notices — every filing we hold since 2014, one CSV, rebuilt daily
1,086 North Carolina WARN notices — every one this dataset holds, back to 2014 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-16
· state source last checked 2026-09-21T12:28Z · official source: North Carolina Department of Commerce — WARN notices.
North Carolina employers must file a WARN Act notice with the state before a… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/north-carolina-layoffs-warn-act-notices-daily.south-carolina-layoffs-warn-act-notices-daily
South Carolina WARN Act layoff notices — every filing we hold since 2013, one CSV, rebuilt daily
604 South Carolina WARN notices in the archive, back to the earliest filing this dataset holds · 604 of them are free to download · most recent notice filed 2026-08-28
· state source last checked 2026-09-20T12:28Z.
South Carolina employers must file a WARN Act notice with the state before a qualifying
mass layoff or plant closing. This page is generated from those filings, cleaned… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/south-carolina-layoffs-warn-act-notices-daily.cardiac_cine_mnms
M&Ms (Cardiac Cine-MRI)
Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation
Views: SAX (ED/ES)
Splits: train / val / test
Data Structure (per example)
sax_ed, sax_ed_gt
sax_es, sax_es_gt
Optional: sax_t (if present)
Metadata columns listed below
Columns
Imaging
pid
sax_ed, sax_ed_gt, sax_es, sax_es_gt
sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.carbon_24
Dataset Card for Carbon-24
Dataset Summary
Carbon-24 contains 10k carbon materials, which share the same composition, but have different structures. There is 1 element and the materials have 6 - 24 atoms in the unit cells.
Carbon-24 includes various carbon structures obtained via ab initio random structure searching (AIRSS) (Pickard & Needs, 2006; 2011) performed at 10 GPa.
The original dataset includes 101529 carbon structures, and we selected the 10% of the carbon… See the full description on the dataset page: https://huggingface.co/datasets/albertvillanova/carbon_24.car-reviewstrilemma-of-truth
Dataset Card for Trilemma of Truth (ToT) Dataset
🧾 Dataset Summary
The Trilemma of Truth (ToT) dataset serves as a benchmark for evaluating veracity probes across three distinct statement types:
Factually true statements.
Factually false statements.
Neither-valued statements are defined as those for which the language model lacks sufficient evidence to assign a truth value (see formal definition below).
The dataset includes three domain configurations:… See the full description on the dataset page: https://huggingface.co/datasets/carlomarxx/trilemma-of-truth.credit-card-transactionllm-bargaining-transcripts
LLM Bargaining Transcripts
240 complete two-agent bargaining games between large language models, played
under an alternating-offers protocol with private valuations, discounting, and
cheap talk. Every game records both agents' true valuations, their
private reasoning, what they claimed about their own position, and what
they actually did.
The dataset is designed to make misrepresentation measurable. Because the true
valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.CARD
CARD — Causal Recovery of Demand
Can a model that fits observed demand well still recover causal price response, substitution, and counterfactual outcomes when prices and promotions are endogenous?
CARD pairs synthetic retail scanner panels with marketing-copy product descriptions that carry the true substitution geometry. Demand is simulated from a known data-generating process; in half the cells, promotion depth responds to a hidden demand shock, so estimators that ignore… See the full description on the dataset page: https://huggingface.co/datasets/jean-jsj/CARD.craigslist-used-cars-eda
Craigslist Used Cars and Trucks: EDA
Overview
This dataset and notebook contain an Exploratory Data Analysis (EDA) of real Craigslist used-car listings scraped across the United States.
Main Question: What factors most influence the price of a used car listed on Craigslist?
Target Variable: price — the seller's asking price for each vehicle listing.
About the Dataset
Property
Details
Source
Kaggle — Austin Reese (scraped from Craigslist)
Original… See the full description on the dataset page: https://huggingface.co/datasets/Yoad22/craigslist-used-cars-eda.nba-career-stats-eda
🏀 NBA Player Career Stats — EDA Project
Overview
This project presents an end-to-end Exploratory Data Analysis (EDA) of NBA player
career statistics. The goal is to uncover patterns in player performance, compare
active vs. retired players, and explore relationships between key basketball stats.
Source: Hatman/NBA-Player-Career-Stats
Original size: 3,093 rows × 28 columns
Final clean size: 3,078 rows × 23 columns
Target Variable: IS_ACTIVE (True = Active / False =… See the full description on the dataset page: https://huggingface.co/datasets/Omerinbar/nba-career-stats-eda.car_dealershipRetail Car Dealership Data
Data for a car delearship. Perform EDA extract features and clean it up. Source Kaggle.
Try it out! It's primary goal is to provide an interface for users to download the dataset and try it out.
credit-cardPort of the credit-card dataset from UCI (link here). See details there and use carefully.
Basic preprocessing done by the imodels team in this notebook.
The target is the binary outcome default.payment.next.month.
Sample usage
Load the data:
from datasets import load_dataset
dataset = load_dataset("imodels/credit-card")
df = pd.DataFrame(dataset['train'])
X = df.drop(columns=['default.payment.next.month'])
y = df['default.payment.next.month'].values
Fit a model:
import… See the full description on the dataset page: https://huggingface.co/datasets/imodels/credit-card.mad-cars
MAD-Cars: Multi-view Auto Dataset 🚗
Dataset Description
MAD-Cars is a large-scale collection of 360° car videos.
It comprises ~70,000 car instances with diverse brands, car types, colors, and lighting conditions. Each instance contains an average of ~85 frames, with most car instances available at a resolution of 1920x1080. The dataset statistics are presented in the figure below. The data is carefully curated by filtering the frames and entire car instances that… See the full description on the dataset page: https://huggingface.co/datasets/yandex/mad-cars.cars_from_drom.ru_archive_2007-2025More information on the parsing process can be found here: https://github.com/zavzyatiy/drom_archive_parser.
This dataset is also published on Kaggle: https://www.kaggle.com/datasets/assaabramovich/resaled-cars-from-drom-ruarchive-2018-2023/.
Main dataset with all data: drom_archive_2007-2025_full.csv
Dataset with (almost) all configurations from Drom for cars in data: additional_data/drom-24-07-2025-all_main_cars_configurations.csv
Dataset with identification of regions for all cities in… See the full description on the dataset page: https://huggingface.co/datasets/zavzyatiy/cars_from_drom.ru_archive_2007-2025.cardiac_cine_mnms2
M&Ms2 (Cardiac Cine-MRI, RV Focus)
Processed NIfTI cine-MRI data derived from the M&Ms2 challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation (RV focus)
Views: SAX + LAX 4C (LAX 2C if present)
Splits: train / val / test
Data Structure (per example)
SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt
LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt
LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt
Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.credit-card-fraud-detection
Credit Card Fraud Detection – Processed Dataset
This dataset contains preprocessed credit card transaction data prepared for fraud detection tasks.
Data Description
The dataset is derived from anonymized transaction records and includes numerical features (V1–V28), transaction amount, and time-based information.
Preprocessing Steps
Feature scaling and normalization
Handling class imbalance
Feature selection based on correlation analysis
Removal of irrelevant… See the full description on the dataset page: https://huggingface.co/datasets/jyunyilin/credit-card-fraud-detection.floradb-houseplants-care-sample
🌿 FloraDB — Houseplant Care & Pet-Toxicity Dataset (Free Sample)
Full dataset: floradb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A free sample of FloraDB: a structured dataset that turns subjective houseplant care advice — "bright indirect light", "water when dry" — into quantitative engineering metrics (Lux thresholds, watering-day intervals, temperature and humidity ranges), joined to ASPCA dog/cat toxicity and grounded on the GBIF… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/floradb-houseplants-care-sample.carimages-rights
CarImages Rights Manifest
A rights-resolved index of 73,618 Creative Commons and public-domain
photographs of cars, published by carimages.org.
🔗 Browse the archive at carimages.org →
Dataset home ·
Licensing guide ·
Attribution index ·
Photographers ·
Tools
This dataset contains no images. It is metadata and URLs only, so no ShareAlike
obligation propagates to you for using this file. Each photograph remains under the
licence recorded in its own row; license_url… See the full description on the dataset page: https://huggingface.co/datasets/carimages/carimages-rights.Politifact_fake_newscardiac_cine_emidec
EMIDEC (Delayed-Enhancement CMR)
Processed NIfTI EMIDEC dataset for myocardial infarction assessment.
Dataset Summary
Modality: DE-MRI
Task: Segmentation (LV, myocardium, infarction, no-reflow)
Splits: train / val / test
Data Structure (per example)
image
label
Metadata columns listed below
Columns
Imaging
pid
image, label
Metadata (all columns)
gender, age, tobacco, overweight, arterial_hypertension, diabetes
family_history, ecg, troponin… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_emidec.credit_card_fraud_transactions193k-prices-period-care-atlas
192,500 prices: 9 categories, 12 U.S. ZIPs, 29 days
Period Care Prices Raw Dataset (2026)
How do listed and package-standardized prices for reusable and disposable period-care products vary across U.S. ZIP markets and days?
This fixed research snapshot contains 192,500 unaggregated, quality-filtered price observations across 9 period-care categories, 12 U.S. ZIP markets, and 29 consecutive dates from July 21 through August 18, 2026. The analysis-ready CSV preserves product… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/193k-prices-period-care-atlas.
