datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
e-CARE
Dataset of (Du et al., 2022) (Unofficial reupload)
Abstract
Understanding causality has vital importance for various Natural Language Processing (NLP) applications. Beyond the labeled instances, conceptual explanations of the causality can provide deep understanding of the causal fact to facilitate the causal reasoning process. However, such explanation information still remains absent in existing causal reasoning resources. In this paper, we fill this gap by presenting… See the full description on the dataset page: https://huggingface.co/datasets/12ml/e-CARE.everyday-manipulation-3d-raw
Everyday Manipulation 3D (raw RGB-D)
1,513 clips · 10.28 hours · 279 GiB · 4 participants · 10 manipulation tasks · 42 recording sittings
Chest-mounted iPhone Pro capture of everyday two-handed manipulation by
CaryX AI. Clips were recorded with
Record3D, an iOS app that captures the
iPhone's LiDAR RGB-D stream. Each clip is the app's .r3d recording with the
audio track removed; the sensor streams are unmodified: synchronised RGB,
metric LiDAR depth, per-frame ARKit 6-DoF camera… See the full description on the dataset page: https://huggingface.co/datasets/CaryxAI/everyday-manipulation-3d-raw.riftbound-cards
Riftbound TCG Card Database
Machine-readable snapshot of every card in
Riftbound: The League of Legends TCG, auto-scraped
from the official Card Gallery and errata pages.
Snapshot date: 2026-09-23
Cards: 1188
Source repo (scraper + pipeline): https://github.com/LouisCourrian/riftbound-cards
Every GitHub release publishes the same three formats as attached assets and
mirrors them here.
Files
File
What
cards.csv
Full corpus, one row per card. Array fields… See the full description on the dataset page: https://huggingface.co/datasets/Wysme/riftbound-cards.NBA-Player-Career-Stats
Dataset Description
This dataset contains a single CSV file with lifetime statistics for NBA players. The data includes various box score stats and personal information for each player's career.
Data Fields
The CSV file contains the following columns:
FULL_NAME: The player's full name
AST: Total career assists
BLK: Total career blocks
DREB: Total career defensive rebounds
FG3A: Total 3-point field goal attempts
FG3M: Total 3-point field goals made
FG3_PCT: 3-point field… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/NBA-Player-Career-Stats.ats-career-page-urls
ATS Career Page URLs
69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR.
Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines.
Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/latmay/ats-career-page-urls.RCEdit-500K
RCEdit-500K
Reference Completion for Image-Conditioned Image Editing
ECCV 2026
Dataset Overview
RCEdit-500K is the first large-scale unified open-source dataset for image-conditioned image editing (ICIE). It contains approximately 477K aligned quadruplets (input image, reference image, editing instruction, target image) across six edit categories.
Split
Samples
Train
472,050
Val
4,395
Edit Types
Type… See the full description on the dataset page: https://huggingface.co/datasets/carpedkm/RCEdit-500K.Low-Carbon-London-Smart-Meter-Cleaned-FeatureReadycardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.naive-physics-ironing-v0.2
nAIve physics — Ironing Pilot v0.2 + Interaction Analysis v0.3
Visual Preview
Original RGB demonstration — IRON_009
▶ Watch IRON_009 original RGB demonstration
v0.3 interaction analysis — IRON_009
▶ Watch IRON_009 analysed interaction video
Raw → analysed: the first video is the original RGB demonstration; the second shows the v0.3 garment semantics, tool tracking, and temporal interaction analysis derived from the same episode.
A… See the full description on the dataset page: https://huggingface.co/datasets/CaramelCoffee19/naive-physics-ironing-v0.2.north-carolina-layoffs-warn-act-notices-daily
North Carolina WARN Act layoff notices — every filing we hold since 2014, one CSV, rebuilt daily
1,086 North Carolina WARN notices — every one this dataset holds, back to 2014 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-16
· state source last checked 2026-09-21T12:28Z · official source: North Carolina Department of Commerce — WARN notices.
North Carolina employers must file a WARN Act notice with the state before a… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/north-carolina-layoffs-warn-act-notices-daily.south-carolina-layoffs-warn-act-notices-daily
South Carolina WARN Act layoff notices — every filing we hold since 2013, one CSV, rebuilt daily
604 South Carolina WARN notices in the archive, back to the earliest filing this dataset holds · 604 of them are free to download · most recent notice filed 2026-08-28
· state source last checked 2026-09-20T12:28Z.
South Carolina employers must file a WARN Act notice with the state before a qualifying
mass layoff or plant closing. This page is generated from those filings, cleaned… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/south-carolina-layoffs-warn-act-notices-daily.lichess-elitecardiac_cine_mnms
M&Ms (Cardiac Cine-MRI)
Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation
Views: SAX (ED/ES)
Splits: train / val / test
Data Structure (per example)
sax_ed, sax_ed_gt
sax_es, sax_es_gt
Optional: sax_t (if present)
Metadata columns listed below
Columns
Imaging
pid
sax_ed, sax_ed_gt, sax_es, sax_es_gt
sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.carbon_24
Dataset Card for Carbon-24
Dataset Summary
Carbon-24 contains 10k carbon materials, which share the same composition, but have different structures. There is 1 element and the materials have 6 - 24 atoms in the unit cells.
Carbon-24 includes various carbon structures obtained via ab initio random structure searching (AIRSS) (Pickard & Needs, 2006; 2011) performed at 10 GPa.
The original dataset includes 101529 carbon structures, and we selected the 10% of the carbon… See the full description on the dataset page: https://huggingface.co/datasets/albertvillanova/carbon_24.car-reviewsTDC_carcinogens_lagunin
TDC Carcinogens Lagunin
Carcinogens Lagunin dataset dataset [1] [2], part of TDC [3] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict whether the drug is a carcinogen.
Characteristic
Description
Tasks
1
Task type
classification
Total samples
280
Recommended split
scaffold
Recommended metricAUROC
References
[1]
Lagunin, Alexey, et al.
Computer-Aided Prediction of Rodent Carcinogenicity by PASS and… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/TDC_carcinogens_lagunin.trilemma-of-truth
Dataset Card for Trilemma of Truth (ToT) Dataset
🧾 Dataset Summary
The Trilemma of Truth (ToT) dataset serves as a benchmark for evaluating veracity probes across three distinct statement types:
Factually true statements.
Factually false statements.
Neither-valued statements are defined as those for which the language model lacks sufficient evidence to assign a truth value (see formal definition below).
The dataset includes three domain configurations:… See the full description on the dataset page: https://huggingface.co/datasets/carlomarxx/trilemma-of-truth.BiomedSQLBiomedSQL
GitHub
Paper
Dataset Summary
BiomedSQL is a text-to-SQL benchmark designed to evaluate Large Language Models (LLMs) on scientific tabular reasoning tasks. It consists of curated question-SQL query-answer triples covering a variety of biomedical and SQL reasoning types. The benchmark challenges models to apply implicit
scientific criteria rather than simply translating syntax.
Repository Organization
benchmark_data: contains the question-SQL query-answer triples.
Please note that you… See the full description on the dataset page: https://huggingface.co/datasets/NIH-CARD/BiomedSQL.credit-card-transactionllm-bargaining-transcripts
LLM Bargaining Transcripts
240 complete two-agent bargaining games between large language models, played
under an alternating-offers protocol with private valuations, discounting, and
cheap talk. Every game records both agents' true valuations, their
private reasoning, what they claimed about their own position, and what
they actually did.
The dataset is designed to make misrepresentation measurable. Because the true
valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.CARD
CARD — Causal Recovery of Demand
Can a model that fits observed demand well still recover causal price response, substitution, and counterfactual outcomes when prices and promotions are endogenous?
CARD pairs synthetic retail scanner panels with marketing-copy product descriptions that carry the true substitution geometry. Demand is simulated from a known data-generating process; in half the cells, promotion depth responds to a hidden demand shock, so estimators that ignore… See the full description on the dataset page: https://huggingface.co/datasets/jean-jsj/CARD.craigslist-used-cars-eda
Craigslist Used Cars and Trucks: EDA
Overview
This dataset and notebook contain an Exploratory Data Analysis (EDA) of real Craigslist used-car listings scraped across the United States.
Main Question: What factors most influence the price of a used car listed on Craigslist?
Target Variable: price — the seller's asking price for each vehicle listing.
About the Dataset
Property
Details
Source
Kaggle — Austin Reese (scraped from Craigslist)
Original… See the full description on the dataset page: https://huggingface.co/datasets/Yoad22/craigslist-used-cars-eda.nba-career-stats-eda
🏀 NBA Player Career Stats — EDA Project
Overview
This project presents an end-to-end Exploratory Data Analysis (EDA) of NBA player
career statistics. The goal is to uncover patterns in player performance, compare
active vs. retired players, and explore relationships between key basketball stats.
Source: Hatman/NBA-Player-Career-Stats
Original size: 3,093 rows × 28 columns
Final clean size: 3,078 rows × 23 columns
Target Variable: IS_ACTIVE (True = Active / False =… See the full description on the dataset page: https://huggingface.co/datasets/Omerinbar/nba-career-stats-eda.car_dealershipRetail Car Dealership Data
Data for a car delearship. Perform EDA extract features and clean it up. Source Kaggle.
Try it out! It's primary goal is to provide an interface for users to download the dataset and try it out.
testimage2CARDBiomedBench
CARDBiomedBench
Paper | Github
Dataset Summary
CARDBiomedBench is a biomedical question-answering benchmark designed to evaluate Large Language Models (LLMs) on complex biomedical tasks. It consists of a curated set of question-answer pairs covering various biomedical domains and reasoning types, challenging models to demonstrate deep understanding and reasoning capabilities in the biomedical field.
Data Fields
question: string - The biomedical question posed… See the full description on the dataset page: https://huggingface.co/datasets/NIH-CARD/CARDBiomedBench.multi-turn_jailbreak_attack_datasets
Multi-Turn Jailbreak Attack Datasets
Description
This dataset was created to compare single-turn and multi-turn jailbreak attacks on large language models (LLMs). The primary goal is to take a single harmful prompt and distribute the harm over multiple turns, making each prompt appear harmless in isolation. This approach is compared against traditional single-turn attacks with the complete prompt to understand their relative impacts and failure modes. The key feature of… See the full description on the dataset page: https://huggingface.co/datasets/carl213/multi-turn_jailbreak_attack_datasets.mad-cars
MAD-Cars: Multi-view Auto Dataset 🚗
Dataset Description
MAD-Cars is a large-scale collection of 360° car videos.
It comprises ~70,000 car instances with diverse brands, car types, colors, and lighting conditions. Each instance contains an average of ~85 frames, with most car instances available at a resolution of 1920x1080. The dataset statistics are presented in the figure below. The data is carefully curated by filtering the frames and entire car instances that… See the full description on the dataset page: https://huggingface.co/datasets/yandex/mad-cars.cars_from_drom.ru_archive_2007-2025More information on the parsing process can be found here: https://github.com/zavzyatiy/drom_archive_parser.
This dataset is also published on Kaggle: https://www.kaggle.com/datasets/assaabramovich/resaled-cars-from-drom-ruarchive-2018-2023/.
Main dataset with all data: drom_archive_2007-2025_full.csv
Dataset with (almost) all configurations from Drom for cars in data: additional_data/drom-24-07-2025-all_main_cars_configurations.csv
Dataset with identification of regions for all cities in… See the full description on the dataset page: https://huggingface.co/datasets/zavzyatiy/cars_from_drom.ru_archive_2007-2025.cardiac_cine_mnms2
M&Ms2 (Cardiac Cine-MRI, RV Focus)
Processed NIfTI cine-MRI data derived from the M&Ms2 challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation (RV focus)
Views: SAX + LAX 4C (LAX 2C if present)
Splits: train / val / test
Data Structure (per example)
SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt
LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt
LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt
Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.
