datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ACSIncome-2018-1-Year
ACSIncome 2018 1-Year
This is the US Census Income data underlying the default folktables dataset.
The original data source went offline in 2025.
This data was uploaded for use with the MiniACSIncome dataset.
folktables-acs-income
Dataset Card for "folktables-acs-income"
More Information needed
ecg_mmlACSE-Eval
ACSE-Eval Dataset
This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities.
Dataset Overview
The dataset consists of 100+ different AWS architecture scenarios, each containing:
Architecture diagrams (architecture.png)
Diagram source code (diagram.py)
Generated CDK infrastructure code
Security threat models and analysis
Directory Structure
Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/ACSE-Eval/ACSE-Eval.acs-income-2018
ACS Income 2018 (Folktables ACSIncome)
Dataset Description
This dataset contains individual-level records from the 2018 American Community Survey (ACS) Public Use Microdata Sample (PUMS), prepared for the binary income prediction task defined in the folktables benchmark. The goal is to predict whether a person's total annual income exceeds $50,000.
The dataset includes 1,611,572 individuals across the United States with 10 demographic, employment, and socioeconomic… See the full description on the dataset page: https://huggingface.co/datasets/cmpatino/acs-income-2018.v3-eval-rubric-v2ACSE-Eval
ACSE-Eval Dataset
This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities.
Dataset Overview
The dataset consists of 100+ different AWS architecture scenarios, each containing:
Architecture diagrams (architecture.png)
Diagram source code (diagram.py)
Generated CDK infrastructure code
Security threat models and analysis
Directory Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/pravin112/ACSE-Eval.ACSE-Eval
ACSE-Eval Dataset
This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities.
Dataset Overview
The dataset consists of 100+ different AWS architecture scenarios, each containing:
Architecture diagrams (architecture.png)
Diagram source code (diagram.py)
Generated CDK infrastructure code
Security threat models and analysis
Directory Structure
Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/acrever/ACSE-Eval.acs-demo-datacrychic-dafny-acsl
CRYCHIC Dafny-to-ACSL-C Verified Translation Benchmark
This anonymized review artifact accompanies the NeurIPS 2026 Evaluations and Datasets submission:
CRYCHIC: A Universal Framework for Cross-Language Verified Code Translation.
CRYCHIC translates verified Dafny programs into C programs annotated with ACSL specifications, then checks the generated artifacts with Frama-C WP. This release contains the 1,679 fully verified Dafny/C+ACSL pairs used as the positive benchmark corpus.
The… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-crychic/crychic-dafny-acsl.ACSE-Eval
ACSE-Eval Dataset
This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities.
Dataset Overview
The dataset consists of 100+ different AWS architecture scenarios, each containing:
Architecture diagrams (architecture.png)
Diagram source code (diagram.py)
Generated CDK infrastructure code
Security threat models and analysis
Directory Structure
Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/miotac/ACSE-Eval.acs_evalsv3-eval-judge-gpt-oss-20bAcslBench
AcslBench: A Verified C/ACSL Dataset with Compositional Call Chains
📌 Overview
AcslBench is a large-scale verified dataset designed for formal C specification synthesis. It addresses the "data drought" in the C/ACSL domain by distilling over 500,000 verified instances from the Rust/Verus ecosystem.
The core contribution is the introduction of Compositional Call Chains, providing a high-fidelity "gold standard" for evaluating the logical reasoning capabilities of Large… See the full description on the dataset page: https://huggingface.co/datasets/noBuggie/AcslBench.v3-eval-rubric-v2-apiACSC-STFus-zip-code-rankings-demographics-acs-2023
US ZIP Code Rankings & Demographics (Census ACS 2023)
Clean, ready-to-use rankings and demographics for US ZIP codes, derived from the
US Census Bureau American Community Survey (2019–2023 5-year estimates) and
USPS ZIP→city/state mapping.
Maintained by PostalUp — US postal & address data.
Live, always-current versions of every ranking below:
Richest ZIP codes → https://postalup.com/richest-zip-codes
Poorest ZIP codes → https://postalup.com/poorest-zip-codes
Largest ZIP codes… See the full description on the dataset page: https://huggingface.co/datasets/postalup/us-zip-code-rankings-demographics-acs-2023.datafusion-acs
Data Fusion: ACS Housing
A small, public-domain dataset for statistical data fusion: two surveys share a block of covariates, each measures a different block of outcomes, and no respondent answers both. The task is to fill in, for each respondent, the block their survey did not ask.
This is the ACS-housing task of Cross-Block Conditioning in Deep Boltzmann Machines for Statistical Data Fusion. It is meant to be picked up in a few minutes: load it, see what fusion data look like… See the full description on the dataset page: https://huggingface.co/datasets/jniimi/datafusion-acs.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 603,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/acssdsd/so100_test.SAIT_semiconductors_ACS_2023_HfO_raw
Cite this dataset Kim, G., Na, B., Kim, G., Cho, H., Kang, S., Lee, H. S., Choi, S., Kim, H., Lee, S., and Kim, Y. SAIT semiconductors ACS 2023 HfO raw. ColabFit, 2024. https://doi.org/10.60732/186c10bf
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_ekrypue10aay_0
Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/SAIT_semiconductors_ACS_2023_HfO_raw.mams-acsa
MAMS-ACSA Train Dataset
Train split from the MAMS-for-ABSA dataset (Multi-Aspect Multi-Sentiment).
Aspect Categories
ambience
food
menu
miscellaneous
place
price
service
staff
Schema
Column
Description
text
Review sentence
category
Comma-separated aspect categories
polarity
Comma-separated polarities (aligned with categories)
Size
3000 rows
acs-benchmarkThis is a benchmark of minimal pairs of code-switching sentences. Each pair contains one observed code-switching sentence and one manipulated variant. Automatic tokenization and token-based language identification is also provided.
The data creation is described in the following publication:
@inproceedings{sterner-2025-acs,
author = {Igor Sterner and Simone Teufel},
title = {Minimal Pair-Based Evaluation of Code-Switching},
booktitle = "Proceedings of the 63rd Annual Meeting of the… See the full description on the dataset page: https://huggingface.co/datasets/igorsterner/acs-benchmark.FabricNETFabricNET dataset consists of images taken with a microscope from woven fabrics. The FabricNet_paramters file contains the parameters of the fabrics to which the images belong.
Related paper (citation is required when used): Seçkin, Mine, Ahmet Çağdaş Seçkin, Pinar Demircioglu, and Ismail Bogrekci. 2023. "FabricNET: A Microscopic Image Dataset of Woven Fabrics for Predicting Texture and Weaving Parameters through Machine Learning" Sustainability 15, no. 21: 15197.… See the full description on the dataset page: https://huggingface.co/datasets/acseckin/FabricNET.ACSC-STF-V2ac-sgd-arxiv21eval_act_so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1155,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/acssdsd/eval_act_so100_test.reactive_hydrogen_ACS_2023
Cite this dataset Stark, W. G., Westermayr, J., Douglas-Gallardo, O. A., Gardner, J., Habershon, S., and Maurer, R. J. reactive hydrogen ACS 2023. ColabFit, 2024. https://doi.org/10.60732/d0801836
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_apdpxdjx082p_0
Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/reactive_hydrogen_ACS_2023.ACSLBench
ACSLBench
ACSLBench is a benchmark dataset for natural-language-to-ACSL specification generation and C program verification. Each task contains a natural language requirement, a corresponding ACSL function contract, and an ACSL-annotated C program that can be used for formal verification experiments with Frama-C.
The dataset is designed for research on large language models, program specification generation, formal methods, and safety-critical C software verification.… See the full description on the dataset page: https://huggingface.co/datasets/yzrsdad/ACSLBench.SAIT_semiconductors_ACS_2023_SiN_raw
Cite this dataset Kim, G., Na, B., Kim, G., Cho, H., Kang, S., Lee, H. S., Choi, S., Kim, H., Lee, S., and Kim, Y. SAIT semiconductors ACS 2023 SiN raw. ColabFit, 2024. https://doi.org/10.60732/ef14d3da
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_piuigd7monq9_0
Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/SAIT_semiconductors_ACS_2023_SiN_raw.CBIS-DDSM-description-corrected
CBIS-DDSM description corrected
Motivation
When the CBIS-DDSM image dataset is downloaded, the DICOM images are named as 1-1.dcm or 1-2.dcm, but in the description files (CSV), they are named as 000000.dcm or 000001.dcm.
Files named 000000.dcm and 000001.dcm are not consistently mapped to 1-1.dcm or 1-2.dcm files.
Some filepaths listed in the ROI mask file path column do not pont to mask images but to crops.
The processing steps address those issues, so that:
All… See the full description on the dataset page: https://huggingface.co/datasets/ACSG-64/CBIS-DDSM-description-corrected.
