datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
accessibility-atlas
Accessibility Atlas
42 datasets covering disability demographics, employment gaps, web accessibility compliance, assistive technology patents, special education, healthcare, housing discrimination, transportation, government benefits, and more -- across 15 categories.
Built for researchers, journalists, policy analysts, and anyone building tools for disabled communities.
What's Inside
US Disability Demographics (Census Bureau)
8 datasets from the American… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/accessibility-atlas.worldbank-project-documents
Dataset Card for World Bank Project Documents
Dataset Summary
This is a dataset of documents related to World Bank development projects in the period 1947-2020. The dataset includes
the documents used to propose or describe projects when they are launched, and those in the review. The documents are indexed
by the World Bank project ID, which can be used to obtain features from multiple publicly available tabular datasets.
Supported Tasks and Leaderboards
No… See the full description on the dataset page: https://huggingface.co/datasets/lukesjordan/worldbank-project-documents.us-military-veteran-analysis
US Military & Veteran Analysis by State
50 states: veterans, firearms, PTSD, suicide rates, VA healthcare
State-level integration of veteran demographics, firearm ownership, mental health indicators, and VA healthcare utilization from Census ACS, RAND, ATF, CDC, VA, and DoD.
Dataset Structure
See demo_notebook.ipynb for data exploration examples.
Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/us-military-veteran-analysis.bluesky-alt-text
Bluesky Alt Text: Frozen April 2026 Snapshot
Frozen historical snapshot. These files were collected in April 2026 and
will not be extended into a longitudinal series. The 279K-row corpus was
deliberately selected from accounts with high alt-text adoption, so it must
not be interpreted as a representative platform adoption estimate. The
14.5-hour Jetstream file is a short observed window, not a durable census.
Rows contain author handles and DIDs, post URIs and CIDs, post… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bluesky-alt-text.FirstProofGradingBench
FirstProofGradingBench
16 research-level problem–proof pairs with 24 expert-derived error annotations,
for evaluating automatic proof verifiers, derived from the second batch of the
First Proof project.
Each row is an AI-generated proof attempt at a research mathematics problem that
no frontier model could solve at the time of release, together with the errors
expert mathematicians found when refereeing it. A verifier is given the problem
statement and the proof, and must return… See the full description on the dataset page: https://huggingface.co/datasets/LukeBailey181Pub/FirstProofGradingBench.us-housing-affordability-crisis
US Housing Affordability Crisis by County
3,222 US counties: rent burden, income ratios, minimum wage hours needed
3,222 US counties with rent burden analysis, affordability calculations, and minimum wage comparisons derived from Census ACS 2022 data.
Dataset Structure
See demo_notebook.ipynb for data exploration examples.
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("lukeslp/us-housing-affordability-crisis")
# Or load… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/us-housing-affordability-crisis.corporate-governance-reasoning
Dataset Summary
The corporate-governance-reasoning dataset was designed to test a model's ability to reason about executive/board/shareholder proposals to alter companies' corporate governance structures. While there are multiple legal datasets, none are focused specifically on reasoning tasks. On the other hand, reasoning (ratiocination and ability to make connections to precedent) is a core part of the practice of the law in the real world. We focus on a specific aspect of the… See the full description on the dataset page: https://huggingface.co/datasets/LukeIrwin/corporate-governance-reasoning.math-reasoning-questionsgithub-issuesgithub-issues-embeddingsqwen3-80b-regen-prefectblendemergent-misalignment-questionstone_agnostic_questionstestossbaby-agi-dataset-v0
BabyAGI (Dataset)
The initial demonstration dataset follows the Huggingface dataset spec, with the raw data split into two components, trajectory images and trajectory metadata. The metadata is stored in the raw dataset, and the images are stored on S3. The data is loaded using the dataloader defined in baby_agi_dataset.py.
Data Layout:
├── data
│ ├── metadata_0.json
│ ├── metadata_1.json
│ └── ...
├-- baby_agi_dataset.py
Metadata Format (.json)
[
{… See the full description on the dataset page: https://huggingface.co/datasets/lukemann/baby-agi-dataset-v0.test_cotqwen3-235-regen-perfect_blenddeepfabric-github-mcp
deepfabric-github-mcp
Dataset generated with DeepFabric.
point_flock-miner-1charon-corpusllama-lora-testRH-TESTRH-AITest
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/LukeWais/RH-AITest.llama-aa-dataset
