datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-ml-foundations-book-collection
Introduction
I put this collection together after spending a lot of time reading what I think are some of the best books on AI, machine learning, deep learning, probabilistic modeling, optimization, reinforcement learning, transformers, LLMs, validation, and fairness. I want to share this with the community for one simple reason: I want to give people a structured path through the books that actually help them understand things deeply, instead of sending them through random… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/ai-ml-foundations-book-collection.Lumina-Math-Foundations-1B
Lumina-Math-Foundations-1B
Lumina-Math-Foundations-1B is an industrial-scale foundational mathematical reasoning dataset in Indonesian,
Dataset Summary
Language: Indonesian (id) with LaTeX mathematical formulas.
Scale: 1 Billion Synthetic High-Fidelity Mathematical Reasoning instances.---
Data Schema & Field Breakdown
Field Name
Type
Description
problem
string
100% pure human natural language problem statement with standard LaTeX math… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/Lumina-Math-Foundations-1B.measurement-db
The AI Measurement Data Bank
Measurement Data Bank is a curated collection of standardized, item-level AI
evaluation results for measurement-science analysis.
Schema and versioned downloads
These six benchmark datasets use schema version 3. The GitHub repository contains the corresponding builders, source manifests, tests, and curation records.
Grading criteria are stored in items.parquet; traces link to observations through response_id. The response table on… See the full description on the dataset page: https://huggingface.co/datasets/aims-foundations/measurement-db.safety-irt
⚠️ Content Warning: This dataset contains sensitive prompts and model
responses that include harmful, offensive, and dangerous content. It is intended
for safety research.
Dataset Card for Safety-IRT
Data for "Why Do Safety Guardrails Degrade Across Languages?"
Contains 1.9M graded responses from 61 model configurations across 10 languages,
along with anchor selections, judge validation data, and native speaker translation ratings.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/aims-foundations/safety-irt.ai-ml-foundations-book-collection
Introduction
I put this collection together after spending a lot of time reading what I think are some of the best books on AI, machine learning, deep learning, probabilistic modeling, optimization, reinforcement learning, transformers, LLMs, validation, and fairness. I want to share this with the community for one simple reason: I want to give people a structured path through the books that actually help them understand things deeply, instead of sending them through random courses… See the full description on the dataset page: https://huggingface.co/datasets/invincible-jha/ai-ml-foundations-book-collection.cognitive_foundationsai-ml-foundations-book-collection
Introduction
I put this collection together after spending a lot of time reading what I think are some of the best books on AI, machine learning, deep learning, probabilistic modeling, optimization, reinforcement learning, transformers, LLMs, validation, and fairness. I want to share this with the community for one simple reason: I want to give people a structured path through the books that actually help them understand things deeply, instead of sending them through random courses… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/ai-ml-foundations-book-collection.moral_stories_foundations
Moral Stories, labelled by moral foundation
Why moral foundations? Moral Foundations Theory is one of the few maps of human values that is
calibrated against real people: it was drawn from cross-cultural survey studies and aims to hold
across societies, not just Western ones. That breadth is what matters here. If we want to study
how a constructed intelligence, a kind of moral alien, reasons about right and wrong, we need a
measure that generalises beyond any one culture. Moral… See the full description on the dataset page: https://huggingface.co/datasets/wassname/moral_stories_foundations.ai-ml-foundations-book-collection
Introduction
I put this collection together after spending a lot of time reading what I think are some of the best books on AI, machine learning, deep learning, probabilistic modeling, optimization, reinforcement learning, transformers, LLMs, validation, and fairness. I want to share this with the community for one simple reason: I want to give people a structured path through the books that actually help them understand things deeply, instead of sending them through random… See the full description on the dataset page: https://huggingface.co/datasets/Dinamitrii/ai-ml-foundations-book-collection.FoundationStereoWeightsAI-Governance-Foundations
AI Governance Foundations
A practical educational series on AI governance, risk, compliance, security, accountability, and the controls required to govern increasingly capable AI systems.
The series starts with foundational concepts and progressively moves into deeper technical and governance topics.
Modules
Module 01 — What Is AI Governance?
Introduces AI governance from the ground up: authority, accountability, risk control, and evidence.… See the full description on the dataset page: https://huggingface.co/datasets/jeremy0621rose/AI-Governance-Foundations.ai-ml-foundations-book-collection
Introduction
I put this collection together after spending a lot of time reading what I think are some of the best books on AI, machine learning, deep learning, probabilistic modeling, optimization, reinforcement learning, transformers, LLMs, validation, and fairness. I want to share this with the community for one simple reason: I want to give people a structured path through the books that actually help them understand things deeply, instead of sending them through random… See the full description on the dataset page: https://huggingface.co/datasets/prashant-AI-ML/ai-ml-foundations-book-collection.RC_Foundations_Dataset_V1
PN Engineering Datasets
DATA DICTIONARY – Professional Specification
This document defines the structure and attributes of all datasets released
under PN Engineering Datasets.
Dataset Metadata
dataset_name: RC_Foundations_Dataset_V1
dataset_version: V1
element_type: Reinforced concrete foundations
pdf_count: 25
png_count: 25
File Structure
PDF files:
• Flattened
• Anonymous
• Metadata removed
• OCR-ready
PNG files:
• 1200 DPI resolution
• Clean and uniform background
• High… See the full description on the dataset page: https://huggingface.co/datasets/PNEngineeringDatasets/RC_Foundations_Dataset_V1.ml-foundations-high-reasoning-480kml-foundations-vtuber-wikiintrinsic-intelligence-foundations
🌿 Intrinsic Intelligence Foundations
Toward truly autonomous and benevolent intelligence — beyond externally imposed objectives.
Intrinsic Intelligence Foundations is a structured, math-aware JSONL corpus built from K. Takahashi’s theoretical preprints (Fractal Category Theory / PF–UGV / “no-meta” autonomy line).It is designed to help LLMs understand mathematical structure, category-theoretic formalisms, and equation-level reasoning, while exposing an explicit architecture… See the full description on the dataset page: https://huggingface.co/datasets/kadubon/intrinsic-intelligence-foundations.measurement-db-embedcognitive_foundations_human_v4Stratmeyer-Analytica-Foundations
Stratmeyer Analytica: Foundational Frameworks
Author: Adam Ian Stratmeyer, J.D.Institution: Stratmeyer Analytica
Overview
This repository hosts the foundational documents of the Observable Function framework and the Helpful-Harmless Paradox analysis. These documents serve as the primary doctrinal texts for understanding machine agency, institutional constraints, and the structural contradictions in modern AI deployment.
Documents
1. Observable Function… See the full description on the dataset page: https://huggingface.co/datasets/aistrat/Stratmeyer-Analytica-Foundations.The-Foundations-of-I-AMfast_foundationstereoopenart-painterly-foundations
OpenArt — Painterly Foundations
openart-painterly-foundations is the painterly foundations subject collection of the OpenArt family
of open, public-domain art datasets: 13,786 works (9,107 paintings/illustrations · 4,596
photographed objects · 83 unclassified), each paired with a structured VLM caption plus
medium, attribution and inscription metadata.
The painting-forward core of OpenArt — the 2-D fine-art techniques that anchor the openbrush brand: drawing, etching, oil… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-painterly-foundations.ml-foundations-vtubercommentsreddita
ml-foundations-vrchatlegendsml-foundations-vtuberredditpostsecosystem
AI Evaluation Ecosystem Simulation Dataset
Hugging Face dataset repository: aims-foundations/ecosystem.
Simulation outputs supporting the AI Evaluation Ecosystem paper. Each run is a stochastic
simulation of an AI evaluation ecosystem (providers, evaluators, consumers, regulators,
funders, media) over 40 monthly rounds. This release contains 119 LLM-mode runs (agent policies: claude-opus-4-6, claude-sonnet-4-6, gpt-5.5-2026-04-23) and 250 heuristic-mode runs (rule-based agent… See the full description on the dataset page: https://huggingface.co/datasets/aims-foundations/ecosystem.ml-foundations-vtuber-wiki-longml-foundations-vrctermsworldsfoundation-sec-v2-datasetmy_dataset_repo
