foundations
ai-ml-foundations-book-collection
Introduction
I put this collection together after spending a lot of time reading what I think are some of the best books on AI, machine learning, deep learning, probabilistic modeling, optimization, reinforcement learning, transformers, LLMs, validation, and fairness. I want to share this with the community for one simple reason: I want to give people a structured path through the books that actually help them understand things deeply, instead of sending them through random… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/ai-ml-foundations-book-collection.Lumina-Math-Foundations-1B
Lumina-Math-Foundations-1B
Lumina-Math-Foundations-1B is an industrial-scale foundational mathematical reasoning dataset in Indonesian,
Dataset Summary
Language: Indonesian (id) with LaTeX mathematical formulas.
Scale: 1 Billion Synthetic High-Fidelity Mathematical Reasoning instances.---
Data Schema & Field Breakdown
Field Name
Type
Description
problem
string
100% pure human natural language problem statement with standard LaTeX math… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/Lumina-Math-Foundations-1B.measurement-db
The AI Measurement Data Bank
Measurement Data Bank is a curated collection of standardized, item-level AI
evaluation results for measurement-science analysis.
Schema and versioned downloads
These six benchmark datasets use schema version 3. The GitHub repository contains the corresponding builders, source manifests, tests, and curation records.
Grading criteria are stored in items.parquet; traces link to observations through response_id. The response table on… See the full description on the dataset page: https://huggingface.co/datasets/aims-foundations/measurement-db.safety-irt
⚠️ Content Warning: This dataset contains sensitive prompts and model
responses that include harmful, offensive, and dangerous content. It is intended
for safety research.
Dataset Card for Safety-IRT
Data for "Why Do Safety Guardrails Degrade Across Languages?"
Contains 1.9M graded responses from 61 model configurations across 10 languages,
along with anchor selections, judge validation data, and native speaker translation ratings.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/aims-foundations/safety-irt.ai-ml-foundations-book-collection
Introduction
I put this collection together after spending a lot of time reading what I think are some of the best books on AI, machine learning, deep learning, probabilistic modeling, optimization, reinforcement learning, transformers, LLMs, validation, and fairness. I want to share this with the community for one simple reason: I want to give people a structured path through the books that actually help them understand things deeply, instead of sending them through random courses… See the full description on the dataset page: https://huggingface.co/datasets/invincible-jha/ai-ml-foundations-book-collection.cognitive_foundations
