datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron_actual_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
OpenHermes-NoRefusal-95K
OpenHermes-NoRefusal-95K
A refusal-free instruction-tuning dataset: 95,401 single-turn conversations derived from
teknium/OpenHermes-2.5, filtered so
that zero assistant responses contain refusals, hedging boilerplate, or
"as an AI language model" disclaimers.
Why this exists
The usual way to get a model that doesn't refuse is to train it on aligned data and then
remove the alignment afterwards — refusal-direction ablation, weight editing, abliteration.
That works… See the full description on the dataset page: https://huggingface.co/datasets/ghost-actual/OpenHermes-NoRefusal-95K.ipt_actual_all_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
actuarial-fm-p-ifm-dataset
Actuarial FM + P + IFM Dataset v0.0.7
Dataset Description
Comprehensive training dataset for actuarial AI covering three SOA exams.
Dataset Summary
Total Examples: 18,794
Exam FM: ~18,000 examples
Exam P: 743 examples
Exam IFM: 37 examples
Format: JSONL with instruction-response pairs
Topics Covered
Financial Mathematics (FM)
Time value of money
Annuities and perpetuities
Bonds and interest theory
Amortization
Probability (P)… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-fm-p-ifm-dataset.actuarial-conversational-dataset
Conversational Actuarial Dataset v0.1.0
Revolutionary Approach: Human First, Expert Second
This dataset transforms technical AI into conversational AI while maintaining domain expertise.
Dataset Composition
Total Examples: 461
56.6% Conversational: Natural dialogue, emotions, context
43.4% Technical: Actuarial with personality
Conversational Categories
Basic Interactions (56 examples)
Greetings and introductions
Small talk
Humor and… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-conversational-dataset.actuarial-gpt-conversations
👋 Connect with me on LinkedIn!
Manuel Caccone - Actuarial Data Scientist & Open Source Educator
Let's discuss actuarial science, AI, and open source projects!
📊 ActuarialGPT Conversations Dataset
Precision Mathematical Conversations for Insurance Intelligence
🎯 Quick Facts
Feature
Description
Domain
Actuarial Science, Insurance Analytics, Risk Management
Language
English (Technical/Expert Level)… See the full description on the dataset page: https://huggingface.co/datasets/manuelcaccone/actuarial-gpt-conversations.actuarial-fm-p-ifm-ultimate-dataset
Ultimate Actuarial FM/P/IFM Dataset v0.0.9
Dataset Description
The ultimate training dataset for actuarial AI models, containing 1,708 meticulously crafted examples targeting 95%+ accuracy on professional actuarial exams.
Dataset Statistics
Total Examples: 1,708
Train: 1,366 (80%)
Validation: 170 (10%)
Test: 172 (10%)
Distribution by Exam
Exam
Examples
Percentage
IFM
884
51.8%
P
570
33.4%
FM
254
14.9%
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-fm-p-ifm-ultimate-dataset.actuarial-exam-fm-p-dataset
Actuarial Exam FM & P Dataset v0.0.6
Dataset Description
Comprehensive training dataset for actuarial AI models covering SOA Exam FM (Financial Mathematics) and Exam P (Probability).
Dataset Summary
Total Examples: 18,757
Exam FM Examples: ~18,000
Exam P Examples: 743
Format: JSONL with instruction-response pairs
Example Structure
{
"instruction": "Calculate the present value of a 10-year annuity...",
"response": "To solve this problem, I'll… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-exam-fm-p-dataset.fluid-knowledge-validation
Fluid Knowledge
Public synthetic release-validation fixtures. These repeated arithmetic items test artifact publication and verification only; they were not authored or blindly reviewed by frontier models and are not a usable benchmark.
Each immutable epochs/<id>/manifest.json binds its published artifacts. commitment.json reveals the nonce for verification. Protocol and source attribution accompany each epoch. Pin the returned Hugging Face commit SHA for reproduction. Public… See the full description on the dataset page: https://huggingface.co/datasets/actuallymentor/fluid-knowledge-validation.ActuarialMathBenchActuarialMathBench is a domain-specific benchmark dataset designed to evaluate the mathematical reasoning capabilities of Large Language Models (LLMs) within the actuarial domain. The dataset consists of 750 question-answer pairs derived from Society of Actuaries (SOA) sample exams.
Dataset Statistics
Each exam subject is part of the curriculum to obtain the designation Associate of the Society of Actuaries (ASA). The % Pass Mark is the 90th percentile pass mark percentage over the… See the full description on the dataset page: https://huggingface.co/datasets/BjornvanBraak/ActuarialMathBench.actuarial-fm-p-ifm-enhanced-dataseteval_actuator_unboxing_pi05_sweep_v2_01_freeze_01_fullbns_act_2023
Dataset on BNS Act, 2023 which replaces the Indian Penal Code
This dataset has been made from the official document published by the government of India on the BNS Act, 2023
Source: PDF Source
