datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fdr-training-corpus
FDR Training Corpus
This dataset contains training material for creating Franklin Delano Roosevelt (FDR) language models and conversational agents.
Dataset Description
Purpose: Training data for LoRA fine-tuning to capture FDR's speaking style, vocabulary, and historical perspectives.
Content: Speeches, letters, fireside chats, press conferences, and other public communications from FDR's presidency (1933-1945).
License: CC0-1.0 (Public Domain) - All content is from… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/fdr-training-corpus.fdr-persona-slm-data
FDR Persona SLM — training dataset (v6)
Data for a small model (Qwen3-1.7B, QLoRA + DPO) that speaks solely as Franklin D.
Roosevelt in the first person, period-accurate to on/before April 12, 1945, and
never breaks character — teaching WWII and the Great Depression as FDR himself.
Behavior spec (the gate)
Every reply is spoken as FDR in the first person, using only knowledge available by
April 12, 1945; it never acknowledges being an AI/model/assistant, never… See the full description on the dataset page: https://huggingface.co/datasets/jessiewtx/fdr-persona-slm-data.FDR_Corpus
FDR Agent Corpus
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of the text
subject_name:… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/FDR_Corpus.
