datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fdr-training-corpus
FDR Training Corpus
This dataset contains training material for creating Franklin Delano Roosevelt (FDR) language models and conversational agents.
Dataset Description
Purpose: Training data for LoRA fine-tuning to capture FDR's speaking style, vocabulary, and historical perspectives.
Content: Speeches, letters, fireside chats, press conferences, and other public communications from FDR's presidency (1933-1945).
License: CC0-1.0 (Public Domain) - All content is from… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/fdr-training-corpus.HardMultiQAFDR_Corpus
FDR Agent Corpus
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of the text
subject_name:… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/FDR_Corpus.fdr-training-corpus-text-format
