datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vintage-ft-v1
Vintage fine-tuning
This is a fully synthetic dataset generated using TypeWriter-7B, Talkie-13B, MonadGPT and other LLMs.
The language is English.
It is designed to be time-limited to the year 1900.
There is no knowledge about airplanes, atomic bombs, antibiotics or computers. See a big list of banned terms in the banned.txt file.
Some modern words may have leaked in the data from modern LLMs like Claude, Gemma, etc., even if the processing pipeline is aggressively dropping any… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/vintage-ft-v1.tiny-vintage-completions
Tiny vintage completions
Synthetic vintage texts, with a cutoff date for year 1900.
Based on unique 2-3 word seeds, extracted from croqaz/Vintage-v1, croqaz/Vintage-v2 and Haykgrigorian/English-historical-corpus-1800-1875.
Check the files seeds1.txt and seeds2.txt.
Generated by TypeWriter-7B-base and Talkie-13B-base completions.
Citation
If you find this dataset valuable, please consider citing:
@misc{Tiny-vintage-completions,
title = {Tiny vintage completions}… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/tiny-vintage-completions.vintage-gsm8k
Vintage GSM8K
A full-size adaptation of OpenAI GSM8K for models whose knowledge ends on December 31, 1930. Post-cutoff context was minimally rewritten while preserving complete written reasoning, calculations, and final answers.
Train rows: 7,473
Test rows: 1,319
Contextually rewritten rows: 1,007
Schema: id, question, answer
Source: OpenAI GSM8K, main
Source revision: 740312add88f781978c0658806c59bc2815b9866
License: MIT
The official train/test sizes, source order, and stable… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/vintage-gsm8k.vintage-gsm8k-filtered
Vintage GSM8K (Filtered)
A filtered derivative of OpenAI GSM8K for models whose knowledge ends on December 31, 1930. Rows requiring contextual rewriting were removed rather than modified.
Train rows: 6,600
Test rows: 1,185
Schema: id, question, answer
Source: OpenAI GSM8K, main
Source revision: 740312add88f781978c0658806c59bc2815b9866
License: MIT
Filtering changes the official split sizes. Source order and stable IDs are preserved among retained rows.
Solutions retain their… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/vintage-gsm8k-filtered.vintage-ft-v2
Vintage fine-tuning-2
Contains:
080-ai/mcq_ps_v1
11-47/Archangel_grabrail_25k_mindstate_dataset
11-47/Archangel_michael_mindstate_25k
11-47/archimedes_mindset_25k
11-47/high_priest_occult_50k -- only 10 unique entries
11-47/leonardo_da_vinci_mindframe_instruction_dataset -- 2,002 entries
11-47/newton_mindset_training_dataset -- 5,735 entries
ambrosfitz/Openstax_american_yawp
ambrosfitz/OR_training_full
ambrosfitz/Philosophy
croqaz/commonsense-v1
croqaz/vintage-exam-qa
grade… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/vintage-ft-v2.vintage-photography-captions
Dataset Card for Vintage Photograph Captions Recaption
This dataset contains 445,271 recaptioned vintage photographs, derived from the vintage-photography-450k-high-quality-captions dataset. It provides high-quality bilingual (English and Chinese) captions, aesthetic scores, and other metadata generated using the Qwen2-VL model.
This dataset is a recaptioned version of SilentAntagonist/vintage-photography-450k-high-quality-captions. The original dataset contained 456,006… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/vintage-photography-captions.VintageCookingRecipes
🍽️ Vintage American Recipes Dataset (1940–1999)
This curated dataset features a selection of vintage American recipes extracted from unpublished church cookbooks, bulletins, and local magazines spanning 1940 to 1999. Each recipe is cleaned and structured in JSON format, including fields such as title, ingredients, instructions, category, and year.
Over the years, I’ve been collecting a large archive of 10,000+ pre-internet recipes from cookbooks, newspapers, churches, schools, and… See the full description on the dataset page: https://huggingface.co/datasets/dataDatasic/VintageCookingRecipes.
