vintage
Datasets
All datasets matching “vintage”vintage-v1
Vintage dataset
Composed of documents from:
biglam/hmd_newspapers
biglam/lampeter_corpus
TheBritishLibrary/blbooks
dell-research-harvard/AmericanStories
dell-research-harvard/NewsWire
storytracer/LoC-PD-Books
some books from https://ota.ox.ac.uk/
The perfect score is larger than 0 (zero). The score is composed of how far the metrics are from 100,
either smaller of larger than 100, except for quality which can be larger than 0.
Short documents
Filtered documents… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/vintage-v1.vintage-v2
Vintage dataset
Contains:
CLMET v3.1 -- 335 books
ECCO -- 3,102 books
EEBO -- 34,937 books
EVANS -- 5,013 books
All books are cleaned (best effort) from XML to Markdown.
Original sources:
https://fedora.clarin-d.uni-saarland.de/clmet/clmet.html
https://textpartnership.net/pages/faq.html
https://dropbox.com/sh/inyy253jvytoxcu/AAAT9nMPo0sa5aXEgs5SC5aGa?dl=0 -- Evans TCP bulk files
https://app.box.com/s/zj7pzfokxde4glrhebxsavbbxzyr3ogz -- Evans TCP bulk files… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/vintage-v2.vintage-adsvintage-exam-qa
Vintage exams Q&A
This is a fine-tuning dataset, extracted from vintage books using code (not with AI).
These 6 books from Archive.org are included:
Advanced question book 1883
Common school examiner and review 1890
New common school question book 1888
New common school question book 1900
Recreations in the common school studies 1885
School Room Search Light 1895
Four books have separate sections for the questions and answers, so they had to be matched (you don't need to do… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/vintage-exam-qa.vintage-conversations
Vintage Conversations
These conversations are extracted from Gutenberg books using code and LLM extraction.
They are guaranteed to be identical to the original text (every dialog text is matched against the original book letter by letter), but it's not guaranteed that all dialogs are extracted, or that the speaker is correct.
The nearby words of the same speaker are grouped together, eg:
“He has a thirst for travelling; perhaps he may turn out a Bruce or a
Mungo Park,” said Mr.… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/vintage-conversations.winedb-fine-wines-and-vintages
🍷 WineDB — Fine Wine & Vintages Sample Dataset
Full dataset: winedb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A curated free sample of the WineDB dataset: highly normalized relational tables tracking fine wine producers, cuvees, exact vintage varietal blend percentages (SUM <= 100.001, enforced by SQLite triggers), alcohol content (ABV %), organoleptic tasting descriptors, and secondary market valuation indices. Prefer SQLite? The same… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/winedb-fine-wines-and-vintages.
