datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
london-llm-1800
Historic London LLM Dataset (1800-1875)
This dataset consists of approximately 90GB of cleaned English text derived from historical documents ranging from 1800 to 1875. It serves as the training corpus for the "Time Capsule LLM" project.
Dataset Details
Time Period: 1800 - 1875
Language: English
Source: Digitized historical texts (Internet Archive)
Tokenization: Custom BPE Tokenizer (vocab size 32,000)
Usage
This dataset is intended for research in digital… See the full description on the dataset page: https://huggingface.co/datasets/postgrammar/london-llm-1800.london_venues_synthetic
London Venues Synthetic Dataset 🇬🇧
Project Overview
This dataset contains 10,000 synthetic rows of fictional venues in London, designed to train and test a Semantic Search & Recommendation System.
The goal of this project was to solve the "problem" in recommendation engines. Real-world user reviews are often messy, sparse, or lack specific "intent" or "vibe" contexts (e.g., explicitly mentioning "good for studying" or "cosy cafe"). By generating synthetic data, we… See the full description on the dataset page: https://huggingface.co/datasets/uleeberber/london_venues_synthetic.
