datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gallica_literary_fictions
Dataset Card for Literary fictions of Gallica
Dataset Summary
The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.bangla-literary-corpus
Bangla Literary Corpus - বাংলা সাহিত্য সংকলন
Bangla Literary Corpus is a large-scale, chapter-segmented Bengali literature dataset with over 3,200 books including classical novels (উপন্যাস), short stories (ছোটগল্প), essays (প্রবন্ধ), and poetry collections (কবিতা). Curated specifically for Bengali NLP and Indic language modeling, this corpus provides high-register, linguistically diverse Bangla text optimized for large language model (LLM) pretraining, continual pre-training… See the full description on the dataset page: https://huggingface.co/datasets/0xmezbah/bangla-literary-corpus.
