datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.apollo_english_books_translated_to_dutch_with_geminiflash15
Data description
Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
bavarian-books-ocred-v0.1
📚 🥨 OCR'ed Bavarian Books
Due to the lack of high quality resources for Bavarian, I've started this dataset repo for OCR'ing Bavarian books.
The dataset is based on the Bavarian Books dataset.
🧮 OCR
The current form of this dataset uses the awesome Tesseract library for OCR'ing the Bavarian books.
We use the following Fraktur model:
wget https://github.com/tesseract-ocr/tessdata/raw/refs/heads/main/script/Fraktur.traineddata
📄 Dataset Format
Here's an… See the full description on the dataset page: https://huggingface.co/datasets/bavarian-nlp/bavarian-books-ocred-v0.1.booksmatrix-books-all-0000classified-books_all-0000book_sum_sort
