CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ilsp /ancient-modern_greek_translations Dataset Card for Ancient-Modern Greek translations The Ancient-Modern Greek translations dataset includes 100 sentences of Ancient Greek texts manually translated into Modern Greek. Original texts and translations have been extracted from the web sources cited below. Δημοσθένους, Ὑπὲρ τῆς Ῥοδίων ἐλευθερίας (ell: Δημοσθένους, Υπέρ της Ελευθερίας των Ροδίων; eng: Demosthenes, On the Liberty of the Rhodians). Μτφρ. Β.Η. Τσακατίκας. χ.χ. Λόγοι του Δημοσθένη. Γ' Ολυνθιακός, Υπέρ της… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/ancient-modern_greek_translations.texttranslationn<1K2 likes691 downloads2y agoHugging Face02gvlassis /ancient_greek_theatre Ancient Greek Theatre About 🏛️ A collection of 24 Ancient Greek plays, translated in Modern Greek. Description The dataset is basically shakespearefirstfolio, which was itself a better Tiny Shakespeare, but for Ancient Greek theatre plays and Modern Greek. Remarkably, and similarly to Karpathy, we observe that language models trained from scratch on this tiny dataset can produce samples that look eerily close to those written by Ancient Greek playwrights. You… See the full description on the dataset page: https://huggingface.co/datasets/gvlassis/ancient_greek_theatre.texttext-generationn<1K2 likes421 downloads2y agoHugging Face03Ericu950 /AncientGreek Ancient Greek Corpus — re-OCR'd, LLM-repaired A large, deduplicated Ancient Greek corpus, ~361M words in 2,145,799 records. The bulk of it is public-domain page images that we re-OCR'd ourselves with a vision–language model fine-tuned for polytonic Greek (Angleraud et al., 2026), rather than reusing the existing text layers, which roughly quadrupled the usable yield. The noisier output was then repaired by Qwen3.6-27B. Provenance source upstream what we did… See the full description on the dataset page: https://huggingface.co/datasets/Ericu950/AncientGreek.textfill-mask1M<n<10M2 likes276 downloads2mo agoHugging Face04eulogikon /ancient-greek-texts Ancient Greek Texts The surviving literary works of ancient Greece — 1,353 authors and 4,055 works (≈47 million words), spanning Homer through late antiquity. Philosophy, history, drama, lyric, medicine, mathematics, rhetoric, and the fragmentary traditions, all in clean Unicode Greek. This is the data store for Eulogikon: the reading site, search, and browse experience live there. This repository holds the corpus as downloadable files. No logins. No fees. No paywalls.… See the full description on the dataset page: https://huggingface.co/datasets/eulogikon/ancient-greek-texts.text1K<n<10K1 likes156 downloads29d agoHugging Face05Urdatorn /AncientGreek-no-sphragis AncientGreek-no-sphragis (second derivative) A contamination-controlled derivative of Ericu950/AncientGreek at revision 6ac90787c669a7e9218d6d4675a029fa3f10ed99, for pretraining models that are evaluated on Urdatorn/sphragis (revision 1e6d8b58d956e84aec7c1c778bef036bd0286fa9) and Urdatorn/sphragis-metre (revision 43af5a8b230af1c62d2cafb69c0e3d4d83b81800). Both quality tiers are retained. The first derivative removed only source lines that equalled a benchmark unit after… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/AncientGreek-no-sphragis.textfill-mask1M<n<10M0 likes154 downloads1d agoHugging Face06tadad /diorisis-ancient-greek Diorisis Ancient Greek Corpus A Hugging Face conversion of Alessandro Vatri and Barbara McGillivray's Diorisis Ancient Greek Corpus for the BigLAM community. Diorisis contains 820 literary texts from Homer through the fifth century CE, with automatic lemma, part-of-speech, and morphological annotations. The conversion combines the original XML headers with the JSON corpus and a checksum-pinned snapshot of the author's public per-file corrections. It retains Beta Code and adds… See the full description on the dataset page: https://huggingface.co/datasets/tadad/diorisis-ancient-greek.tabulartoken-classification100K<n<1M0 likes74 downloads19d agoHugging Face07anonymous-stoicheia /AncientGreek Ancient Greek Corpus — re-OCR'd, LLM-repaired A large, deduplicated Ancient Greek corpus, ~361M words in 2,145,799 records. The bulk of it is public-domain page images that we re-OCR'd ourselves with a vision–language model fine-tuned for polytonic Greek (Angleraud et al., 2026), rather than reusing the existing text layers, which roughly quadrupled the usable yield. The noisier output was then repaired by Qwen3.6-27B. Provenance source upstream what we did… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-stoicheia/AncientGreek.textfill-mask1M<n<10M0 likes72 downloads2mo agoHugging Face08NIKAW /ud-ancient-greek-r2.15 Aggregated Universal Dependencies for Ancient Greek Based on the Github r2.15 release of all Ancient Greek corpora, namely PROIEL, Perseus, PTNK. Parsed with conllu. To ensure compatibility with Hugging Face datasets and the underlying Arrow format, some quirks are present in the dataset: the ID field (token_ids) is a string rather than an integer. This is needed because in rare cases the ID can be a token range like 1-2. because features can contain arbitrary content that is only… See the full description on the dataset page: https://huggingface.co/datasets/NIKAW/ud-ancient-greek-r2.15.text10K<n<100K1 likes35 downloads2y agoHugging Face09cvalore /ancient-greek-ds-fim-EVALtext1K<n<10K0 likes5 downloads3mo agoHugging Face10cvalore /ancient-greek-ds-fim-TRAINtext10K<n<100K0 likes3 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.