datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
diavgeia
Diavgeia — Greek Government Transparency Decisions
Dataset Info
This dataset contains the full text and metadata of public-sector decisions
(αποφάσεις / πράξεις) published on Diavgeia (diavgeia.gov.gr),
the Greek government's transparency portal. Since 2010 (Law 3861/2010), every
Greek public entity is legally required to publish its administrative acts —
budget commitments, expenditure approvals, contracts, appointments, regulatory
acts, and more — on Diavgeia… See the full description on the dataset page: https://huggingface.co/datasets/glossAPI/diavgeia.archetai
Archetai
Dataset Language:
Modern and Ancient Greek (with a small number of foreign-language items, mostly English, German, and French summaries or full volumes inside the same collections)
Dataset Info:
This dataset consists of OCR-extracted text from the digital publications archive of the Archaeological Society at Athens (Η εν Αθήναις Αρχαιολογική Εταιρεία, archetai.gr). The Society, founded in 1837, is one of the oldest learned societies in Greece and the principal Greek… See the full description on the dataset page: https://huggingface.co/datasets/glossAPI/archetai.Ellinika_Keimena_Project_Gutenberg
Ellinika Keimena – Project Gutenberg Greek Text Corpus
Πληροφορίες για τα δεδομένα
Το παρόν dataset περιλαμβάνει 214 κείμενα σε όλες τις ποικιλίες της ελληνικής γλώσσας (Αρχαία, Καθαρεύουσα, Δημοτική, ΚΝΕ).
Τα κείμενα αντλήθηκαν από το Project Gutenberg και αφορούν έργα της ελληνικής γραμματείας και δοκίμια που έχουν αποχαρακτηριστεί ως πνευματική ιδιοκτησία.
Τα αρχεία προσφέρονται σε μορφή parquet μαζί με τα μεταδεδομένα τους όπως συγγραφέας / μεταφραστής, χρονολογία… See the full description on the dataset page: https://huggingface.co/datasets/glossAPI/Ellinika_Keimena_Project_Gutenberg.libiepLIBIEP - Historical Greek School Textbooks
Dataset Language:
Greek (with extensive coverage of historical written varieties: Καθαρεύουσα, Απλή Καθαρεύουσα, Δημοτική, plus a long tail of Αρχαία Ελληνικά, Λατινικά, Γαλλικά, Γερμανικά, Τουρκικά, Αγγλικά, Καραμανλίδικα).
Dataset Info:
This dataset is a structured snapshot of the Historical Collection of School Textbooks of the Institute of Educational Policy (Ινστιτούτο Εκπαιδευτικής Πολιτικής - IEP), the public-law body that advises the Greek… See the full description on the dataset page: https://huggingface.co/datasets/glossAPI/libiep.psepheda
Psepheda (ΨΗΦΙΔΑ) — University of Macedonia Institutional Repository
Dataset Language
Modern Greek (primary), with a minority of English-language academic content and occasional multilingual passages (Slavic/Balkan-studies material includes Cyrillic).
Dataset Info
This dataset is a cleaned, full-text snapshot of ΨΗΦΙΔΑ (Psepheda), the institutional repository and digital library of the University of Macedonia (Πανεπιστήμιο Μακεδονίας, Thessaloniki, Greece), served from… See the full description on the dataset page: https://huggingface.co/datasets/glossAPI/psepheda.elocus
E-Locus
Dataset Language:Greek (primary), with English alternative titles, abstracts and bibliographic terms throughout, and occasional French or German material in older volumes. The content column carries the full Greek text of each thesis; the alt_title and most of the subjects field are English by convention of the source repository.
Dataset Info:This dataset consists of text-extracted PDFs from E-Locus (elocus.lib.uoc.gr), the Institutional Repository of the Library and… See the full description on the dataset page: https://huggingface.co/datasets/glossAPI/elocus.glossapi-greek-nanochat-pretraining-dataset
Glossapi Greek Nanochat Pretraining Dataset
This repository contains the source-separated Greek corpus used to build nanochat Greek pretraining mixtures. It is intentionally not a pre-split train/validation/test export: builders load data/*.parquet, preserve source_dataset, and create deterministic experiment-specific mixes and splits downstream.
Current Snapshot
Total rows: 49474947
Total characters: 248276390721
Included source datasets: 19
Data files: 273… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset.
