Greek
Datasets
All datasets matching “Greek”mmlu_greek
Dataset Card for MMLU Greek
The MMLU Greek dataset is a set of 15858 examples from the MMLU dataset [available from here and here], machine-translated into Greek. The original dataset consists of multiple-choice questions from 57 tasks including elementary mathematics, US history, computer science, law, etc.
Dataset Details
Bias, Risks, and Limitations
This dataset is the result of machine translation.
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/mmlu_greek.GreekMMLU
GreekMMLU
GreekMMLU is a native-sourced benchmark for evaluating massive multitask language understanding in Greek, built from authentic Greek exam-style multiple-choice questions (MCQ) rather than machine-translated English benchmarks.
21,805 questions across 45 subjects
4 high-level groups: STEM, Humanities, Social Sciences, Other
Difficulty/education levels spanning Primary → Secondary → University → Professional (+ an N/A bucket)
Public vs. private split for… See the full description on the dataset page: https://huggingface.co/datasets/dascim/GreekMMLU.greek-cc
Greek Common Crawl
A FineWeb-style Greek-language text dataset extracted from Common Crawl, following the FineWeb-2 recipe adapted for Greek (ell_Grek).
Pipeline source: github.com/alexliap/greek-cc.
Crawl coverage starts at CC-MAIN-2024-22 rather than Common Crawl's earliest snapshots: this project picks up right where the FineWeb-2 dataset's own Greek (ell_Grek) subset leaves off (2013 through April 2024), so it extends FineWeb-2's Greek coverage forward instead of… See the full description on the dataset page: https://huggingface.co/datasets/alexliap/greek-cc.open-greek-corpus-annotations
Open Greek Corpus Annotations
Token-level linguistic annotations for the
Open Greek Corpus:
lemma, part of speech (UD UPOS), and morphology (UD features) for every
served token. Three provenance classes, never confused thanks to per-token
provenance and confidence tiers: gold treebank annotations where an openly
licensed MANUAL treebank covers a work (GLAUx's treebank layers, MACULA
Greek for the NT), GLAUx's own automatic annotation as the middle auto:
class, and model… See the full description on the dataset page: https://huggingface.co/datasets/ciscoriordan/open-greek-corpus-annotations.greek-contentsgreek-corpus-150b
Greek Corpus 150B
A large-scale, deduplicated Greek (Modern Greek, el) text corpus for training and fine-tuning foundation models. It pairs a broad web/knowledge/formal-document pretrain layer with a multilingual-instruction SFT layer, all normalized to a single unified schema and globally deduplicated.
This is part of an ongoing Global Corpus family of per-language foundation-model datasets (Dutch, Turkish, Bulgarian, Greek, …) built on a consistent architecture so that sources… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/greek-corpus-150b.
