datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
greek_legal_ner
Dataset Card for Greek Legal Named Entity Recognition
Dataset Summary
This dataset contains an annotated corpus for named entity recognition in Greek legislations. It is the first of its kind for the Greek language in such an extended form and one of the few that examines legal text in a full spectrum entity recognition.
Supported Tasks and Leaderboards
The dataset supports the task of named entity recognition.
Languages
The language in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/greek_legal_ner.greek-nlu-bench
Greek NLU Benchmark
Frozen (v1.0) Greek language-understanding benchmark — 9,751 items / 18 tasks, one split per task. Native-verified and train/eval-disjoint. Built to detect CPT/SFT gains and regressions on Greek, weighted toward morphology-sensitive tasks.
Item schema
One JSON object per line; identical envelope across tasks:
{
"id": "el-mcq-000042", "task": "mcq", "language": "el", "version": "1.0",
"tags": {"category": "linguistic", "phenomenon":… See the full description on the dataset page: https://huggingface.co/datasets/KIEFERSA/greek-nlu-bench.GPT-4-Self-Instruct-GreekAs per the community's request, here we share a Greek dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Greek. This dataset will be… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Greek.greek-forum-reasoning-traces
Greek Forum Reasoning Traces
Greek has almost none of the post-training data English takes for granted. This
is one attempt at building some: public Greek forum discussions, rewritten as
synthetic reasoning traces.
Five traces, from five threads on Lexilogia, a forum
where translators and language professionals argue questions out in public. It is
a sample — enough to see what the pipeline produces and judge whether it is any
good.
How a discussion becomes a trace
— the… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-forum-reasoning-traces.Greek_Relation_Extractiongreek-apertus-sft
Greek Apertus SFT datasets
The supervised fine-tuning data of the Greek Apertus project: the GlossAPI team of EELLAK (Open Technologies Alliance) continues the pre-training of swiss-ai/Apertus-8B-2509 on Greek text and then trains it on instructions, with a grant from the Swiss AI Initiative. This repository holds every training arm we assembled, exactly as it went (or goes) to the trainer: one messages list per row, chat format, no system turn.
Access is gated: request it and… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-apertus-sft.Greek_name_entity_recognitionAIKIA_Offensive_Greekancient-greek-ds-fim-EVALalpaca_greek_10000gemma_greek_1000greekelgreek_culture_rag_datasetancient-greek-ds-fim-TRAIN
