krimits/greek-hotel-reviews-sentiment
GreekBERT sentiment classifier (Greek hotel-review extension)
Fine-tune of `nlpaueb/bert-base-greek-uncased-v1` on the pinned `DGurgurov/greek_sa` corpus (revision 767a5a0737311829f43cb4f4c90a609b6da4ed50) — the Tsakalidis et al. (2018) Greek sentiment corpus (political tweets, 2015 election period, MIT-licensed).
Domain-transfer framing (stated, not hidden): the training corpus is Twitter political sentiment, not hotel reviews. It is the largest publicly labeled Greek sentiment corpus on the Hub; evaluation of this checkpoint on Greek hotel reviews is declared future work, not an achieved result.
Results (frozen 767-row ordered test split, seed 42)
Sanity: beats the majority baseline by ~39 accuracy points; label mapping verified (0 = negative, 1 = positive, written into the checkpoint config).
An earlier seed of the same config scored 0.9150 test macro-F1; the ~0.6 pp gap between identical-config runs is run-to-run GPU nondeterminism on a small (767-row) test split — see the companion 5-family English benchmark for the variance-discipline story.
Training settings
seed 42 · max_length 160 · batch 32 · lr 2e-5 · weight decay 0.01 · warmup 0.06 · cosine-free linear schedule via Trainer defaults · fp16 · best checkpoint selected on validation macro-F1.
Intended use
Greek-language short-text sentiment (tweets, reviews, comments). Part of the `hotel-review-nlp` project's Greek extension: same provenance discipline as the English pipeline — pinned dataset revision, ordered-split fingerprints, test_logits.npy/test_labels.npy saved in original test order for unified-benchmark compatibility.
Checkpoint integrity fix (v2)
The initial push of this repo omitted the tokenizer files (vocab.txt, tokenizer_config.json, special_tokens_map.json). Any consumer that then called AutoTokenizer.from_pretrained on this repo got a silent fallback to a default BERT vocab: Greek text tokenized into UNKs, and predictions collapsed to the majority class. The three files were added verbatim from the base model and verified (tokens 5/10, unk=0 on probe texts). The reported metrics are unaffected — training and all in-job sanity checks loaded the tokenizer from the base model directly.
Limitations
- Trained on political tweets from 2015: opinions and vocabulary are period- and domain-specific; performance on other Greek domains is untested.
- Cleaned for URLs and @mentions; diacritics and final sigma left untouched (the nlpaueb tokenizer is trained on unnormalized Greek).
- Binary labels only; the source corpus contains no neutral class.
Reproduce
python scripts/train_greek.py --config configs/greek_bert.yaml
python scripts/predict_greek.py --model-dir runs/greek_bert --text "Εξαιρετικό δωμάτιο και εξυπηρετικό προσωπικό"