datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
austrian-german-instructions
AT-Instruct: Austrian German Instructions
500 instruction-response pairs written in Austrian German. Not translated from English or Bundesdeutsch — written from scratch with Austrian vocabulary, institutions, and perspective.
Why this exists
Every German instruction dataset I found was either translated from English (losing all cultural context) or written in Bundesdeutsch. If you fine-tune on those, your model will tell users to go to the "Bürgeramt" — which doesn't… See the full description on the dataset page: https://huggingface.co/datasets/Laborator/austrian-german-instructions.austrian-german-benchmark
AT-Bench: Austrian German Benchmark
300 multiple-choice questions testing whether an LLM actually understands Austrian German — not just German.
The problem
Every German benchmark treats German as one language. But ask GPT what "Obers" means and half the time it guesses wrong. Ask it about the Bezirksgericht and it describes the German court system. Austrian German is an official language variety spoken by 9 million people, and models consistently get it wrong.… See the full description on the dataset page: https://huggingface.co/datasets/Laborator/austrian-german-benchmark.
