datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
liberty-archive-austrian-economics
The Liberty Archive: Austrian Economics Corpus
Austrian economics in bulk: 19,531,994 words across two splits.
Full text of the Creative Commons licensed books in the Liberty Archive, one row
per chapter. 2,582 chapters from 176 books, 7,132,091 words.
Machine transcripts of 2,701 recorded lectures, 12,399,903 words,
sit in a second split with a different and weaker rights basis than the books.
Read "Lecture rights: what is and is not established" before using them. Books and… See the full description on the dataset page: https://huggingface.co/datasets/davidveksler/liberty-archive-austrian-economics.austrian-german-instructions
AT-Instruct: Austrian German Instructions
500 instruction-response pairs written in Austrian German. Not translated from English or Bundesdeutsch — written from scratch with Austrian vocabulary, institutions, and perspective.
Why this exists
Every German instruction dataset I found was either translated from English (losing all cultural context) or written in Bundesdeutsch. If you fine-tune on those, your model will tell users to go to the "Bürgeramt" — which doesn't… See the full description on the dataset page: https://huggingface.co/datasets/Laborator/austrian-german-instructions.
