datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
plazos-normativa-laboral-espana
Normativa laboral y RRHH en España: 300 preguntas con su respuesta y su artículo
Corpus de 300 pares de pregunta y respuesta sobre recursos humanos y normativa laboral española, repartidos en 30 temas y con la referencia normativa en 286 de ellos. Todas las respuestas son autónomas: se entienden sin contexto adicional.
Publicado por Nucleo360, software de recursos humanos para pymes españolas.
Qué cubre
Tema
Preguntas
Registro horario
42
Inspección de… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/plazos-normativa-laboral-espana.preguntas-normativa-laboral-rrhh-espana
Normativa laboral y RRHH en España: 300 preguntas con su respuesta y su artículo
Corpus de 300 pares de pregunta y respuesta sobre recursos humanos y normativa laboral española, repartidos en 30 temas y con la referencia normativa en 286 de ellos. Todas las respuestas son autónomas: se entienden sin contexto adicional.
Publicado por Nucleo360, software de recursos humanos para pymes españolas.
Qué cubre
Tema
Preguntas
Registro horario
42
Inspección de… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/preguntas-normativa-laboral-rrhh-espana.turkish-foreigners-labor-law-benchmark
Turkish Foreigners and Labor Law Benchmark
Turkish Foreigners and Labor Law Benchmark, Türkiye'deki yabancılar ve uluslararası işgücü hukuku alanında büyük dil modellerini ölçmek için hazırlanmış 200 soruluk Türkçe çoktan seçmeli test setidir.
SFT yapılandırması
sft yapılandırması, benchmark'tan ayrı tutulan 2.400 chat örneği içerir:
train: 2.280 örnek
validation: 120 örnek
%35 yabancılar hukuku, %35 uluslararası işgücü, %12 iş hukuku, %8 TBK hizmet sözleşmeleri… See the full description on the dataset page: https://huggingface.co/datasets/enes1863/turkish-foreigners-labor-law-benchmark.polish-labor-code
Rest of your Markdown content starts here...
Polish Labor Code
Original source: https://isap.sejm.gov.pl/isap.nsf/download.xsp/WDU19740240141/U/D19740141Lj.pdf
The dataset contains a structured contents of Polish Labor Code in Markdown format.
The code was generated using docling.
GitHub Repository
austrian-german-instructions
AT-Instruct: Austrian German Instructions
500 instruction-response pairs written in Austrian German. Not translated from English or Bundesdeutsch — written from scratch with Austrian vocabulary, institutions, and perspective.
Why this exists
Every German instruction dataset I found was either translated from English (losing all cultural context) or written in Bundesdeutsch. If you fine-tune on those, your model will tell users to go to the "Bürgeramt" — which doesn't… See the full description on the dataset page: https://huggingface.co/datasets/Laborator/austrian-german-instructions.austrian-german-benchmark
AT-Bench: Austrian German Benchmark
300 multiple-choice questions testing whether an LLM actually understands Austrian German — not just German.
The problem
Every German benchmark treats German as one language. But ask GPT what "Obers" means and half the time it guesses wrong. Ask it about the Bezirksgericht and it describes the German court system. Austrian German is an official language variety spoken by 9 million people, and models consistently get it wrong.… See the full description on the dataset page: https://huggingface.co/datasets/Laborator/austrian-german-benchmark.
