datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STOP
🛑 STOP
This is the repository for STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions, a dataset comprised of 450 offensive progressions designed to target evolving scenarios of bias and quanitfy the threshold of appropriateness. This work was published in the 2024 Main Conference on Empirical Methods in Natural Language Processing and was honoured with the Social Impact Award.
Authors: Robert Morabito, Sangmitra Madhusudan, Tyler McDonald… See the full description on the dataset page: https://huggingface.co/datasets/Robert-Morabito/STOP.pl-legal-instruct-sample
PL Legal Instruct — darmowy sample (50 par)
Polski dataset instrukcyjny z domeny prawnej. Pary instruction-output
generowane z publicznych orzeczen sadow (API SAOS), filtrowane
LLM-as-judge (prog 7/10) i recznie spot-checkowane.
Aktualny rozmiar pelnego korpusu: 3644 par (rosnie codziennie).
Pelne pakiety (kupujesz raz, pobierasz zawsze aktualna wersje)
Standard (149 zl): https://robertian040.gumroad.com/l/pl-legal-standard
Pro (449 zl):… See the full description on the dataset page: https://huggingface.co/datasets/Robert19001/pl-legal-instruct-sample.
