datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
movielens-1m-ratings-standardizedrating_scores_3rules_hhrlhfrating_scores_10rules_hhrlhfalpaca-cleaned-gemini-hun-ratingsEz az adathalmaz úgy keletkezett, hogy a Bazsalanszky/alpaca-cleaned-gemini-hun-n lefuttattam egy llm által támogatott értékelést.
Az értékelő modell a gemini-pro (az ingyenes) volt. A használt kód az alpagasus módosítása: https://github.com/boapps/alpagasus-hu
K-HATERS-Ratingsnfpa-704-hazard-ratings
NFPA 704 fire diamond ratings by chemical
Canonical, always-current version: https://referencesource.org/nfpa-704-hazard-ratings/
Machine-readable: https://referencesource.org/nfpa-704-hazard-ratings/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-06
Stale after: 2028-08-05 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 4
The NFPA 704 'fire diamond' rating for a chemical — the health (blue)… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nfpa-704-hazard-ratings.prompt-difficulty-model-ratings
Prompt Difficulty Model Ratings
Dataset contains approximately 100 000 ChatGPT prompts from agentlans/chatgpt
The prompts were rated for difficulty using the large language models:
allenai/Olmo-3-7B-Instruct
google/gemma-3-12b-it
ibm-granite/granite-4.0-h-tiny
meta-llama/Llama-3.1-8B-Instruct
microsoft/phi-4
mistralai/Ministral-3-8B-Instruct-2512nvidia/NVIDIA-Nemotron-Nano-9B-v2
Qwen/Qwen3-8B
swiss-ai/Apertus-8B-Instruct-2509
tiiuae/Falcon-H1-7B-Instruct
Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-difficulty-model-ratings.atc-tts-mos-ratingscredit-rating-methodologybudapest-v0.1-hun-ratingsEz az adathalmaz úgy keletkezett, hogy a Bazsalanszky/budapest-v0.1-hun-n lefuttattam egy llm által támogatott értékelést.
Az értékelő modell a gemini-pro (az ingyenes) volt. A használt kód az alpagasus módosítása: https://github.com/boapps/alpagasus-hu
rating_scores_5rules_hhrlhfalpaca_hu_2k-ratingsEz az adathalmaz úgy keletkezett, hogy az NYTK/alpaca_hu_2k-n lefuttattam egy llm által támogatott értékelést.
Az értékelő modell a gemini-pro (az ingyenes) volt. A használt kód az alpagasus módosítása: https://github.com/boapps/alpagasus-hu
alpaca_hu_mt-ratingsalpaca-hu-v2-ratingsEz az adathalmaz úgy keletkezett, hogy az alpaca-hu-v2-n lefuttattam egy llm által támogatott értékelést.
Ez első ránézésre elég királyul kiszűri (0-ás ratinget ad) a Gemini random halandzsáira.
Az értékelő modell a gemini-pro (az ingyenes) volt. A használt kód az alpagasus módosítása: https://github.com/boapps/alpagasus-hu
atc-tts-llm-mos-ratings4chan_archive_ShareGPT_with_rating_and_commentsI took adamo1139/4chan_archive_ShareGPT_fixed_newlines_unfiltered and removed broken samples, like those having no two-sided conversation or containing empty responses (most likely someone just posted an image there and that wasn't scraped).
Then I used Hermes 3 8B (mostly W8A8) to add comments to each sample and add a final score, from 0 to 5. The result of this is this dataset. I had to process a few billions tokens to create this.
I now plan to further filter down the dataset and most… See the full description on the dataset page: https://huggingface.co/datasets/adamo1139/4chan_archive_ShareGPT_with_rating_and_comments.sensory-modality-ratingsconcreteness_phrase_ratingsmrc_imageability_ratingsRating1000conreteness_ratingsESG_ratingsultrafeedback-gpt-single-ratingsummary-ratingsrating_scores_5rules_pkuspeaking-ratingprincipled_ratingsatc-tts-mos-large-axite-ratingsCheckin-Ratings
