datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
r_judge_labelled
R-Judge with LLM-Judge Labels
This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains.
Files
File
Description
r_judge_data.csv
Base dataset extracted from R-Judge (568 rows, deduplicated)
r_judge_labelled_anthropic_claude-sonnet-4-6.csv
Base dataset augmented with… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/r_judge_labelled.CSSR-S_labelled_suicidewatch_posts_reddit
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code.
License and Citation
This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following:
@article{patil2025evaluating,
title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.Multilingual-USAS-Labelled-Silver-Wikipedia
Multilingual USAS Silver Labelled Wikipedia Articles
Silver-labelled Wikipedia article text for training USAS semantic taggers and Multi-Word
Expression (MWE) identifiers, covering 8 Wikipedia language sites. The source text comes from the
HuggingFace HuggingFaceFW/finewiki
dataset, restricted to articles rated Good (GA) or Featured (FA) — using the article ID list from
ucrelnlp/wikipedia-ga-fa-ids — and
then sentence split and automatically tagged with USAS semantic tags and… See the full description on the dataset page: https://huggingface.co/datasets/ucrelnlp/Multilingual-USAS-Labelled-Silver-Wikipedia.150k_keyphrases_labelledThis dataset is a list of important keyphrases for academic topics. The keyphrases file contains around 150,000 keyphrases that are given about 300-400 labels, such as algorithm, disease, theorem, lemma, chemial compound, research methods, fields, subfields, topics, ect.
We also have about 2 million additional keyphrases that have been sorted into unigrams, bigrams, trigrams and fourgrams.
The keyphrases were obtain fromed a mixture of webscraping academic databases such as pubmed, wikipedia.… See the full description on the dataset page: https://huggingface.co/datasets/ClovenDoug/150k_keyphrases_labelled.topics_labelledr_judge_labelled
R-Judge with LLM-Judge Labels
This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains.
Files
File
Description
r_judge_data.csv
Base dataset extracted from R-Judge (568 rows, deduplicated)
r_judge_labelled_anthropic_claude-sonnet-4-6.csv
Base dataset augmented with LLM-judge… See the full description on the dataset page: https://huggingface.co/datasets/imerad-kv/r_judge_labelled.1.4b-policy_preference_data_gold_labelled_with_refdataset_cards_with_metadata_labelledlabelled-PubMedQAsummarize_from_feedback_tldr3_labelled_vllm_20k_dpo_costa_1b_fp16.yml_3d94f50_b9ff2summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144_labelledfr_sexism_labelled
Dataset Card for "fr_sexism_labelled"
Based on the Kaggle dataset Sexist Workplace Statements.
This dataset features more than 1100 examples of statements of workplace sexism, roughly balanced between examples of certain sexism and ambiguous or neutral cases (labeled with a “1” and “0” respectively).
The original English dataset has been translated into French via machine translation with the Helsinki-NLP/opus-mt-en-fr model.
stsb-binary-paraphrase-labelled
Paraphrase Detection Dataset (Derived from SetFit/stsb)
Description:
This dataset originates from the SetFit/stsb dataset, which was initially created for semantic textual similarity (STS) tasks with a label range of 0 to 5. It has been adapted for binary paraphrase detection by leveraging the high-accuracy paraphrase classification model viswadarshan06/pd-robert.
Each sentence pair in the original dataset has been re-labeled according to the following binary scheme:
1 →… See the full description on the dataset page: https://huggingface.co/datasets/viswadarshan06/stsb-binary-paraphrase-labelled.summarize_from_feedback_tldr3_labelled_generated_relabel_20k_dpo_costa_1b_fp16.yml_3d94f50_b9ff2extracted_distil_n_batched_labelled_llm_dsmy_ncc_labelled_datasetslabelled_cricketcricket_labelled-2blbooks-labelledeicu-tsc-labelledSPP_labelled_3783rows_baseline_flippedsangeetkar-labelled-data
🎵 Sangeetkar Teacher Labels (CLAP-Generated)
This dataset contains high-fidelity "soft labels" for 6,870 music tracks, specifically curated for training lightweight Music Emotion Recognition (MER) models. The labels were generated using LAION-CLAP (laion/clap-htsat-fused) acting as a Teacher model.
🏗 Project Architecture
The goal of this project is to distill the knowledge of a heavy, billion-parameter transformer (CLAP) into a smaller "Student" model that can run… See the full description on the dataset page: https://huggingface.co/datasets/beastLucifer/sangeetkar-labelled-data.labelled_articlesrecency_condition_labelledlabelled_review_datarecency_condition_labelled_version_2
