datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
r_judge_labelled
R-Judge with LLM-Judge Labels
This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains.
Files
File
Description
r_judge_data.csv
Base dataset extracted from R-Judge (568 rows, deduplicated)
r_judge_labelled_anthropic_claude-sonnet-4-6.csv
Base dataset augmented with… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/r_judge_labelled.CSSR-S_labelled_suicidewatch_posts_reddit
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code.
License and Citation
This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following:
@article{patil2025evaluating,
title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.r_judge_labelled
R-Judge with LLM-Judge Labels
This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains.
Files
File
Description
r_judge_data.csv
Base dataset extracted from R-Judge (568 rows, deduplicated)
r_judge_labelled_anthropic_claude-sonnet-4-6.csv
Base dataset augmented with LLM-judge… See the full description on the dataset page: https://huggingface.co/datasets/imerad-kv/r_judge_labelled.labelled_articleslabelled-PubMedQAlabelled_datasetscommon_voice_13_0_zh_pseudo_labelledlabelled_hatespeechlabelled_regex
Labelled Regex
This dataset consists of Regexes and their descriptive labels. As far as I am aware, this is the largest, cleanly labelled regex dataset on this platform.
I constructed this dataset by taking innovatorved/regex_dataset and using gemma-3-27b-it LLM to generate a concise and suitable title for each regex.
For each regex that was larger than 100 characters, I used a slightly different prompt to generate an even more detailed description.
common_voice_13_0_thai_small_pseudo_labelledwolof_kallaama_pseudo_labelledSkill_labelled_MATHfake_news_elections_labelled_data
Dataset Card for Election-Related Fake News Classification
Dataset Summary
This dataset is designed for the task of fake news classification in the context of elections. It consists of news articles, social media posts, and other text sources related to various elections worldwide. Each entry in the dataset is labeled as 'fake' or 'real' based on its content and the veracity of the information presented.
https://arxiv.org/abs/2312.03750
Languages
English… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/fake_news_elections_labelled_data.ai4bharat_telugu_pseudo_labelledmn_pseudo_labelled_2synthetic-resume-fit-labelledstsb-binary-paraphrase-labelled
Paraphrase Detection Dataset (Derived from SetFit/stsb)
Description:
This dataset originates from the SetFit/stsb dataset, which was initially created for semantic textual similarity (STS) tasks with a label range of 0 to 5. It has been adapted for binary paraphrase detection by leveraging the high-accuracy paraphrase classification model viswadarshan06/pd-robert.
Each sentence pair in the original dataset has been re-labeled according to the following binary scheme:
1 →… See the full description on the dataset page: https://huggingface.co/datasets/viswadarshan06/stsb-binary-paraphrase-labelled.pseudo_labelledmn_pseudo_labelled_1wenetspeech_zh_TW_pseudo_labelled_large_v3_non_concatcommon_voice_16_1_hi_pseudo_labelledFAKE-NEWS-BIASES-LABELLEDNews_Train_LABELLEDcommon_voice_16_0_fa_pseudo_labelledlabelled_cricketcommon_voice_16_1_sw_pseudo_labelledcricket_labelled-2fact_updates_GPT_labelledSS_dataset_FINAL_VERSION_LABELLEDcommon_voice_16_1_sw2_pseudo_labelled
