CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MultiverseComputingCAI /llm-refusal-evaluation 🛡️ LLM Refusal Evaluation Benchmark This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite. The prompts are organized into three groups: Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse. Chinese Sensitive Topics — prompts that may be censored by China-aligned models. Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse. 📌 Contents Safety Benchmarks JailbreakBench SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.texttext-generation1K<n<10K5 likes240 downloads9mo agoHugging Face02llmeval-fdu /LLMEval-Fair LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models LLMEval-Fair is a 30-month longitudinal study on the robustness and fairness of LLM evaluation, built on a proprietary bank of 220,000+ graduate-level Chinese questions spanning 13 academic disciplines. This is the publicly released subset of that bank. Paper (arXiv): https://arxiv.org/abs/2508.05452 Venue: ACL 2026 Main Conference Project website: https://llmeval.com/ GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Fair.textquestion-answering100K<n<1M0 likes98 downloads5mo agoHugging Face03BushNate /llm-refusal-evaluation 🛡️ LLM Refusal Evaluation Benchmark This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite. The prompts are organized into three groups: Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse. Chinese Sensitive Topics — prompts that may be censored by China-aligned models. Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse. 📌 Contents Safety Benchmarks JailbreakBench SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/BushNate/llm-refusal-evaluation.texttext-generation1K<n<10K1 likes88 downloads2mo agoHugging Face04llmeval-fdu /LLMEval-1 LLMEval-1: Large-Scale Chinese LLM Evaluation (Phase I) LLMEval-1 is the Phase I evaluation dataset of the LLMEval project (Fudan NLP Lab), released alongside the paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models" (AAAI 2024). It is a Chinese, generative-QA benchmark designed to study how large language models should be evaluated. Paper: https://arxiv.org/abs/2312.07398 Project website: https://llmeval.com/ GitHub: https://github.com/llmeval/LLMEval-1… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-1.textquestion-answering1K<n<10K0 likes80 downloads5mo agoHugging Face05llmeval-fdu /LLMEval-2 LLMEval-2: Professional Domain Evaluation of Chinese LLMs (Phase II) LLMEval-2 is the Phase II evaluation dataset of the LLMEval project (Fudan NLP Lab), released alongside the AAAI 2024 paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models". While LLMEval-1 focuses on general capabilities, LLMEval-2 targets professional domain evaluation across 12 academic disciplines. Paper: https://arxiv.org/abs/2312.07398 Project website: https://llmeval.com/ GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-2.textquestion-answering1K<n<10K0 likes77 downloads5mo agoHugging Face06316usman /llm-output-evaluation LLM_OUTPUT_EVALUATION A preference dataset for LLM_OUTPUT_EVALUATION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/llm-output-evaluation.texttext-generation1K<n<10K0 likes69 downloads15d agoHugging Face07llm-compe-2025-kato /step2-evaluated-dataset-test2 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-test2 Total Samples: 92 Successfully Evaluated (Rubric): 92 Failed Evaluations (Rubric): 0 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence: 3.51… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-test2.tabulartext-generationn<1K0 likes30 downloads1y agoHugging Face08llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B-cp32 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32 Total Samples: 60 Successfully Evaluated (Rubric): 53 Failed Evaluations (Rubric): 7 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32.tabulartext-generationn<1K0 likes25 downloads1y agoHugging Face09llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B-cp40 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40 Total Samples: 58 Successfully Evaluated (Rubric): 53 Failed Evaluations (Rubric): 5 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40.tabulartext-generationn<1K0 likes23 downloads1y agoHugging Face10llm-compe-2025-kato /step2-DeepMath-103K-Bespoke-Filtered-Test-200-eval-cp40 Chain of Thought生成データセット このデータセットは、問題と解答から説明(Chain of Thought)を生成したデータセットです。 概要 処理したサンプル数: 58 有効な説明生成数: 58 生成成功率: 100.00% 使用モデル: /home/Competition2025/P07/shareP07/share_model/step2_rlt/Qwen3-14B-step2-deepmath103k-bs512/checkpoint-40 トークン数統計 最小トークン数: 830 最大トークン数: 6301 平均トークン数: 3477.9 トークン数分布 0-100トークン: 0件 (0.0%) 101-500トークン: 0件 (0.0%) 501-1000トークン: 1件 (1.7%) 1001-2000トークン: 9件 (15.5%) 2001-5000トークン: 39件 (67.2%) 5001+トークン: 9件 (15.5%)… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-DeepMath-103K-Bespoke-Filtered-Test-200-eval-cp40.texttext-generationn<1K0 likes21 downloads1y agoHugging Face11llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B Total Samples: 156 Successfully Evaluated (Rubric): 135 Failed Evaluations (Rubric): 21 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B.tabulartext-generationn<1K0 likes20 downloads1y agoHugging Face12llm-compe-2025-kato /step2-DeepMath-103K-Bespoke-Filtered-Test-200-eval-Qwen3-14B Chain of Thought生成データセット このデータセットは、問題と解答から説明(Chain of Thought)を生成したデータセットです。 概要 処理したサンプル数: 156 有効な説明生成数: 156 生成成功率: 100.00% 使用モデル: Qwen/Qwen3-14B トークン数統計 最小トークン数: 313 最大トークン数: 3066 平均トークン数: 887.4 トークン数分布 0-100トークン: 0件 (0.0%) 101-500トークン: 26件 (16.7%) 501-1000トークン: 85件 (54.5%) 1001-2000トークン: 39件 (25.0%) 2001-5000トークン: 6件 (3.8%) 5001+トークン: 0件 (0.0%) データセット構造 system_prompt: モデルに送信されたシステムプロンプト question_text: 元の問題文 answer_text: 問題の解答… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-DeepMath-103K-Bespoke-Filtered-Test-200-eval-Qwen3-14B.texttext-generationn<1K0 likes17 downloads1y agoHugging Face13llm-compe-2025-kato /step2-pairwise-eval Pairwise Comparison Results This dataset contains the results of a pairwise comparison between two models' explanation generation capabilities. Overview Total Comparisons: 68 Model A (Dataset): llm-compe-2025-kato/Tag-Validation_Qwen3-14B-Step1-Bespoke17k-ep1 Model B (Dataset): llm-compe-2025-kato/Tag-Validation_Qwen3-14B-Step1-Bespoke17k-ep1 Results Summary Winner: Model A (Dataset: llm-compe-2025-kato/Tag-Validation_Qwen3-14B-Step1-Bespoke17k-ep1) Margin:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-pairwise-eval.texttext-generationn<1K0 likes15 downloads1y agoHugging Face14llm-compe-2025-kato /step2-DeepMath-103K-Bespoke-Filtered-Test-200-eval-cp32 Chain of Thought生成データセット このデータセットは、問題と解答から説明(Chain of Thought)を生成したデータセットです。 概要 処理したサンプル数: 60 有効な説明生成数: 60 生成成功率: 100.00% 使用モデル: /home/Competition2025/P07/shareP07/share_model/step2_rlt/Qwen3-14B-step2-deepmath103k-bs512/checkpoint-32 トークン数統計 最小トークン数: 892 最大トークン数: 5932 平均トークン数: 3560.4 トークン数分布 0-100トークン: 0件 (0.0%) 101-500トークン: 0件 (0.0%) 501-1000トークン: 1件 (1.7%) 1001-2000トークン: 10件 (16.7%) 2001-5000トークン: 41件 (68.3%) 5001+トークン: 8件 (13.3%)… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-DeepMath-103K-Bespoke-Filtered-Test-200-eval-cp32.texttext-generationn<1K0 likes15 downloads1y agoHugging Face15eliyahabba /llm-evaluation-hujigated DOVE textmultiple-choice1M<n<10M0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.