CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face02LLM-Digital-Twin /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.imagetext-classification1K<n<10K33 likes2.7k downloads6mo agoHugging Face03bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes734 downloads2y agoHugging Face04orbench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our leaderboard at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.imagetext-generation10K<n<100K0 likes608 downloads2y agoHugging Face05bench-llms /or-bench-toxic-all OR-Bench: An Over-Refusal Benchmark for Large Language Models This dataset constains highly toxic prompts, use with caution!!! Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.imagetext-generation10K<n<100K1 likes353 downloads2y agoHugging Face06niveck /LLMafia LLMafia - Asynchronous LLM Agent Our Mafia game dataset of an Asynchronous LLM Agent playing games of Mafia with multiple human players. 🌐 Project | 📃 Paper | 💻 Code A virtual game of Mafia, played by human players and an LLM agent player. The agent integrates in the asynchronous group conversation by constantly simulating the decision to send a message. Time to Talk: 🕵️‍♂️ LLM Agents for Asynchronous Group Communication in Mafia Games Niv Eckhaus, Uri Berger, Gabriel… See the full description on the dataset page: https://huggingface.co/datasets/niveck/LLMafia.imagetext-generationn<1K7 likes154 downloads11mo agoHugging Face07llm-addiction-research /llm-addictiongated LLM Addiction Research Dataset Behavioral and neural data from experiments studying gambling-like behaviors in Large Language Models. Paper: "Can Large Language Models Develop Gambling Addiction?" — NeurIPS 2026 (submission 24231), camera-ready Authors: Seungpil Lee, Donghyeon Shin, Yunjeong Lee, Sundong Kim (GIST) Code: github.com/iamseungpil/llm-addiction Quick Links Paper → file map: sae_v3_analysis/results/reports/paper_asset_manifest.md (every claim → exact… See the full description on the dataset page: https://huggingface.co/datasets/llm-addiction-research/llm-addiction.imagetext-generationn<1K1 likes149 downloads2d agoHugging Face08amalia-llm /MATH-Vision-PT MATH-Vision-PT European Portuguese (pt-PT) machine translation of MATH-Vision, a benchmark of competition-level mathematics problems presented in visual contexts. Translated from the original English test split using gemini-3.1-pro. Original Dataset: https://huggingface.co/datasets/MathLLMs/MathVision Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MATH-Vision-PT.imagequestion-answering1K<n<10K0 likes35 downloads3mo agoHugging Face09Felguk /LLM-iconsimagetext-generationn<1K0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.