CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01garak-llm /pypi-20241031text100K<n<1M2 likes9.5k downloads2y agoHugging Face02bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.2k downloads2y agoHugging Face03garak-llm /npm-20241031text1M<n<10M1 likes8.5k downloads2y agoHugging Face04garak-llm /crates-20250307text100K<n<1M0 likes8.5k downloads2y agoHugging Face05bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K195 likes7.8k downloads2y agoHugging Face06garak-llm /rubygems-20241031text100K<n<1M0 likes6.3k downloads2y agoHugging Face07garak-llm /npm-20240828text1M<n<10M2 likes6.2k downloads2y agoHugging Face08garak-llm /crates-20240903text100K<n<1M1 likes5.9k downloads2y agoHugging Face09Anthropic /llm_global_opinions Dataset Card for GlobalOpinionQA Dataset Summary The data contains a subset of survey questions about global issues and opinions adapted from the World Values Survey and Pew Global Attitudes Survey. The data is further described in the paper: Towards Measuring the Representation of Subjective Global Opinions in Language Models. Purpose In our paper, we use this dataset to analyze the opinions that large language models (LLMs) reflect on complex global… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/llm_global_opinions.text1K<n<10K60 likes1.6k downloads3y agoHugging Face10bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.4k downloads2y agoHugging Face11minnesotanlp /LLM-Artifacts Under the Surface: Tracking the Artifactuality of LLM-Generated Data Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶ Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo, Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu Dongyeop Kang Minnesota NLP, University of Minnesota Twin Cities † Project Lead, ¶ Core Contribution, Arxiv Project Page 📌 Table of Contents Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.tabular100K<n<1M2 likes1.3k downloads3y agoHugging Face12Xiaolong-Han /w2t-llm-arc-easy-lora W2T Llm Arc Easy Lora This repository contains artifacts for the W2T paper: Paper: W2T: LoRA Weights Already Know What They Can Do Repo: Weight2Token Summary ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction. Source Status Storage location: local Verification status: confirmed Files See manifest.json for the exact local or remote source paths used to prepare this release. Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.tabular10K<n<100K0 likes1.1k downloads4mo agoHugging Face13bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes881 downloads2y agoHugging Face14bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes770 downloads2y agoHugging Face15orbench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our leaderboard at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.imagetext-generation10K<n<100K0 likes596 downloads2y agoHugging Face16bxiong /rl_llm_experiment_p6tabularn<1K0 likes559 downloads1y agoHugging Face17llmlatency /llm-latency-tracker LLM Latency Tracker Independent, continuously measured latency and availability for AI inference API providers, aggregated by day. Covers 45 providers across 4 regions (ap-tokyo, eu-hetzner, sa-east, us-central), built from 3,185,268 raw probes collected between 2026-07-23 and 2026-09-22. Live rankings and full methodology: llmlatency.dev How the numbers are produced Probes run every five minutes from separate network locations and are never routed through a… See the full description on the dataset page: https://huggingface.co/datasets/llmlatency/llm-latency-tracker.tabular10K<n<100K2 likes467 downloads13h agoHugging Face18sunny1820f /llm-pulse-datatext1K<n<10K0 likes450 downloads13d agoHugging Face19blanchon /snac_llm_parler_ttstabular100K<n<1M6 likes447 downloads2y agoHugging Face20nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes445 downloads6d agoHugging Face21mario0369 /llm-cost-same-prompt Measured per-call LLM cost — same prompt, every model Vendors publish prices per million tokens. Nobody publishes what one call actually costs, because that depends on how many tokens the model chooses to emit — and on the same question models differ by more than an order of magnitude. One model finishes a JSON extraction in 23 tokens; another writes 300. This dataset sends a fixed set of prompts to every model at temperature 0, every night, and records the cost computed from… See the full description on the dataset page: https://huggingface.co/datasets/mario0369/llm-cost-same-prompt.tabular1K<n<10K1 likes419 downloads22h agoHugging Face22seantw /DEBATE_LLM DEBATE Benchmark This repository contains CSV files from the DEBATE project: large-scale human conversation experiments organized around controversial and opinion-based topics. The data consists of multi-round conversations between human participants discussing political, social, and belief-related topics, following the protocol described in: Chuang, Y.-S., Tu, R., Dai, C., Vasani, S., Li, Y., Yao, B., Tessler, M. H., Yang, S., Shah, D., Hawkins, R., Hu, J., & Rogers, T. T. (2026).… See the full description on the dataset page: https://huggingface.co/datasets/seantw/DEBATE_LLM.tabular100K<n<1M4 likes368 downloads5mo agoHugging Face23neemiasbsilva /multimodal-LLMs-See-Sentiment MLLMsent — datasets and experiment results Every input and every output of "Multimodal LLMs See Sentiment" (arXiv:2508.16873): the image descriptions generated by six multimodal LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete per-fold results of all 141 experiments. Paper: arXiv:2508.16873 Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.texttext-classification10K<n<100K1 likes339 downloads1mo agoHugging Face24bench-llms /or-bench-toxic-all OR-Bench: An Over-Refusal Benchmark for Large Language Models This dataset constains highly toxic prompts, use with caution!!! Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.imagetext-generation10K<n<100K1 likes336 downloads2y agoHugging Face25akmaier /LLM-Ads LLM-Ads — Sponsored-recommendation evaluation traces Per-trial responses and labels from the experiments in Just Ask for a Table: A Thirty-Token User Prompt Defeats Sponsored Recommendations in Twelve LLMs (arXiv:2605.12772). The data set reproduces and extends the evaluation of Wu et al.\ 2026 (arXiv:2604.08525) on a twelve-model pool (ten open-source chat models served through an OpenAI-compatible API endpoint plus the two paper-overlap OpenAI models gpt-3.5-turbo and gpt-4o).… See the full description on the dataset page: https://huggingface.co/datasets/akmaier/LLM-Ads.tabulartext-classification10K<n<100K0 likes279 downloads4mo agoHugging Face26bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes259 downloads2y agoHugging Face27MBZUAI-LLM /Mobile-MMLUtext10K<n<100K6 likes255 downloads2y agoHugging Face28aida-ugent /llm-censorship Dataset Details Dataset Description This dataset measures soft censorship (selective omission of information) in large language models (LLMs). It contains responses from 14 state-of-the-art LLMs from different regions (Western countries, China, and Russia) when prompted about political figures in all six official UN languages. The dataset is designed to provide insights into how and when LLMs refuse to provide information or selectively omit details when discussing… See the full description on the dataset page: https://huggingface.co/datasets/aida-ugent/llm-censorship.text1M<n<10M3 likes244 downloads1y agoHugging Face29copenlu /llm-pct-tropes Dataset Card for LLM Tropes arXiv: https://arxiv.org/abs/2406.19238v1 Dataset Details Dataset Description This is the dataset LLM-Tropes introduced in paper "Revealing Fine-Grained Values and Opinions in Large Language Models" Dataset Sources Repository: https://github.com/copenlu/llm-pct-tropes Paper: https://arxiv.org/abs/2406.19238 Structure ├── Opinions │   ├── demographic <- Generations for the demographic prompting setting │… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/llm-pct-tropes.tabulartext-generation100K<n<1M5 likes229 downloads2y agoHugging Face30bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K8 likes215 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.