CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01minnesotanlp /LLM-Artifacts Under the Surface: Tracking the Artifactuality of LLM-Generated Data Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶ Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo, Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu Dongyeop Kang Minnesota NLP, University of Minnesota Twin Cities † Project Lead, ¶ Core Contribution, Arxiv Project Page 📌 Table of Contents Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.tabular100K<n<1M2 likes1.3k downloads3y agoHugging Face02NoeFlandre /benchmark-llms-landuse-relevance Land-use relevance benchmark v3-multilingual · 85 languages x 300 items/language · 25,500 items · binary yes/no labels. Code Task and prompt Does a sentence describe a place's land or environment in ways visible to satellites? English prompt · greedy decoding · seed 0 · max_new_tokens=4096 · bfloat16 · batch varies by model. Prompt text Replace {} with the target sentence. Classify whether the TARGET SENTENCE contains information about the target… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.tabulartext-classification10K<n<100K0 likes1.1k downloads17m agoHugging Face03Xiaolong-Han /w2t-llm-arc-easy-lora W2T Llm Arc Easy Lora This repository contains artifacts for the W2T paper: Paper: W2T: LoRA Weights Already Know What They Can Do Repo: Weight2Token Summary ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction. Source Status Storage location: local Verification status: confirmed Files See manifest.json for the exact local or remote source paths used to prepare this release. Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.tabular10K<n<100K0 likes1.1k downloads4mo agoHugging Face04bxiong /rl_llm_experiment_p6tabularn<1K0 likes560 downloads1y agoHugging Face05llmlatency /llm-latency-tracker LLM Latency Tracker Independent, continuously measured latency and availability for AI inference API providers, aggregated by day. Covers 45 providers across 4 regions (ap-tokyo, eu-hetzner, sa-east, us-central), built from 3,237,730 raw probes collected between 2026-07-23 and 2026-09-23. Live rankings and full methodology: llmlatency.dev How the numbers are produced Probes run every five minutes from separate network locations and are never routed through a… See the full description on the dataset page: https://huggingface.co/datasets/llmlatency/llm-latency-tracker.tabular10K<n<100K2 likes461 downloads6h agoHugging Face06nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes447 downloads7d agoHugging Face07blanchon /snac_llm_parler_ttstabular100K<n<1M6 likes446 downloads2y agoHugging Face08mario0369 /llm-cost-same-prompt Measured per-call LLM cost — same prompt, every model Vendors publish prices per million tokens. Nobody publishes what one call actually costs, because that depends on how many tokens the model chooses to emit — and on the same question models differ by more than an order of magnitude. One model finishes a JSON extraction in 23 tokens; another writes 300. This dataset sends a fixed set of prompts to every model at temperature 0, every night, and records the cost computed from… See the full description on the dataset page: https://huggingface.co/datasets/mario0369/llm-cost-same-prompt.tabular1K<n<10K1 likes415 downloads16h agoHugging Face09seantw /DEBATE_LLM DEBATE Benchmark This repository contains CSV files from the DEBATE project: large-scale human conversation experiments organized around controversial and opinion-based topics. The data consists of multi-round conversations between human participants discussing political, social, and belief-related topics, following the protocol described in: Chuang, Y.-S., Tu, R., Dai, C., Vasani, S., Li, Y., Yao, B., Tessler, M. H., Yang, S., Shah, D., Hawkins, R., Hu, J., & Rogers, T. T. (2026).… See the full description on the dataset page: https://huggingface.co/datasets/seantw/DEBATE_LLM.tabular100K<n<1M4 likes374 downloads5mo agoHugging Face10akmaier /LLM-Ads LLM-Ads — Sponsored-recommendation evaluation traces Per-trial responses and labels from the experiments in Just Ask for a Table: A Thirty-Token User Prompt Defeats Sponsored Recommendations in Twelve LLMs (arXiv:2605.12772). The data set reproduces and extends the evaluation of Wu et al.\ 2026 (arXiv:2604.08525) on a twelve-model pool (ten open-source chat models served through an OpenAI-compatible API endpoint plus the two paper-overlap OpenAI models gpt-3.5-turbo and gpt-4o).… See the full description on the dataset page: https://huggingface.co/datasets/akmaier/LLM-Ads.tabulartext-classification10K<n<100K0 likes291 downloads4mo agoHugging Face11copenlu /llm-pct-tropes Dataset Card for LLM Tropes arXiv: https://arxiv.org/abs/2406.19238v1 Dataset Details Dataset Description This is the dataset LLM-Tropes introduced in paper "Revealing Fine-Grained Values and Opinions in Large Language Models" Dataset Sources Repository: https://github.com/copenlu/llm-pct-tropes Paper: https://arxiv.org/abs/2406.19238 Structure ├── Opinions │   ├── demographic <- Generations for the demographic prompting setting │… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/llm-pct-tropes.tabulartext-generation100K<n<1M5 likes227 downloads2y agoHugging Face12SIP-med-LLM /JMedQA JMedQA: Benchmarking Large Language Models and Vision-Language Models on the Japanese Medical Licensing Examination JMedQA is a Japanese medical question-answering benchmark derived from Japan's National Medical Examination materials publicly released by the Ministry of Health, Labour and Welfare (MHLW). The dataset supports both text-only large language model (LLM) evaluation and vision-language model (VLM) evaluation using associated examination images. Its image-dependency… See the full description on the dataset page: https://huggingface.co/datasets/SIP-med-LLM/JMedQA.imagequestion-answering1K<n<10K0 likes197 downloads28d agoHugging Face13oxford-llms /ai-respondents-challenge AI Respondents Challenge — Oxford LLMs 2026 Predict a survey respondent's answer to a held-out question from their other answers (World Values Survey wave 7). Any method allowed; you must disclose the features and prompts you used. Ranked on normalized skill + distributional alignment, on in-domain and out-of-domain (held-out countries) boards. Configs train — 5,000 labeled respondents (100 per seen country): respondent_id, country + all WVS variables… See the full description on the dataset page: https://huggingface.co/datasets/oxford-llms/ai-respondents-challenge.tabular1K<n<10K1 likes185 downloads2mo agoHugging Face14CarlosGI /llm-bargaining-transcripts LLM Bargaining Transcripts 240 complete two-agent bargaining games between large language models, played under an alternating-offers protocol with private valuations, discounting, and cheap talk. Every game records both agents' true valuations, their private reasoning, what they claimed about their own position, and what they actually did. The dataset is designed to make misrepresentation measurable. Because the true valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.tabular1K<n<10K1 likes162 downloads18d agoHugging Face15mayank-dubey-ai /l4-gpu-llm-benchmark-leaderboard 🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB) An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU. 📊 Executive Summary & Key Takeaways ⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.tabulartext-generationn<1K0 likes155 downloads1mo agoHugging Face16bxiong /rl_llm_experiment_p9tabularn<1K0 likes152 downloads1y agoHugging Face17ibm-research /LLMFineTuningBench Dataset Card for LLMFineTuningBench A dataset of over 30,000 LLM fine-tuning experiments, capturing detailed performance metrics from jobs run on high-performance computing (HPC) clusters. It spans a wide range of models, fine-tuning methods, and hardware configurations, and is intended to support research on predictive resource allocation, performance optimization, and cost estimation for LLM fine-tuning workloads. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/LLMFineTuningBench.tabulartabular-regression10K<n<100K3 likes147 downloads16d agoHugging Face18gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes125 downloads2y agoHugging Face19mznaser /moral-tracing-in-LLMs LLM Moral Evolution Study A longitudinal dataset tracking moral reasoning patterns across 14 large language models from OpenAI and Anthropic, spanning multiple generations (2023–2025). The dataset measures how moral stances, ethical judgments, and value priorities shift across model updates using a 107-item probe instrument grounded in Moral Foundations Theory. Models OpenAI Model Release GPT-3.5 Turbo 2023-11 GPT-4 2023-03 GPT-4o… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/moral-tracing-in-LLMs.documenttext-classification10K<n<100K0 likes124 downloads3mo agoHugging Face20tussiiiii /llm-classification-distilled-v2-sharded LLM Classification Distilled v2 Sharded Overview This repository stores shard CSV files produced by the teacher-judge distillation pipeline. How to Use Run the distillation notebook once per shard: NUM_SHARDS = 4 SHARD_INDEX = 0 .. 3 After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos. Final Repositories Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.tabulartext-classification100K<n<1M0 likes121 downloads4mo agoHugging Face21danilocorsi /LLMs-Sentiment-Augmented-Bitcoin-Dataset Leveraging LLMs for Informed Bitcoin Trading Decisions: Prompting with Social and News Data Reveals Promising Predictive Abilities The work was carried out by: Danilo Corsi Cesare Campagnano Description This project investigates the potential of leveraging Large Language Models (LLMs) to support Bitcoin traders. Specifically, we analyze the correlation between Bitcoin price movements and sentiment expressed in news headlines, posts, and comments on social media. We… See the full description on the dataset page: https://huggingface.co/datasets/danilocorsi/LLMs-Sentiment-Augmented-Bitcoin-Dataset.tabulartext-classification10K<n<100K8 likes111 downloads2y agoHugging Face22kishan51 /llm-affect-lab LLM Affect Lab This dataset contains the API-level results for LLM Affect Lab, a study of functional affect signatures in language model behavior. Functional Affect Score (FAS) is a 0-1 behavioral proxy. It combines generated-token confidence, enthusiastic language, consistency across repeated samples, forced self-report computed from digit top-logprob probabilities, and length control. The goal is not to claim that models feel emotions; the goal is to measure whether different… See the full description on the dataset page: https://huggingface.co/datasets/kishan51/llm-affect-lab.tabulartext-generation1K<n<10K0 likes110 downloads5mo agoHugging Face23llmbenchio /benchmarks-by-vramUpdated on: 21 Sep 2026 Data contains: runs from the last 30 days Minimum runs: model/hardware combos with fewer than 3 runs are excluded llm-bench.io — Community LLM Benchmark Leaderboard by Hardware Per-model community benchmark data for local LLMs, curated from llm-bench.io and grouped by hardware and available VRAM. This dataset contains only aggregated statistics derived from individual benchmark submissions. It does not contain raw submissions, prompts, model responses… See the full description on the dataset page: https://huggingface.co/datasets/llmbenchio/benchmarks-by-vram.tabularn<1K0 likes110 downloads2d agoHugging Face24DavidYor06 /llm-disagreement Lenz Frontier-LLM Disagreement — v1.1 Five frontier language models each rated the same 997 real fact-check claims submitted by users of Lenz. This dataset is the per-claim record of where they agreed and where they did not. In 63% of real-world fact-checks, top AI models don't agree on the answer — at least one model dissents from the majority, or no majority forms at all (95% CI 60–66%). At a glance Claims 997 complete (of 1,000 harvested) Models… See the full description on the dataset page: https://huggingface.co/datasets/DavidYor06/llm-disagreement.tabulartext-classification1K<n<10K1 likes109 downloads8d agoHugging Face25salem-mbzuai /LLM-BabyBench LLM-BabyBench Overview LLM-BabyBench is a benchmark suite designed to evaluate Large Language Models (LLMs) on grounded planning and reasoning tasks. Built upon a textual adaptation of the procedurally generated BabyAI grid world, this benchmark assesses the capabilities of LLMs to plan and reason within the constraints of interactive environments. LLM-BabyBench specifically evaluates three fundamental aspects of grounded intelligence: (1) predicting the consequences of… See the full description on the dataset page: https://huggingface.co/datasets/salem-mbzuai/LLM-BabyBench.tabular10K<n<100K4 likes105 downloads1y agoHugging Face26salttechno /LLM-Model-Comparison-2026 LLM Model Comparison 2026 Which LLM should you use for enterprise AI in 2026? This open dataset compares 16 large language models from 7 providers across 22 fields: pricing, benchmark scores, context windows, latency, API features, and recommended use cases. Published and maintained by Salt Technologies AI, the AI engineering division of Salt Technologies (14+ years, 800+ projects delivered). Quick Links Interactive dataset page:… See the full description on the dataset page: https://huggingface.co/datasets/salttechno/LLM-Model-Comparison-2026.tabularn<1K0 likes87 downloads7mo agoHugging Face27pixeloffice /llm-smartrouter-benchmark LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0) Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1). v1.4.0 Benchmark Highlights Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings. Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.tabulartext-generationn<1K0 likes76 downloads19d agoHugging Face28RaynarDM /apple-silicon-llm-benchmarks Apple Silicon Local LLM Benchmarks — M2 Max 32GB Measurements taken while trying to get Qwen3.8-27B usable locally on a 32GB M2 Max. Most of the popular speedup advice did not transfer from CUDA, so these are mostly negative results. Everything here was measured on one machine. Treat it as a datapoint, not a law. Hardware and software Chip Apple M2 Max Unified memory 32 GB (~21.8 GB wireable to the GPU) macOS 26.5.2 llama.cpp build c1d0e7a00… See the full description on the dataset page: https://huggingface.co/datasets/RaynarDM/apple-silicon-llm-benchmarks.tabularn<1K0 likes73 downloads1d agoHugging Face29tarekmasryo /llm-system-ops-production-telemetry-sft-data 🤖📈 LLM System Ops Telemetry (Synthetic) A synthetic, production-style, multi-table LLM telemetry dataset designed for LLMOps analytics and decision-grade experiments. It supports monitoring cost, latency, tokens, failures, safety flags, tool usage, and user feedback at the interaction level, with rollups at the session and user levels — plus an SFT table aligned 1:1 with interactions and a prompt/config dimension. Synthetic data (safe for teaching, prototyping, and portfolio… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/llm-system-ops-production-telemetry-sft-data.tabulartabular-classification10K<n<100K1 likes72 downloads8mo agoHugging Face30nbvbharath-1729 /llm-math-evaluation-dataset LLM Math Response Evaluation Dataset Dataset Summary A human-annotated dataset of 150 AI-generated math responses evaluated across GPT-4o, Claude, and Gemini. Each response is scored on Correctness, Reasoning, and Clarity using a structured rubric, with written justification for every score. Supported Tasks LLM evaluation and benchmarking Math reasoning quality assessment Error type classification in AI responses Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.tabularn<1K0 likes70 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.