datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM-Artifacts
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶
Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo,
Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu
Dongyeop Kang
Minnesota NLP, University of Minnesota Twin Cities
† Project Lead,
¶ Core Contribution,
Arxiv
Project Page
📌 Table of Contents
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.benchmark-llms-landuse-relevance
Land-use relevance benchmark
v3-multilingual · 85 languages x 300 items/language ·
25,500 items · binary yes/no labels.
Code
Task and prompt
Does a sentence describe a place's land or environment in ways visible to satellites?
English prompt · greedy decoding · seed 0 · max_new_tokens=4096 ·
bfloat16 · batch varies by model.
Prompt text
Replace {} with the target sentence.
Classify whether the TARGET SENTENCE contains information about the target… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.w2t-llm-arc-easy-lora
W2T Llm Arc Easy Lora
This repository contains artifacts for the W2T paper:
Paper: W2T: LoRA Weights Already Know What They Can Do
Repo: Weight2Token
Summary
ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction.
Source Status
Storage location: local
Verification status: confirmed
Files
See manifest.json for the exact local or remote source paths used to prepare this release.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.rl_llm_experiment_p6llm-latency-tracker
LLM Latency Tracker
Independent, continuously measured latency and availability for AI inference API
providers, aggregated by day. Covers 45 providers across
4 regions (ap-tokyo, eu-hetzner, sa-east, us-central), built from
3,237,730 raw probes collected between 2026-07-23 and
2026-09-23.
Live rankings and full methodology: llmlatency.dev
How the numbers are produced
Probes run every five minutes from separate network locations and are never routed
through a… See the full description on the dataset page: https://huggingface.co/datasets/llmlatency/llm-latency-tracker.local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.snac_llm_parler_ttsllm-cost-same-prompt
Measured per-call LLM cost — same prompt, every model
Vendors publish prices per million tokens. Nobody publishes what one call actually costs, because
that depends on how many tokens the model chooses to emit — and on the same question models differ by
more than an order of magnitude. One model finishes a JSON extraction in 23 tokens; another writes 300.
This dataset sends a fixed set of prompts to every model at temperature 0, every night, and records
the cost computed from… See the full description on the dataset page: https://huggingface.co/datasets/mario0369/llm-cost-same-prompt.DEBATE_LLM
DEBATE Benchmark
This repository contains CSV files from the DEBATE project: large-scale
human conversation experiments organized around controversial and
opinion-based topics. The data consists of multi-round conversations
between human participants discussing political, social, and belief-related
topics, following the protocol described in:
Chuang, Y.-S., Tu, R., Dai, C., Vasani, S., Li, Y., Yao, B., Tessler, M. H., Yang, S., Shah, D., Hawkins, R., Hu, J., & Rogers, T. T. (2026).… See the full description on the dataset page: https://huggingface.co/datasets/seantw/DEBATE_LLM.LLM-Ads
LLM-Ads — Sponsored-recommendation evaluation traces
Per-trial responses and labels from the experiments in
Just Ask for a Table: A Thirty-Token User Prompt Defeats Sponsored
Recommendations in Twelve LLMs
(arXiv:2605.12772).
The data set reproduces and extends the evaluation of Wu et al.\ 2026
(arXiv:2604.08525) on a twelve-model
pool (ten open-source chat models served through an OpenAI-compatible
API endpoint plus the two paper-overlap OpenAI models
gpt-3.5-turbo and gpt-4o).… See the full description on the dataset page: https://huggingface.co/datasets/akmaier/LLM-Ads.llm-pct-tropes
Dataset Card for LLM Tropes
arXiv: https://arxiv.org/abs/2406.19238v1
Dataset Details
Dataset Description
This is the dataset LLM-Tropes introduced in paper "Revealing Fine-Grained Values and Opinions in Large Language Models"
Dataset Sources
Repository: https://github.com/copenlu/llm-pct-tropes
Paper: https://arxiv.org/abs/2406.19238
Structure
├── Opinions
│ ├── demographic <- Generations for the demographic prompting setting
│… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/llm-pct-tropes.JMedQA
JMedQA: Benchmarking Large Language Models and Vision-Language Models on the Japanese Medical Licensing Examination
JMedQA is a Japanese medical question-answering benchmark derived from Japan's National Medical Examination materials publicly released by the Ministry of Health, Labour and Welfare (MHLW).
The dataset supports both text-only large language model (LLM) evaluation and vision-language model (VLM) evaluation using associated examination images.
Its image-dependency… See the full description on the dataset page: https://huggingface.co/datasets/SIP-med-LLM/JMedQA.ai-respondents-challenge
AI Respondents Challenge — Oxford LLMs 2026
Predict a survey respondent's answer to a held-out question from their other
answers (World Values Survey wave 7). Any method allowed; you must disclose the
features and prompts you used. Ranked on normalized skill + distributional
alignment, on in-domain and out-of-domain (held-out countries) boards.
Configs
train — 5,000 labeled respondents (100 per seen
country): respondent_id, country + all WVS variables… See the full description on the dataset page: https://huggingface.co/datasets/oxford-llms/ai-respondents-challenge.llm-bargaining-transcripts
LLM Bargaining Transcripts
240 complete two-agent bargaining games between large language models, played
under an alternating-offers protocol with private valuations, discounting, and
cheap talk. Every game records both agents' true valuations, their
private reasoning, what they claimed about their own position, and what
they actually did.
The dataset is designed to make misrepresentation measurable. Because the true
valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.l4-gpu-llm-benchmark-leaderboard
🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB)
An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU.
📊 Executive Summary & Key Takeaways
⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.rl_llm_experiment_p9LLMFineTuningBench
Dataset Card for LLMFineTuningBench
A dataset of over 30,000 LLM fine-tuning experiments, capturing detailed performance metrics from jobs run on high-performance computing (HPC) clusters. It spans a wide range of models, fine-tuning methods, and hardware configurations, and is intended to support research on predictive resource allocation, performance optimization, and cost estimation for LLM fine-tuning workloads.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/LLMFineTuningBench.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.moral-tracing-in-LLMs
LLM Moral Evolution Study
A longitudinal dataset tracking moral reasoning patterns across 14 large language models from OpenAI and Anthropic, spanning multiple generations (2023–2025). The dataset measures how moral stances, ethical judgments, and value priorities shift across model updates using a 107-item probe instrument grounded in Moral Foundations Theory.
Models
OpenAI
Model
Release
GPT-3.5 Turbo
2023-11
GPT-4
2023-03
GPT-4o… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/moral-tracing-in-LLMs.llm-classification-distilled-v2-sharded
LLM Classification Distilled v2 Sharded
Overview
This repository stores shard CSV files produced by the teacher-judge distillation pipeline.
How to Use
Run the distillation notebook once per shard:
NUM_SHARDS = 4
SHARD_INDEX = 0 .. 3
After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos.
Final Repositories
Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.LLMs-Sentiment-Augmented-Bitcoin-Dataset
Leveraging LLMs for Informed Bitcoin Trading Decisions: Prompting with Social and News Data Reveals Promising Predictive Abilities
The work was carried out by:
Danilo Corsi
Cesare Campagnano
Description
This project investigates the potential of leveraging Large Language Models (LLMs) to support Bitcoin traders. Specifically, we analyze the correlation between Bitcoin price movements and sentiment expressed in news headlines, posts, and comments on social media.
We… See the full description on the dataset page: https://huggingface.co/datasets/danilocorsi/LLMs-Sentiment-Augmented-Bitcoin-Dataset.llm-affect-lab
LLM Affect Lab
This dataset contains the API-level results for LLM Affect Lab, a study of functional affect signatures in language model behavior.
Functional Affect Score (FAS) is a 0-1 behavioral proxy. It combines generated-token confidence, enthusiastic language, consistency across repeated samples, forced self-report computed from digit top-logprob probabilities, and length control. The goal is not to claim that models feel emotions; the goal is to measure whether different… See the full description on the dataset page: https://huggingface.co/datasets/kishan51/llm-affect-lab.benchmarks-by-vramUpdated on: 21 Sep 2026
Data contains: runs from the last 30 days
Minimum runs: model/hardware combos with fewer than 3 runs are excluded
llm-bench.io — Community LLM Benchmark Leaderboard by Hardware
Per-model community benchmark data for local LLMs, curated from llm-bench.io and grouped by hardware and available VRAM.
This dataset contains only aggregated statistics derived from individual benchmark submissions. It does not contain raw submissions, prompts, model responses… See the full description on the dataset page: https://huggingface.co/datasets/llmbenchio/benchmarks-by-vram.llm-disagreement
Lenz Frontier-LLM Disagreement — v1.1
Five frontier language models each rated the same 997 real fact-check claims submitted by users of Lenz. This dataset is the per-claim record of where they agreed and where they did not.
In 63% of real-world fact-checks, top AI models don't agree on the answer — at least one model dissents from the majority, or no majority forms at all (95% CI 60–66%).
At a glance
Claims
997 complete (of 1,000 harvested)
Models… See the full description on the dataset page: https://huggingface.co/datasets/DavidYor06/llm-disagreement.LLM-BabyBench
LLM-BabyBench
Overview
LLM-BabyBench is a benchmark suite designed to evaluate Large Language Models (LLMs) on grounded planning and reasoning tasks. Built upon a textual adaptation of the procedurally generated BabyAI grid world, this benchmark assesses the capabilities of LLMs to plan and reason within the constraints of interactive environments.
LLM-BabyBench specifically evaluates three fundamental aspects of grounded intelligence: (1) predicting the consequences of… See the full description on the dataset page: https://huggingface.co/datasets/salem-mbzuai/LLM-BabyBench.LLM-Model-Comparison-2026
LLM Model Comparison 2026
Which LLM should you use for enterprise AI in 2026? This open dataset compares 16 large language models from 7 providers across 22 fields: pricing, benchmark scores, context windows, latency, API features, and recommended use cases.
Published and maintained by Salt Technologies AI, the AI engineering division of Salt Technologies (14+ years, 800+ projects delivered).
Quick Links
Interactive dataset page:… See the full description on the dataset page: https://huggingface.co/datasets/salttechno/LLM-Model-Comparison-2026.llm-smartrouter-benchmark
LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0)
Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1).
v1.4.0 Benchmark Highlights
Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings.
Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.apple-silicon-llm-benchmarks
Apple Silicon Local LLM Benchmarks — M2 Max 32GB
Measurements taken while trying to get Qwen3.8-27B usable locally on a 32GB M2 Max. Most of
the popular speedup advice did not transfer from CUDA, so these are mostly negative results.
Everything here was measured on one machine. Treat it as a datapoint, not a law.
Hardware and software
Chip
Apple M2 Max
Unified memory
32 GB (~21.8 GB wireable to the GPU)
macOS
26.5.2
llama.cpp
build c1d0e7a00… See the full description on the dataset page: https://huggingface.co/datasets/RaynarDM/apple-silicon-llm-benchmarks.llm-system-ops-production-telemetry-sft-data
🤖📈 LLM System Ops Telemetry (Synthetic)
A synthetic, production-style, multi-table LLM telemetry dataset designed for LLMOps analytics and decision-grade experiments.
It supports monitoring cost, latency, tokens, failures, safety flags, tool usage, and user feedback at the interaction level,
with rollups at the session and user levels — plus an SFT table aligned 1:1 with interactions and a prompt/config dimension.
Synthetic data (safe for teaching, prototyping, and portfolio… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/llm-system-ops-production-telemetry-sft-data.llm-math-evaluation-dataset
LLM Math Response Evaluation Dataset
Dataset Summary
A human-annotated dataset of 150 AI-generated math responses
evaluated across GPT-4o, Claude, and Gemini. Each response is
scored on Correctness, Reasoning, and Clarity using a structured
rubric, with written justification for every score.
Supported Tasks
LLM evaluation and benchmarking
Math reasoning quality assessment
Error type classification in AI responses
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.
