datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.llmfit-benchmarks
llmfit Real-World LLM Inference Benchmarks
An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit.
The initial release contains 1,501 normalized observations:
1,010 unique external-community observations from the repository's 2026-08-10 snapshot.
491 llmfit-community observations from 61… See the full description on the dataset page: https://huggingface.co/datasets/axjns/llmfit-benchmarks.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.PCBSchemaGen-Benchmarks
PCBSchemaGen Benchmarks
Two benchmark suites for LLM-driven PCB schematic synthesis, from the paper
PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Board (PCB) Schematic Design with Structured Verification.
Correctness in this domain is not defined by unit tests: there are no per-task golden references,
and SPICE does not validate schematic-level correctness. Instead, each task is scored by a
deterministic structural verifier against real-IC pin- and… See the full description on the dataset page: https://huggingface.co/datasets/Hzou9/PCBSchemaGen-Benchmarks.llm-benchmarks-capabilities-2020-2026
📊 LLM Benchmarks & Capabilities 2020–2026
The most comprehensive open dataset tracking the evolution of Large Language Models — from GPT-3 to GPT-5.5, Claude Opus 4.7, Gemini 3.5, and beyond.
🧭 Overview
This dataset captures the complete LLM landscape from 2020 to 2026 across five dimensions:
🤖 113 models from 25+ organizations
📈 17 benchmarks tracking capability growth over time
💰 Monthly API pricing showing 100x+ cost reductions
⚙️ Training compute… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/llm-benchmarks-capabilities-2020-2026.TMT-Benchmarks
TMT-Benchmarks
Benchmark dataset for the TemporalMesh Transformer (TMT) paper.
Paper: TemporalMesh Transformer: Dynamic Graph Attention with Temporal Decay and Adaptive Depth RoutingDOI: 10.5281/zenodo.20287197Author: Vigneshwar LK
Key Result
Full TMT achieves PPL 29.4 vs Vanilla 42.1 -- a 30.2% perplexity reduction while using only 48% of the compute (2.1x efficiency gain).
Dataset Subsets
This dataset contains 5 subsets, each measuring a… See the full description on the dataset page: https://huggingface.co/datasets/vigneshwar234/TMT-Benchmarks.benchmarks
Dataset Card for Dataset Name
jailbreak analysis
Qwen3.6-27B-OTQ-GGUF-benchmarks
Qwen3.6-27B OTQ GGUF Benchmark Reproducibility
This dataset contains the compact paired benchmark evidence used by zlaabsi/Qwen3.6-27B-OTQ-GGUF.
It is a reproducibility dataset, not a leaderboard dataset. The rows are small practical release signals run on pinned task IDs with prompt format qwen3-no-think, deterministic decoding and local scoring rules.
Contents
Path
Meaning
data/paired_samples.jsonl
Flattened 232-row paired sample table with prompts, task… See the full description on the dataset page: https://huggingface.co/datasets/zlaabsi/Qwen3.6-27B-OTQ-GGUF-benchmarks.
