datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.runux-tpu-v5e-benchmarks
⚡ RunuX-AI — TPU v5e Inference Benchmarks
Achieving 3× Throughput & 3× Energy Reduction on Google TPU v5e
Xavier Callens · Socrate AI Lab (Non-Profit)
Reproducible benchmark data & scripts — No proprietary code included
🎯 What Is This?
This repository contains benchmark results and Apache-2.0 reproduction scripts for comparing LLM inference performance across 5 frameworks on Google TPU v5e. The goal is to enable independent verification of our claims… See the full description on the dataset page: https://huggingface.co/datasets/callensxavier/runux-tpu-v5e-benchmarks.PCBSchemaGen-Benchmarks
PCBSchemaGen Benchmarks
Two benchmark suites for LLM-driven PCB schematic synthesis, from the paper
PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Board (PCB) Schematic Design with Structured Verification.
Correctness in this domain is not defined by unit tests: there are no per-task golden references,
and SPICE does not validate schematic-level correctness. Instead, each task is scored by a
deterministic structural verifier against real-IC pin- and… See the full description on the dataset page: https://huggingface.co/datasets/Hzou9/PCBSchemaGen-Benchmarks.packrat-benchmarks
PackRat v2 Benchmarks
Version: 2.0.0
Date: 2026-04-10
Tokenizer: tiktoken cl100k_base (GPT-4 / Claude compatible)
Platform: Node.js v25.6.1, Windows 11
Summary
Metric
Result
Round-trip accuracy
100% (144/144 tests)
Token savings (avg)
2.4%
Token savings (best)
17.3% (path/URL-heavy files)
Byte savings (avg)
2.5%
Search speedup
12.03x
Codebook entries
72 (auto-learned)
Negative-savings entries
0
Comparison: PackRat vs MemPalace… See the full description on the dataset page: https://huggingface.co/datasets/kevo666/packrat-benchmarks.PIPE-Cypher-benchmarks
PIPE-Cypher Benchmarks
This dataset release contains the public-proxy benchmark exports used by
PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher
Systems.
Paper: https://arxiv.org/abs/2606.08481
Repository: https://github.com/suraj-ranganath/PIPE-Cypher/
Dataset repo: https://huggingface.co/datasets/suraj-ranganath/PIPE-Cypher-benchmarks
Exports
finbench_snb_full_qwen9b
Total examples: 3000
Split counts: {"dev": 296, "test":… See the full description on the dataset page: https://huggingface.co/datasets/suraj-ranganath/PIPE-Cypher-benchmarks.Qwen3.6-27B-OTQ-GGUF-benchmarks
Qwen3.6-27B OTQ GGUF Benchmark Reproducibility
This dataset contains the compact paired benchmark evidence used by zlaabsi/Qwen3.6-27B-OTQ-GGUF.
It is a reproducibility dataset, not a leaderboard dataset. The rows are small practical release signals run on pinned task IDs with prompt format qwen3-no-think, deterministic decoding and local scoring rules.
Contents
Path
Meaning
data/paired_samples.jsonl
Flattened 232-row paired sample table with prompts, task… See the full description on the dataset page: https://huggingface.co/datasets/zlaabsi/Qwen3.6-27B-OTQ-GGUF-benchmarks.
