datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
inference-benchmarkerfast-autoregressive-inference-gp-trainK4sovereign-shadow-inference-bench
Sovereign Shadow Inference Bench
A public, versioned evidence surface for independent Hugging Face shadow inference beside Sovereign's primary OpenRouter/Revolver route.
What this dataset proves
The seed record in data/shadow_receipts.jsonl was produced by one real Hugging Face Inference Providers request. It records provider/model identity, request bounds, latency, hashes, literal-match outcome, source revision, and an immutable receipt hash.
What it… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-shadow-inference-bench.inference-results
LEGEX Inference Results
System outputs for the LEGEX benchmark, six review-table extraction
runs (four systems; Harvey and Legora twice each) on the case packets of
all 19 jurisdictions:
am, au, be, br, ch, de, es, fr, ge, hk, in, np, nz, ph, rs, sg, tw, uk, us
The systems
Abbreviation
model field value
Notes
harvey
harvey
Harvey Vault Review, a commercial review-table product. Outputs exported from the production system on 2026-05-18 and 2026-06-30… See the full description on the dataset page: https://huggingface.co/datasets/legexbenchmark/inference-results.speculators_benchmarks_tool_callQwen3.5-0.8B-responsesqwen3.8-27b-inference-benchmark-4090
Qwen3.8-27B Inference Benchmark on RTX 4090 48GB
中文说明 · GitHub benchmark repository
Structured performance and accuracy results for four real Qwen3.8-27B serving configurations on an NVIDIA RTX 4090 48 GB workstation. A dual-GPU llama.cpp BF16 reference additionally used an RTX 3090 24 GB.
This dataset is the analysis-friendly companion to the full benchmark repository. It publishes aggregate tables, 140 normalized per-request performance records, accuracy scores, sanitized… See the full description on the dataset page: https://huggingface.co/datasets/pxzleo/qwen3.8-27b-inference-benchmark-4090.cs2-action-inference-test
CS2 战术 Action 推理测试集
本测试集用于 WAN I2V 的战术动作定性测试。每个小类只保留 1 张真实比赛 POV 第一帧,以及两种英文文本条件;本版不提供 GT 视频。第一帧来源依据 parse-dem 的 events.csv、game_events.csv 或逐 tick 状态对齐到 opencs2_matches* 视频。
数据约定
共 45 个 case、9 个大类。
每个 case 只有一张 832x480 的 first_frame.png,作为 WAN I2V 条件图;不裁剪或复制 GT clip。首帧优先选择正常持械、水平视角、无遮挡且较开阔的画面。
prompt.txt 是完整英文 prompt,包含首帧可见环境、初始持械状态、画面保持要求和整段唯一动作变化。
chunk_prompts.json 固定包含 5 个英文 prompt,依次描述期望生成视频的 0-1、1-2、2-3、3-4、4-5 秒。
metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-action-inference-test.OB-Inference-Microtasks
Inference Microtasks
29 synthetic microtasks with reference answers across meeting-notes lookup,
support-ticket triage, and contract-terms extraction. The public set accompanies
the OpenBenchmarks Inference Benchmark,
which measures single-user delay on short, deliberately easy structured tasks.
Dataset contents
Configuration
Rows
Task
contract-terms-extraction
10
Extract commercial terms from a technology contract excerpt.
meeting-notes-lookup
13… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-Inference-Microtasks.HALO-Gemini-3-Flash-AppWorld
Dataset Card: Gemini 3 Flash Traces on AppWorld (test-normal)
Dataset Overview
This dataset contains agent execution traces of Gemini 3 Flash running on the AppWorld benchmark, specifically evaluated on the test-normal dataset split. The traces capture the full span-level execution detail of the model interacting with AppWorld's simulated app ecosystem.
Field
Value
Model
Gemini 3 Flash
Benchmark
AppWorld
Split
test-normal
Total Traces
168
Total Spans
3… See the full description on the dataset page: https://huggingface.co/datasets/inference-net/HALO-Gemini-3-Flash-AppWorld.Qwen3.5-4B-responsesfast-autoregressive-inference-gp-trainK16inference-audit
Inference Audit: Provider Delivery and Metering
This dataset contains 3,932 controlled observations from OpenAI-compatible endpoints
serving openai/gpt-oss-120b through 18 pinned providers. The runs measure what an API returned
and reported at the HTTP boundary: delivery, parameter compliance, token accounting, caching,
streaming behavior, latency, and repeatability.
The records do not identify a model from its outputs, prove billing fraud, or establish why
two endpoints differ.… See the full description on the dataset page: https://huggingface.co/datasets/nuckcrews/inference-audit.RetroDFM-R-inferenceMultiling_inference_8lang_en2xx_xx2en_dataLongbench_Samples_Specdecai-inference-2026
AI Inference & Compute 2026
AI inference infrastructure, compute costs. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c =… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-inference-2026.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.fast-autoregressive-inference-scm-train5gbarxiv-sample-affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air-inference-results-enriched
affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air arXiv author affiliation inference results
Author names and institutional affiliations extracted from arXiv preprints with the affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air LoRA, enriched with ROR identifiers.
Dataset Structure
Each record contains the following fields:
Field
Type
Description
doi
string
DOI for the preprint
title
string
Preprint title
arxiv_id
stringarXiv identifier… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-sample-affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air-inference-results-enriched.dataset_inference
Dataset Inference V2: Detect Datasets, Not Strings
This repository contains data from 22 different domains of the PILE, divided into train and val sets. The data is in the form of a JSON file, with each entry containing the raw text, as well as various kinds of perturbations applied to it. The dataset is used to facilitate privacy research in language models, where the perturbed data can be used as reference detect the presence of a particular dataset in the training data of a… See the full description on the dataset page: https://huggingface.co/datasets/bxiong/dataset_inference.kimi-cyber-reasoning
Kimi Cyber Reasoning
997 chain-of-thought records covering 13 cybersecurity disciplines and 4 systems engineering domains, distilled from the Kimi K3 reasoning model via API. Every record provides an explicit step-by-step <think> reasoning trace followed by a technical resolution, unified code diff fix, or structured tool invocation.
The dataset was curated as an anchor set for training, healing, and specializing compact reasoning models on systems security and tool calling… See the full description on the dataset page: https://huggingface.co/datasets/inferenceport-ai/kimi-cyber-reasoning.SearchAgentDemoTracesdataset-viber-image-generation-preference-inference-endpoints-battle-flux
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.arxiv-author-affiliation-extraction-inference-inputs-metadatafast-autoregressive-inference-eeglaguna-xs-ultrachat-responsesrepro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Qwen3.5-9B-responsesevery-eval-ever-demo
