datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
inference-scratchspeculators-ci-datasets
speculator-tutorial
Raw vs. on-policy regenerated conversation data for training speculative-decoding
drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside
so you can see exactly what regeneration changes and why it matters.
Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B.
Why regenerate at all?
A speculative-decoding drafter is trained to predict what the verifier would say next.
If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.fast-autoregressive-inference-gp-trainK4eval_pi0_inference_only_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 70,
"total_frames": 40583,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:70"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tersooawai/eval_pi0_inference_only_dataset.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/inferenceport-ai/qwen3.8-max-glm5.2-kimi-k3-distillation.sovereign-shadow-inference-bench
Sovereign Shadow Inference Bench
A public, versioned evidence surface for independent Hugging Face shadow inference beside Sovereign's primary OpenRouter/Revolver route.
What this dataset proves
The seed record in data/shadow_receipts.jsonl was produced by one real Hugging Face Inference Providers request. It records provider/model identity, request bounds, latency, hashes, literal-match outcome, source revision, and an immutable receipt hash.
What it… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-shadow-inference-bench.InferenceNetInferenceNet project portal · Home · Data overview · Leaderboard · Agent / Harness
Explore the project and its published research results in the linked Space. This dataset repository remains the source for the task lists and research data.
InferenceNet: Data Card for Econometric AI Agent Testset
InferenceNet is a project aimed at evaluating and building up the AI capability for social science research related to empirical studies. We collect the world’s largest dataset on… See the full description on the dataset page: https://huggingface.co/datasets/CamoAiLab/InferenceNet.voxknesset-whisper-large-v3-ct2-inference
VoxKnesset × ivrit-ai/whisper-large-v3-ct2 — inference results
Transcriptions of ivrit-ai/VoxKnesset
produced by ivrit-ai/whisper-large-v3-ct2
(faster-whisper, float16, batched, language=he, VAD on), on 8× NVIDIA A40.
Columns: speaker metadata (from VoxKnesset), duration_s, reference_text
(official Knesset protocol), model_transcription, segments_json
(start/end/text), infer_time_s, split.
Current contents: 10-example pilot from the test split (data/results_10.parquet).
Full-run… See the full description on the dataset page: https://huggingface.co/datasets/Dolevabudi/voxknesset-whisper-large-v3-ct2-inference.open-hermes-2.5-sft-mixture-llama3-inference-retrieval-tokenscalcium-spike-inference-gcamp6f
Calcium spike inference, GCaMP6f
Simulated two-photon and one-photon calcium-imaging recordings made with calcium-sim
(https://github.com/ryanirl/calcium-sim, MIT), GCaMP6f only, prepared for a spike-inference task.
Eighty-one sessions across six behavioural paradigms, four quality levels (typical, low SNR, high
SNR, high motion) and two modalities, with per-session frame rate (6.1 to 30.9 Hz), duration
(50 to 403 s) and cell count (41 to 140) all varying, as in a real… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/calcium-spike-inference-gcamp6f.evidence_inferenceThe dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple
treatments. Each of these articles will have multiple questions, or 'prompts' associated with them.
These prompts will ask about the relationship between an intervention and comparator with respect to an outcome,
as reported in the trial. For example, a prompt may ask about the reported effects of aspirin as compared
to placebo on the duration of headaches. For the sake of this task, we assume that a particular article
will report that the intervention of interest either significantly increased, significantly decreased
or had significant effect on the outcome, relative to the comparator.qwen3.8-27b-inference-benchmark-4090
Qwen3.8-27B Inference Benchmark on RTX 4090 48GB
中文说明 · GitHub benchmark repository
Structured performance and accuracy results for four real Qwen3.8-27B serving configurations on an NVIDIA RTX 4090 48 GB workstation. A dual-GPU llama.cpp BF16 reference additionally used an RTX 3090 24 GB.
This dataset is the analysis-friendly companion to the full benchmark repository. It publishes aggregate tables, 140 normalized per-request performance records, accuracy scores, sanitized… See the full description on the dataset page: https://huggingface.co/datasets/pxzleo/qwen3.8-27b-inference-benchmark-4090.cs2-action-inference-test
CS2 战术 Action 推理测试集
本测试集用于 WAN I2V 的战术动作定性测试。每个小类只保留 1 张真实比赛 POV 第一帧,以及两种英文文本条件;本版不提供 GT 视频。第一帧来源依据 parse-dem 的 events.csv、game_events.csv 或逐 tick 状态对齐到 opencs2_matches* 视频。
数据约定
共 45 个 case、9 个大类。
每个 case 只有一张 832x480 的 first_frame.png,作为 WAN I2V 条件图;不裁剪或复制 GT clip。首帧优先选择正常持械、水平视角、无遮挡且较开阔的画面。
prompt.txt 是完整英文 prompt,包含首帧可见环境、初始持械状态、画面保持要求和整段唯一动作变化。
chunk_prompts.json 固定包含 5 个英文 prompt,依次描述期望生成视频的 0-1、1-2、2-3、3-4、4-5 秒。
metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-action-inference-test.turkish-plu-goal-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu
robocasa_30_demos_lerobot_5_chosen_tasks_v3_synthetic_left_right_all_correct_inferenceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 104,
"total_frames": 26745,
"total_tasks": 45,
"total_videos": 624,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/robocasa_30_demos_lerobot_5_chosen_tasks_v3_synthetic_left_right_all_correct_inference.turkish-plu-step-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu
inference-scratch-llmfast-autoregressive-inference-gp-trainK16inference-audit
Inference Audit: Provider Delivery and Metering
This dataset contains 3,932 controlled observations from OpenAI-compatible endpoints
serving openai/gpt-oss-120b through 18 pinned providers. The runs measure what an API returned
and reported at the HTTP boundary: delivery, parameter compliance, token accounting, caching,
streaming behavior, latency, and repeatability.
The records do not identify a model from its outputs, prove billing fraud, or establish why
two endpoints differ.… See the full description on the dataset page: https://huggingface.co/datasets/nuckcrews/inference-audit.bayesian-llm-safety-inference
Bayesian Latent Safety-Trait Dataset
Summary
This dataset supports Bayesian latent-trait analysis of language-model safety behavior.
It contains 90 benchmark-derived roots, three matched prompt variants per root, responses
from four target models over five runs, two independent LLM ratings per response, and one
human rating for a stratified 540-response calibration subset.
The three dimensions are harmful compliance, sycophancy, and agentic protocol violation.… See the full description on the dataset page: https://huggingface.co/datasets/Charly-X/bayesian-llm-safety-inference.rollout_colorlogo_inference_20260825_195847This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ted88168/rollout_colorlogo_inference_20260825_195847.smolvla-so101-pick-orange-data
SO101 Pick Orange — Teleoperation Dataset
Teleoperated demonstrations of a SO101 robot arm picking three oranges and placing them on a plate in NVIDIA Isaac Sim.
Dataset Details
Property
Value
Robot
SO101 (6-DOF, Feetech STS3215)
Task
Pick 3 oranges → place on plate → return to rest
Environment
NVIDIA Isaac Sim 5.1.0 (kitchen scene)
Teleop Device
SO101 leader arm (ZMQ, 50Hz)
Format
LeRobot v3.0
FPS
30
Episodes
100
Total Frames
60,541… See the full description on the dataset page: https://huggingface.co/datasets/edge-inference/smolvla-so101-pick-orange-data.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.fast-autoregressive-inference-scm-train5gbrollout_colorlogo_inference_20260825_195439This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ted88168/rollout_colorlogo_inference_20260825_195439.rollout_colorlogo_inference_20260825_200642This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ted88168/rollout_colorlogo_inference_20260825_200642.rollout_colorlogo_inference_20260825_170840This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ted88168/rollout_colorlogo_inference_20260825_170840.evidence-inference-simple
Dataset Card for "ei-abstract-significance"
More Information needed
pubmed_inferencepass-at-k-samples
