datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Evidence_Inference_v2
Evidence Inference 2.0
Dataset Description
Links
Homepage:
Github Pages
Repository:
Github
Paper:
arXiv
Contact (Original Authors):
Jay DeYoung (deyoung.j@northeastern.edu)
Contact (Curator):
Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
The dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple treatments. Each of these articles will have multiple questions, or 'prompts'… See the full description on the dataset page: https://huggingface.co/datasets/araag2/Evidence_Inference_v2.OB-Inference-Microtasks
Inference Microtasks
29 synthetic microtasks with reference answers across meeting-notes lookup,
support-ticket triage, and contract-terms extraction. The public set accompanies
the OpenBenchmarks Inference Benchmark,
which measures single-user delay on short, deliberately easy structured tasks.
Dataset contents
Configuration
Rows
Task
contract-terms-extraction
10
Extract commercial terms from a technology contract excerpt.
meeting-notes-lookup
13… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-Inference-Microtasks.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.
