datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
auditkit-testrun-factual-consistency
auditkit-testrun-factual-consistency
Built using AuditKIT — evaluate any model on any dataset and any task.
Method
evaluate
Model
<auditkit.model.vllm_gen.VLLMModel object at 0x7c1b15bf5010>
Artifact
run
Published
2026-09-01 14:24 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/auditkit-testrun-factual-consistency")
factual-multiagent-roleplay-ft-ru
march228/factual-multiagent-roleplay-ft-ru
Небольшой русскоязычный synthetic finetuning dataset для обучения модели следованию ролевым системным инструкциям личности при сохранении фактической опоры на контекст.
Что это за датасет
Этот набор сделан как instruction / finetuning dataset, а не как benchmark.
В каждой записи есть:
плотный system с персоной и тоном;
context, на который нужно опираться;
пользовательский question;
внутренние thoughts;
финальный answer.… See the full description on the dataset page: https://huggingface.co/datasets/march228/factual-multiagent-roleplay-ft-ru.how_lms_answer_one_to_many_factual_queries
One-to-Many Factual Queries Datasets
This is the official dataset used in our EMNLP 2025 paper Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries.
The dataset includes six subsets named {dataset_name}_template_{i}, where dataset_name is country_cities, artist_songs, or actor_movies, and each dataset has three prompt templates (i = 1, 2, 3).
The {model_name}_step_{i} split in each subset contains the data used for analyzing model_name's behavior at… See the full description on the dataset page: https://huggingface.co/datasets/LorenaYannnnn/how_lms_answer_one_to_many_factual_queries.google-gemma-2-9b-it__llm-quality-factual-mini__019e3b975288
google/gemma-2-9b-it on llm.quality.factual-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
10
N Ok
10
Ok Rate
1
Accuracy
1
Accuracy P05
1
Accuracy P50
1
Accuracy P95
1
TTFT P50
16.8943
ms
Total P50 Ms
142.032
Tokens Out Total
205
Run configuration
Model: google/gemma-2-9b-it @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100 80GB HBM3
Driver:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/google-gemma-2-9b-it__llm-quality-factual-mini__019e3b975288.qwen-qwen2-vl-7b-instruct__llm-quality-factual-mini__019e3b85fc73
Qwen/Qwen2-VL-7B-Instruct on llm.quality.factual-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
10
N Ok
10
Ok Rate
1
Accuracy
1
Accuracy P05
1
Accuracy P50
1
Accuracy P95
1
TTFT P50
29.4403
ms
Total P50 Ms
176.794
Tokens Out Total
150
Run configuration
Model: Qwen/Qwen2-VL-7B-Instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100 80GB HBM3… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-vl-7b-instruct__llm-quality-factual-mini__019e3b85fc73.deepseek-ai-deepseek-coder-v2-lite-instruct__llm-quality-factual-mini__019e3b6f9bdd
deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on llm.quality.factual-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
10
N Ok
10
Ok Rate
1
Accuracy
1
Accuracy P05
1
Accuracy P50
1
Accuracy P95
1
TTFT P50
45.0118
ms
Total P50 Ms
359.225
Tokens Out Total
373
Run configuration
Model: deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct @ unknown00
Engine: vllm vunknownQuantization:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__llm-quality-factual-mini__019e3b6f9bdd.qwen-qwen2-5-coder-7b-instruct__llm-quality-factual-mini__019e3b3fbe50
Qwen/Qwen2.5-Coder-7B-Instruct on llm.quality.factual-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
10
N Ok
10
Ok Rate
1
Accuracy
1
Accuracy P05
1
Accuracy P50
1
Accuracy P95
1
TTFT P50
13.9553
ms
Total P50 Ms
90.957
Tokens Out Total
217
Run configuration
Model: Qwen/Qwen2.5-Coder-7B-Instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-5-coder-7b-instruct__llm-quality-factual-mini__019e3b3fbe50.qwen-qwen2-5-7b-instruct__llm-quality-factual-mini__019e3b397723
Qwen/Qwen2.5-7B-Instruct on llm.quality.factual-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
10
N Ok
10
Ok Rate
1
Accuracy
1
Accuracy P05
1
Accuracy P50
1
Accuracy P95
1
TTFT P50
14.1209
ms
Total P50 Ms
328.7876
Tokens Out Total
421
Run configuration
Model: Qwen/Qwen2.5-7B-Instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100 80GB HBM3… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-5-7b-instruct__llm-quality-factual-mini__019e3b397723.meta-llama-llama-3-1-8b-instruct__llm-quality-factual-mini__019e3b300ded
meta-llama/Llama-3.1-8B-Instruct on llm.quality.factual-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
10
N Ok
10
Ok Rate
1
Accuracy
1
Accuracy P05
1
Accuracy P50
1
Accuracy P95
1
TTFT P50
14.0798
ms
Total P50 Ms
101.8105
Tokens Out Total
219
Run configuration
Model: meta-llama/Llama-3.1-8B-Instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-8b-instruct__llm-quality-factual-mini__019e3b300ded.meta-llama-llama-3-1-70b-instruct__llm-quality-factual-mini__019e3b922429
meta-llama/Llama-3.1-70B-Instruct on llm.quality.factual-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
10
N Ok
10
Ok Rate
1
Accuracy
1
Accuracy P05
1
Accuracy P50
1
Accuracy P95
1
TTFT P50
32.6144
ms
Total P50 Ms
435.9529
Tokens Out Total
267
Run configuration
Model: meta-llama/Llama-3.1-70B-Instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-70b-instruct__llm-quality-factual-mini__019e3b922429.mistralai-mistral-7b-instruct-v0-3__llm-quality-factual-mini__019e3b45a60d
mistralai/Mistral-7B-Instruct-v0.3 on llm.quality.factual-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
10
N Ok
10
Ok Rate
1
Accuracy
1
Accuracy P05
1
Accuracy P50
1
Accuracy P95
1
TTFT P50
11.7003
ms
Total P50 Ms
473.4047
Tokens Out Total
627
Run configuration
Model: mistralai/Mistral-7B-Instruct-v0.3 @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/mistralai-mistral-7b-instruct-v0-3__llm-quality-factual-mini__019e3b45a60d.microsoft-phi-3-5-mini-instruct__llm-quality-factual-mini__019e3b5d1176
microsoft/Phi-3.5-mini-instruct on llm.quality.factual-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
10
N Ok
10
Ok Rate
1
Accuracy
1
Accuracy P05
1
Accuracy P50
1
Accuracy P95
1
TTFT P50
11.2101
ms
Total P50 Ms
499.6194
Tokens Out Total
943
Run configuration
Model: microsoft/Phi-3.5-mini-instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/microsoft-phi-3-5-mini-instruct__llm-quality-factual-mini__019e3b5d1176.
