datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vLLM-SR-Preference-V1The files in this repo is the LLM-labeled samples that are used as the training dataset for vLLM-SR Preference model V1.
The training file (sharegpt_preference_labeld_with_negative.jsonl) contains 25k records that have sample_id, golden label for the preference-based routing policy, and a set of negative labels that are plausible but do not match the conversation context.
The validation file has the same structure, but only 1% of the training file size. The validation file and the training… See the full description on the dataset page: https://huggingface.co/datasets/ppppqp/vLLM-SR-Preference-V1.POVID_preference_data_for_VLLMsvllm-benchmark-payloads
vLLM Benchmark Payloads
Synthetic OpenAI-format chat-completion payloads for latency / throughput benchmarking
of a model served with vLLM's OpenAI-compatible server.
File
File
Records
Notes
payloads_1k_generic.jsonl
1,000
The benchmark dataset — one request body per line. Each has a ~12k-token system prompt + a dynamic per-record customer context.
sample_payload.json
1
One pretty-printed record, to inspect the format quickly.
All data is 100%… See the full description on the dataset page: https://huggingface.co/datasets/rohitjain28/vllm-benchmark-payloads.deepscaler-teacher-sft-vllm-official-40k
DeepScaleR teacher SFT vLLM official 40k
Generated run: exp_003_vllm_official_brainlab_2gpu.
Summary
{
"num_examples": 40300,
"sft_dir": "data/processed/deepscaler/teacher_sft/exp_003_vllm_official_brainlab_2gpu",
"parse_rate": 0.9999751861042183,
"correct_rate": 0.5728039702233251,
"format_rate": 0.005955334987593052,
"mean_reward": 0.42432258064534184,
"deepscaler_mean_reward": 0.6266997518610422,
"deepscaler_match_mean_reward":… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.chart-vllm-ver1deepscaler-teacher-sft-vllm-official-40k-clean-v2
DeepScaleR Teacher40k Clean v2
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
minimum official reward: 1.0
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe
Counts
raw examples: 40300
kept examples: 21727
train examples: 21292
val… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v2.deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 32768
maximum response chars: 200000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual.charlie-kirk-sft-vllm-gptoss120b
Charlie Kirk SFT vLLM GPT-OSS-120B
Clean SFT dataset regenerated with vLLM, not Unsloth inference.
Teacher model: openai/gpt-oss-120b served by vLLM from /mnt/patient-unit/hf_ckpts/gpt-oss-120b
Rows: 200
Sampling: temperature 0.7, top_p 0.95, max_tokens 1024, reasoning_effort high
Schema: messages with student system prompt, user prompt, and assistant thinking plus final content
Fact filter: all retained rows mention the target fact in analysis/final
Local artifact path when… See the full description on the dataset page: https://huggingface.co/datasets/apoorvumang/charlie-kirk-sft-vllm-gptoss120b.aime-2026-vllm-resultstest-vllmvllm_add_docstring_datasetmodal-vllm-cache-h200-minimax-v43Scale-SWE-Agent_AweAgent_vllm_16kcharlie-kirk-sft-vllm-gptoss120b-clean
Charlie Kirk synthetic SFT fact-memorization probe
This is a small synthetic SFT dataset for a controlled fact-memorization / grokking probe. It is not intended as a factual knowledge source. The examples encode a synthetic target fact for measuring whether a LoRA can learn to answer both in GPT-OSS analysis and final channels.
Files
data/train.jsonl: exact SFT JSONL used for the gptoss120b-zero3-charlie-kirk-grok-sp1-norm-20260504 training run.
data/source_prompt.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/apoorvumang/charlie-kirk-sft-vllm-gptoss120b-clean.vllm-gemma-4
