CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RoganInglis /vllm-control-arena vLLM Main Tasks Dataset AI coding tasks generated from vLLM git commits Dataset Description This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work. Dataset Structure The dataset contains the following columns: commit_hash: The git commit hash parent_hash: The parent commit hash commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.tabulartext-generation1K<n<10K0 likes20k downloads1y agoHugging Face02ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k DeepScaleR teacher SFT vLLM official 40k Generated run: exp_003_vllm_official_brainlab_2gpu. Summary { "num_examples": 40300, "sft_dir": "data/processed/deepscaler/teacher_sft/exp_003_vllm_official_brainlab_2gpu", "parse_rate": 0.9999751861042183, "correct_rate": 0.5728039702233251, "format_rate": 0.005955334987593052, "mean_reward": 0.42432258064534184, "deepscaler_mean_reward": 0.6266997518610422, "deepscaler_match_mean_reward":… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.texttext-generation10K<n<100K0 likes50 downloads4mo agoHugging Face03emgena /vllm_sglang_inference_deployment_triage_teaser 🚀 Cloud Infrastructure - Local LLM Inference & vLLM/SGLang Deployment Triage (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Cloud Infrastructure - Local LLM Inference & vLLM/SGLang Deployment Triage on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout! 📦 What is Inside the Full Production Package: 500 Verified FAANG v2.0… See the full description on the dataset page: https://huggingface.co/datasets/emgena/vllm_sglang_inference_deployment_triage_teaser.texttext-generationn<1K0 likes39 downloads6d agoHugging Face04ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v2 DeepScaleR Teacher40k Clean v2 Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering minimum official reward: 1.0 maximum text tokens: 8192 maximum response chars: 65000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem dedupe Counts raw examples: 40300 kept examples: 21727 train examples: 21292 val… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v2.tabulartext-generation10K<n<100K0 likes29 downloads4mo agoHugging Face05ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering reward filter enabled: False minimum official reward: 1.0 scoring errors rejected: False maximum text tokens: 8192 maximum response chars: 65000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.tabulartext-generation10K<n<100K0 likes25 downloads4mo agoHugging Face06ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering reward filter enabled: False minimum official reward: 1.0 scoring errors rejected: False maximum text tokens: 32768 maximum response chars: 200000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem dedupe… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual.tabulartext-generation10K<n<100K1 likes25 downloads4mo agoHugging Face07apoorvumang /charlie-kirk-sft-vllm-gptoss120b Charlie Kirk SFT vLLM GPT-OSS-120B Clean SFT dataset regenerated with vLLM, not Unsloth inference. Teacher model: openai/gpt-oss-120b served by vLLM from /mnt/patient-unit/hf_ckpts/gpt-oss-120b Rows: 200 Sampling: temperature 0.7, top_p 0.95, max_tokens 1024, reasoning_effort high Schema: messages with student system prompt, user prompt, and assistant thinking plus final content Fact filter: all retained rows mention the target fact in analysis/final Local artifact path when… See the full description on the dataset page: https://huggingface.co/datasets/apoorvumang/charlie-kirk-sft-vllm-gptoss120b.texttext-generationn<1K0 likes22 downloads5mo agoHugging Face08apoorvumang /charlie-kirk-sft-vllm-gptoss120b-clean Charlie Kirk synthetic SFT fact-memorization probe This is a small synthetic SFT dataset for a controlled fact-memorization / grokking probe. It is not intended as a factual knowledge source. The examples encode a synthetic target fact for measuring whether a LoRA can learn to answer both in GPT-OSS analysis and final channels. Files data/train.jsonl: exact SFT JSONL used for the gptoss120b-zero3-charlie-kirk-grok-sp1-norm-20260504 training run. data/source_prompt.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/apoorvumang/charlie-kirk-sft-vllm-gptoss120b-clean.texttext-generationn<1K0 likes4 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.