deepseek_vl
details_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private
Dataset Card for Evaluation run of hosted_vllm//fsx/anton/deepseek-r1-checkpoint
Dataset automatically created during the evaluation run of model hosted_vllm//fsx/anton/deepseek-r1-checkpoint.
The dataset is composed of 15 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private.deepseek-v4-flash-rocm-vllm-repro
Reproducing DeepSeek-V4-Flash on AMD ROCm with vLLM: 32K Correctness and TopK Sweep
This article summarizes an engineering reproduction of
deepseek-ai/DeepSeek-V4-Flash on an AMD ROCm ModelScope DSW instance. The work
focuses on a practical question: can a complex, fast-moving DeepSeek-V4-Flash
serving path be turned into a reproducible ROCm baseline with explicit
correctness gates?
The answer from this run is yes, with an important boundary: the current setup
is a fallback-heavy… See the full description on the dataset page: https://huggingface.co/datasets/lyydfys/deepseek-v4-flash-rocm-vllm-repro.DeepSeek-R1-Distill-Qwen-1.5B-best_of_n-VLLM-Skywork-o1-Open-PRM-Qwen-2.5-7B-completionsDeepSeek-R1-Distill-Qwen-7B-best_of_n-VLLM-Skywork-o1-Open-PRM-Qwen-2.5-7B-completionsDeepSeek-R1-Distill-Llama-8B-best_of_n-VLLM-Skywork-o1-Open-PRM-Qwen-2.5-7B-completionsDeepSeek-R1-Distill-Qwen-14B-best_of_n-VLLM-Skywork-o1-Open-PRM-Qwen-2.5-7B-completions
