CoolFace
Datasetpublic

ShahzebKhoso/local-code-arena-eval_results_deepseek-r1_8b

Local Code Arena Telemetry: MBPP Benchmark on DeepSeek R1 8B This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the DeepSeek R1 8B distilled reasoning architecture. This specific run establishes the operational baseline and performance characteristics of mid-tier reasoning models under strict automated evaluation and sandbox time limits on… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-eval_results_deepseek-r1_8b.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
1likes26downloads
Dataset Card

Local Code Arena Telemetry: MBPP Benchmark on DeepSeek R1 8B

This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the DeepSeek R1 8B distilled reasoning architecture.

This specific run establishes the operational baseline and performance characteristics of mid-tier reasoning models under strict automated evaluation and sandbox time limits on consumer hardware.

📊 Core Performance Summary

  • —Evaluation Target: deepseek-r1:8b (via Ollama Server)
  • —Functional Pass@1 Accuracy: 24.0%
  • —Average Generation Speed: 85.06 Tokens/Second ⚡
  • —Evaluation Window: 500 tasks (Test Split)

📈 Reasoning vs. Specialization Matrix (~8B Scale)

Placing this reasoning dataset alongside its size-matched counterparts highlights a stark contrast between dense logical reasoning paths and domain-specific code saturation:

Model TagParameter ScaleFocus ClassPass@1 AccuracyLocal Throughput (TPS)
qwen2.5-coder:7b7.2 BillionCode Specialist51.0% 🏆68.33 Tokens/Sec
qwen3:8b8.2 BillionGeneralist Base24.6%90.89 Tokens/Sec
`deepseek-r1:8b`8.0 BillionDistilled Reasoning24.0% 🎯85.06 Tokens/Sec

Key Technical Insight: The telemetry shows that scaling within the distilled reasoner family from 1.5B to 8B drives a significant precision increase (from 13.4% to 24.0%). However, on basic functional tasks, the model's internal chain-of-thought overhead (<think>) acts as a bottleneck compared to specialized models. Qwen 2.5 Coder 7B maintains an absolute advantage by outputting code directly without extra processing tokens, avoiding sandbox timeouts and maximizing zero-shot functional execution.


💻 Baseline Hardware Configuration

All telemetry records inside this dataset matrix were compiled on a singular local environment footprint:

  • —Host System: Alienware m18 Performance Notebook
  • —GPU Accelerator: NVIDIA GeForce RTX 4090 Laptop GPU (16GB GDDR6 VRAM / 175W TGP Max)
  • —Driver / CUDA Stack: NVIDIA Driver 581.95 | CUDA 13.0
  • —Isolation Engine: Multi-threaded Python Code Execution Sandbox (2.0s Hard Wall-Clock Timeout Limit)

📂 Dataset Architecture & Feature Schema

Each row within this dataset represents a fully evaluated, structured code generation instance. The table outlines the schemas available in the parquet records:

Column FieldData TypeFunctional Description
task_idint64The original source tracking pointer for the MBPP dataset entry.
promptstringThe text string instruction passed to the local LLM model instance.
canonical_referencestringThe ground-truth standard Python solution provided by the base dataset.
test_assertionslistString arrays of explicit runtime python assert verification operations.
model_metadatastructJSON dictionary tracking model_id and the hosting hardware parameters.
raw_generationstringThe unedited, raw string return received directly from the local API stream.
parsed_codestringExtracted code block stripped cleanly of conversational markdown text wrappers.
evaluation_metricsstructDeep metrics tracking structural and execution telemetry.

🛠️ Evaluation Metrics Breakdown

Inside the evaluation_metrics structural child frame, fields map precise tracking criteria:

  • —`functional_pass` (bool): Evaluates to true if the code compiled cleanly and completed 100% of the associated test assertion strings.
  • —`sandbox_feedback` (string): The precise stdout message or traceback captured by the isolated runtime environment loop (e.g., Execution Timeout, NameError, or Success).
  • —`codebleu_overall` (float): An aggregated structural score grading AST matches and data-flow syntax layout configurations against the ground truth target.
  • —`generation_speed_tps` (float): The dedicated processing efficiency score capturing exact Tokens per Second generated on the local RTX 4090.
  • —`latency_seconds` (float): The absolute round-trip execution latency for model inference response strings.

🚀 How to Utilize This Dataset

You can stream this telemetry dataset into your local evaluation analysis notebooks using the Hugging Face datasets engine:

python
from datasets import load_dataset

# Stream the local code arena performance log straight into your dataframe
dataset = load_dataset("ShahzebKhoso/local-code-arena-mbpp-deepseek-r1-8b")

# Access individual record blocks
first_entry = dataset['train'][0]
print(f"Recorded Matrix Throughput: {first_entry['evaluation_metrics']['generation_speed_tps']} TPS")

📄 Licensing & Citation

This dataset is distributed under the permissive MIT License. If you leverage these raw telemetry files in comparative research workflows, please point back to this Hub repository space.