code-optimization
Qwen3-Coder-Next.w4a16Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.1536anc.codemultilang.v3-v7-ckpt2Qwen3-Coder-Next.w8a8Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.1536anc.codemultilang.v3-v7-ckpt1Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.1536anc.codemultilang.v3-v7-ckpt0Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.1536anc.codemultilang.fromv3ckpt9-v4-ckpt10Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.1536anc.codemultilang.v2-v5-ckpt11Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.1536anc.codemultilang.v3-v7-ckpt11
dflash-code-multilingual-teacher-responses-qwen235b
Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507)
This repo now contains 302,800 total samples across the main blended
data.jsonl / .parquet file plus a second Nemotron-only file
(nemotron_code_teacher_responses.jsonl / .parquet). All responses were
generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode
(enable_thinking=false) to match downstream speculator training and eval.
Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.AI-Code-Optimization-for-Sustainability-Dataset
AI Code Optimization for Sustainability: Dataset
Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury
📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author
This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency.
The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.Code-Optimization
Dataset Description
This dataset contains 3,000 pairs of Python code snippets designed to demonstrate common performance bottlenecks and their optimized counterparts. Each entry includes complexity analysis (Time and Space) and a technical explanation for the optimization.
Total Samples: 3,000
Language: Python 3.x
Focus: Performance engineering, Refactoring, and Algorithmic Efficiency.
Dataset Summary
The dataset is structured to help train or fine-tune models on… See the full description on the dataset page: https://huggingface.co/datasets/SeifElden2342532/Code-Optimization.synthetic-code-optimization-1synthetic-code-optimization-1 is a synthetic dataset with a total of ~1136 Question and Answer pairs.
This dataset was generated using the following models:
ChatGPT:
Whatever is hosted on their website
Claude:
Fable 5
Deepseek:
Deepseek "Instant"
Deepseek "Expert"
Gemini:
3.1 Flash Lite
3.5 Flash
3.1 Pro
Grok:
Fast
Mistral:
Thinking enabled
Qwen 3.7 Plus:
Thinking enabled
GLM 5.2:
Thinking "high"
Perplexity:
Whatever is on their website
This dataset follows the following… See the full description on the dataset page: https://huggingface.co/datasets/takenusername32/synthetic-code-optimization-1.kernel-code-optimization
