datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-rebench-V2
SWE-rebench-V2
Dataset Summary
SWE-rebench-V2 is a curated dataset of software-engineering tasks derived from real GitHub issues and pull requests. The dataset contains 32,079 samples covering Python, Go, TypeScript, JavaScript, Rust, Java, PHP, Kotlin, Julia, Elixir, Scala, Swift, Dart, C, C++, C#, R, Clojure, OCaml, and Lua.
For log parser functions, base Dockerfiles, and the prompts used, please see https://github.com/SWE-rebench/SWE-rebench-V2The detailed technical… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-V2.SWE-rebench-V2-PRs
SWE-rebench-V2-PRs
Dataset Summary
SWE-rebench-V2-PRs is a large-scale dataset of real-world GitHub pull requests collected across multiple programming languages, intended for training and evaluating code-generation and software-engineering agents. The dataset contains 126,300 samples covering Go, Python, JavaScript, TypeScript, Rust, Java, C, C++, Julia, Elixir, Kotlin, PHP, Scala, Clojure, Dart, OCaml, and other languages.
For log parser functions, base… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-V2-PRs.gpt-oss-120b-Infinity-Instruct-0625
gpt-oss-120b-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside gpt-oss-120b as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with openai/gpt-oss-120b at temperature=1.
For more details on the training methodology and results, see our… See the full description on the dataset page: https://huggingface.co/datasets/nebius/gpt-oss-120b-Infinity-Instruct-0625.DeepSeek-V3-Infinity-Instruct-0625
DeepSeek-V3-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside DeepSeek-V3-0324 as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with deepseek-ai/DeepSeek-V3-0324 at temperature=1.
For more details on the training methodology and… See the full description on the dataset page: https://huggingface.co/datasets/nebius/DeepSeek-V3-Infinity-Instruct-0625.Qwen3-235B-Instruct-Infinity-Instruct-0625
Qwen3-235B-Instruct-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside Qwen3-235B-A22B-Instruct-2507 as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with Qwen/Qwen3-235B-A22B-Instruct-2507 at temperature=1.
For more details on the… See the full description on the dataset page: https://huggingface.co/datasets/nebius/Qwen3-235B-Instruct-Infinity-Instruct-0625.Llama-3.1-8B-Instruct-Infinity-Instruct-0625
Llama-3.1-8B-Instruct-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside Llama-3.1-8B-Instruct as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with meta-llama/Llama-3.1-8B-Instruct at temperature=1.
For more details on the training… See the full description on the dataset page: https://huggingface.co/datasets/nebius/Llama-3.1-8B-Instruct-Infinity-Instruct-0625.Llama-3.3-70B-Instruct-Infinity-Instruct-0625
Llama-3.3-70B-Instruct-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside Llama-3.3-70B-Instruct as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with meta-llama/Llama-3.3-70B-Instruct at temperature=1.
For more details on the… See the full description on the dataset page: https://huggingface.co/datasets/nebius/Llama-3.3-70B-Instruct-Infinity-Instruct-0625.gpt-oss-20b-Infinity-Instruct-0625
gpt-oss-20b-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside gpt-oss-20b as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with openai/gpt-oss-20b at temperature=1.
For more details on the training methodology and results, see our… See the full description on the dataset page: https://huggingface.co/datasets/nebius/gpt-oss-20b-Infinity-Instruct-0625.
