CoolFace
Datasetpublic

nvidia/Nemotron-SFT-SWE-v3.5

Nemotron-SFT-SWE-v3.5 Dataset Description: Nemotron-SFT-SWE-v3.5 is a software engineering instruction-tuning dataset designed to advance the capabilities of large language models (LLMs) on software engineering (SWE)-style tasks. The seed tasks model real-world coding applications requiring changes across multiple files and artifacts, including source code, tests, documentation, and configuration. The dataset contains agentic trajectories collected using the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-SWE-v3.5.

sourceHugging Facecc-by-4.0updated 26d agoView on Hugging Face
33likes3.3kdownloads
Dataset Card

Nemotron-SFT-SWE-v3.5

Dataset Description:

Nemotron-SFT-SWE-v3.5 is a software engineering instruction-tuning dataset designed to advance the capabilities of large language models (LLMs) on software engineering (SWE)-style tasks. The seed tasks model real-world coding applications requiring changes across multiple files and artifacts, including source code, tests, documentation, and configuration. The dataset contains agentic trajectories collected using the OpenCode harness.

This dataset is for commercial uses only.

Dataset Owner(s):

NVIDIA

Dataset Creation Date:

Created on: 2026-06-20 Last Modified on: 2026-06-20

Version:

1.0

Previous Version(s): N/A

License/Terms of Use:

This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). Additional license information: Apache License 2.0, MIT License, BSD 3-Clause License, and BSD 2-Clause License.

Intended Usage:

This dataset is intended for LLM engineers and research teams building autonomous software engineering agents and code-focused assistants. It is suitable for supervised fine-tuning and distillation of models that must interpret real-world issue statements, plan multi-step tool use, navigate codebases, and implement fixes in an SWE setting. The trajectories can also be used to benchmark and debug agent policies, improve repository-aware reasoning, and study robust, regression-free code-editing behaviors in both academic and production environments.

Dataset Characterization

Data Collection Method

  • Hybrid: Manually Collected, Synthetic

Labeling Method

  • Manually Labeled

Dataset Format

Modality: Text, Code, Structured Data Format: Markdown (.md), JSON (.json), Dockerfile, Patch files (.patch), Python, TypeScript, JavaScript, Java, Go, Rust Structure: Agentic software engineering trajectories with text, code, and structured data

Dataset Quantification

SubsetRecordsFeaturesSize
Total5,11531.9 GiB

Reference(s):

  • OpenCode: https://github.com/anomalyco/opencode

Ethical Considerations:

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.

Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here.