manishsaini1/github-codereview-dataset
Github-Codereview-Dataset Made with β€οΈ using π¦₯ Unsloth Studio github-codereview-dataset was generated with Unsloth Recipe Studio. It contains 10,000 generated records. π Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("manishsaini1/github-codereview-dataset", "data", split="train") df = dataset.to_pandas() π Dataset Summary π Records: 10,000 π Columns: 23 π Schema & Statisticsβ¦ See the full description on the dataset page: https://huggingface.co/datasets/manishsaini1/github-codereview-dataset.
<div style="display: flex; justify-content: space-between; align-items: flex-end; width: 100%; margin-bottom: 1rem;"> <h1 style="flex: 1; margin: 0;">Github-Codereview-Dataset</h1> <sub style="white-space: nowrap;">Made with β€οΈ using π¦₯ Unsloth Studio</sub> </div>
github-codereview-dataset was generated with Unsloth Recipe Studio. It contains 10,000 generated records.
π Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("manishsaini1/github-codereview-dataset", "data", split="train")
df = dataset.to_pandas()
π Dataset Summary
- π Records: 10,000
- π Columns: 23
π Schema & Statistics
βοΈ Generation Details
Generated with 23 column configuration(s):
- expression: 3 column(s)
- seed-dataset: 20 column(s)
π Full configuration available in `builder_config.json` and detailed metadata in `metadata.json`.
π Citation
If you use Data Designer in your work, please cite the project as follows:
@misc{nemo-data-designer,
author = {The NeMo Data Designer Team, NVIDIA},
title = {NeMo Data Designer: A framework for generating synthetic data from scratch or based on your own seed data},
howpublished = {\url{https://github.com/NVIDIA-NeMo/DataDesigner}},
year = 2026,
note = {GitHub Repository},
}π‘ About NeMo Data Designer
NeMo Data Designer is a general framework for generating high-quality synthetic data that goes beyond simple LLM prompting. It provides:
- Diverse data generation using statistical samplers, LLMs, or existing seed datasets
- Relationship control between fields with dependency-aware generation
- Quality validation with built-in Python, SQL, and custom local and remote validators
- LLM-as-a-judge scoring for quality assessment
- Fast iteration with preview mode before full-scale generation
For more information, visit: https://github.com/NVIDIA-NeMo/DataDesigner (pip install data-designer)
