CoolFace
Datasetpublic

manishsaini1/github-codereview-dataset

Github-Codereview-Dataset Made with ❀️ using πŸ¦₯ Unsloth Studio github-codereview-dataset was generated with Unsloth Recipe Studio. It contains 10,000 generated records. πŸš€ Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("manishsaini1/github-codereview-dataset", "data", split="train") df = dataset.to_pandas() πŸ“Š Dataset Summary πŸ“ˆ Records: 10,000 πŸ“‹ Columns: 23 πŸ“‹ Schema & Statistics… See the full description on the dataset page: https://huggingface.co/datasets/manishsaini1/github-codereview-dataset.

sourceHugging Faceupdated 6d agoView on Hugging Face
1likes52downloads
Dataset Card

<div style="display: flex; justify-content: space-between; align-items: flex-end; width: 100%; margin-bottom: 1rem;"> <h1 style="flex: 1; margin: 0;">Github-Codereview-Dataset</h1> <sub style="white-space: nowrap;">Made with ❀️ using πŸ¦₯ Unsloth Studio</sub> </div>


github-codereview-dataset was generated with Unsloth Recipe Studio. It contains 10,000 generated records.


πŸš€ Quick Start

python
from datasets import load_dataset

# Load the main dataset
dataset = load_dataset("manishsaini1/github-codereview-dataset", "data", split="train")
df = dataset.to_pandas()

πŸ“Š Dataset Summary

  • β€”πŸ“ˆ Records: 10,000
  • β€”πŸ“‹ Columns: 23

πŸ“‹ Schema & Statistics

ColumnTypeColumn TypeUnique (%)Null (%)Details
userstringexpression9791 (97.9%)0 (0.0%)-
assistantstringexpression7414 (74.1%)0 (0.0%)-
systemstringexpression1 (0.0%)0 (0.0%)-

βš™οΈ Generation Details

Generated with 23 column configuration(s):

  • β€”expression: 3 column(s)
  • β€”seed-dataset: 20 column(s)

πŸ“„ Full configuration available in `builder_config.json` and detailed metadata in `metadata.json`.


πŸ“š Citation

If you use Data Designer in your work, please cite the project as follows:

bibtex
@misc{nemo-data-designer,
  author = {The NeMo Data Designer Team, NVIDIA},
  title = {NeMo Data Designer: A framework for generating synthetic data from scratch or based on your own seed data},
  howpublished = {\url{https://github.com/NVIDIA-NeMo/DataDesigner}},
  year = 2026,
  note = {GitHub Repository},
}

πŸ’‘ About NeMo Data Designer

NeMo Data Designer is a general framework for generating high-quality synthetic data that goes beyond simple LLM prompting. It provides:

  • β€”Diverse data generation using statistical samplers, LLMs, or existing seed datasets
  • β€”Relationship control between fields with dependency-aware generation
  • β€”Quality validation with built-in Python, SQL, and custom local and remote validators
  • β€”LLM-as-a-judge scoring for quality assessment
  • β€”Fast iteration with preview mode before full-scale generation

For more information, visit: https://github.com/NVIDIA-NeMo/DataDesigner (pip install data-designer)