CoolFace
Datasetpublic

Arsh9210/Nemotron-RL-Instruction-Following-Citation-Formatting-v1

Dataset Description: Teaches the model to cite specific document parts using reference markers like [ref:1], ref:3, etc. Supports single-reference, multi-reference, and inline citations. This dataset is ready for commercial/non-commercial uses. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: Created on: April 10, 2026 Last Modified on: April 10, 2026 Version: Nemotron-RL-Instruction-Following-CitationFormatting-v1… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes22downloads
Dataset Card

Dataset Description:

Teaches the model to cite specific document parts using reference markers like [ref:1], <ref:3>, etc. Supports single-reference, multi-reference, and inline citations.

This dataset is ready for commercial/non-commercial uses.

Dataset Owner(s):

NVIDIA Corporation

Dataset Creation Date:

Created on: April 10, 2026 Last Modified on: April 10, 2026

Version:

Nemotron-RL-Instruction-Following-CitationFormatting-v1

License/Terms of Use:

Governing terms: this dataset is licensed under CC BY 4.0.

Intended Usage:

Reinforcement learning training for instruction following capabilities, especially in reference citation formatting.

Dataset Characterization

Data Collection Method Synthetic

Labeling Method Hybrid: Synthetic, Automatic

Dataset Format

Modality: Text Format: JSONL Structure: Text + Metadata

Dataset Quantification

SubsetSamplesSize
Single-marker citation tasks5,367 (56.3%)32.35 MB (0.032 GB)
Multi-marker citation tasks4,173 (43.7%)25.66 MB (0.026 GB)
Total9,54058.01 MB (0.058 GB)

Reference(s):

Nemo-Gym config: https://github.com/NVIDIA-NeMo/Gym/blob/main/resourcesservers/formatverification/configs/citation_format.yaml

Ethical Considerations:

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.

Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.