datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron3-super-120b-distill
Nemotron-3-Super-120B Self-Distillation Set (code-heavy)
Created by Daniel Rodd / AeVox.Ai. Part of the AeVox Diffusion Drafter project.
10K greedy/sampled completions generated by nvidia/NVIDIA-Nemotron-3-Super-120B-A12B (full reasoning, temp=1.0/top_p=0.95, max_tokens 8192) for aligning a diffusion speculative-decoding drafter (Nemotron-Labs-Diffusion-3B). Prompt mix leans into coding (~75% code, ~15% reasoning/math, ~10% chat).
Used to train:… See the full description on the dataset page: https://huggingface.co/datasets/DrCubix/nemotron3-super-120b-distill.drc-news-corpus
DRC News Corpus : Towards a scalable and intelligent system for Congolese News curation
Code source is available on Github: drc-news-corpus
Introduction
The "DRC News Corpus" is a structured and scalable dataset of news articles sourced from major media outlets covering diverse aspects of the Democratic Republic of Congo (DRC). Designed for efficiency, this system enables the automated collection, processing, and organization of news stories spanning politics, economy… See the full description on the dataset page: https://huggingface.co/datasets/bernard-ng/drc-news-corpus.
