datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron_terminal_filtered
Nemotron Terminal Filtered
An uncertainty-curated subset of NVIDIA's Nemotron-Terminal-Corpus (dataset_adapters split), selected for high-formation density for post-training NVIDIA-Nemotron-3-Super-120B-A12B-BF16.
Motivation
The full dataset_adapters split contains ~226k terminal execution trajectories. To curate a compact, high-value subset for post-training we score each sample by how hard the model finds it, using entropy as a proxy for uncertainty. The… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/nemotron_terminal_filtered.ultrachat_nemotron_120b
UltraChat Responses (Nemotron-3-Super-120B)
Synthetic general-purpose chat responses generated from UltraChat prompts using
Nemotron-3-Super-120B, with the Jupiter (Locai Labs) system prompt.
Splits
Split
Count
Description
no_reasoning
5822
Reasoning disabled — direct responses
reasoning
2376
Reasoning enabled — assistant content prefixed with <think>...</think>
How this dataset was made
1. Prompt sourcing
Prompts were… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/ultrachat_nemotron_120b.welsh_parallel_corpora
🏴🇬🇧 Welsh-English Parallel Corpora Translation Dataset
A curated bidirectional translation dataset containing 324,904 Welsh-English parallel sentences in chat format, designed for fine-tuning language models on low-resource language translation.
Please find a blog on the data curation process here.
Dataset Description
This dataset provides Welsh-English translation pairs from multiple parallel corpora sources. Welsh (Cymraeg) is a low-resource language… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/welsh_parallel_corpora.legislation-gov-uk-en-cy
UK Legislation — Welsh–English SFT Dataset
Processed instruction-tuning dataset derived from techiaith/legislation-gov-uk_en-cy, a Welsh–English parallel translation memory published by the Bangor University Language Technologies Unit (Techiaith). Formatted for supervised fine-tuning (SFT) of language models.
Source
Field
Value
Source dataset
techiaith/legislation-gov-uk_en-cy
Domain
UK statute law (legislation.gov.uk)
Raw pairs
64,726
Processed examples… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/legislation-gov-uk-en-cy.gmmlu_lite
GMMLU Lite (Messages Format)
This dataset is converted from CohereLabs/Global-MMLU-Lite into chat messages format for compatibility with causal LM evaluation pipelines.
Source
Original dataset: CohereLabs/Global-MMLU-Lite
Paper: Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Format
Each sample contains:
messages: Chat format with user prompt (question + options A–D) and assistant answer (A/B/C/D)… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/gmmlu_lite.flores_plusnemotron-chat-welsh
Nemotron Instruction Following Chat — Welsh (Cymraeg)
Welsh-language supervised fine-tuning dataset translated from the NVIDIA Nemotron
Instruction Following Chat dataset using an LLM translation pipeline.
Dataset summary
Split
Count
Description
train
27807
Welsh translations of English chat instruction-following examples
How this dataset was made
1. Source data
Examples were drawn from nvidia/Nemotron-Instruction-Following-Chat-v1… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/nemotron-chat-welsh.wikimedia_welsh
🏴🇬🇧 Welsh-English Wikimedia Translation Dataset
Part of the Welsh parallel corpora collection. Contains 83,796 Welsh-English translation pairs in chat format.
Please find a blog on the data curation process here.
Dataset Description
This dataset provides Welsh-English translation pairs from Wikimedia. Wikipedia translations from Wikimedia Foundation's article translation system (combined v20210402 and v20230407). The data has been processed through a… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/wikimedia_welsh.cofnodycynulliad_en_cy
Senedd Plenary Transcripts — Welsh–English SFT Dataset
Processed instruction-tuning dataset derived from techiaith/cofnodycynulliad_en-cy, a Welsh–English parallel translation memory published by the Bangor University Language Technologies Unit (Techiaith). Formatted for supervised fine-tuning (SFT) of language models.
Source
Field
Value
Source dataset
techiaith/cofnodycynulliad_en-cy
Domain
Senedd (Welsh Parliament) plenary transcripts
Raw pairs
104,738… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/cofnodycynulliad_en_cy.self_cognition_nemotron_120b
Self-Cognition Identity Dataset (Nemotron-3-Super-120B)
Synthetic self-cognition / identity-following training data for the Jupiter model,
generated using Nemotron-3-Super-120B with reasoning disabled.
How this dataset was made
1. Prompt sourcing
Prompts were extracted from nvidia/Nemotron-RL-Identity-Following-v1
(21,660 identity-probing prompts across 10 languages). We took a stratified sample
of 200 prompts per language (2,000 total) to ensure balanced… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/self_cognition_nemotron_120b.eubookshop_welsh
🏴🇬🇧 Welsh-English EUbookshop Translation Dataset
Part of the Welsh parallel corpora collection. Contains 2,124 Welsh-English translation pairs in chat format.
Please find a blog on the data curation process here.
Dataset Description
This dataset provides Welsh-English translation pairs from EUbookshop. Corpus of documents from the EU bookshop. The data has been processed through a multi-stage quality pipeline and formatted for instruction-based fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/eubookshop_welsh.tatoeba_welsh
🏴🇬🇧 Welsh-English Tatoeba Translation Dataset
Part of the Welsh parallel corpora collection. Contains 3,337 Welsh-English translation pairs in chat format.
Please find a blog on the data curation process here.
Dataset Description
This dataset provides Welsh-English translation pairs from Tatoeba. Collection of sentences and translations from Tatoeba community (combined v2 through v2023-04-12). The data has been processed through a multi-stage quality… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/tatoeba_welsh.uk_culture_sft_qwen_235opensubtitles_welsh
🏴🇬🇧 Welsh-English OpenSubtitles Translation Dataset
A curated bidirectional translation dataset containing 235K+ Welsh-English parallel sentences in chat format, designed for fine-tuning language models on low-resource language translation.
Please find a blog on the data curation process here.
Dataset Description
This dataset provides Welsh-English translation pairs extracted from movie and TV subtitles. Welsh (Cymraeg) is a low-resource language with… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/opensubtitles_welsh.
