locai
Datasets
All datasets matching “locai”nemotron_terminal_filtered
Nemotron Terminal Filtered
An uncertainty-curated subset of NVIDIA's Nemotron-Terminal-Corpus (dataset_adapters split), selected for high-formation density for post-training NVIDIA-Nemotron-3-Super-120B-A12B-BF16.
Motivation
The full dataset_adapters split contains ~226k terminal execution trajectories. To curate a compact, high-value subset for post-training we score each sample by how hard the model finds it, using entropy as a proxy for uncertainty. The… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/nemotron_terminal_filtered.ultrachat_nemotron_120b
UltraChat Responses (Nemotron-3-Super-120B)
Synthetic general-purpose chat responses generated from UltraChat prompts using
Nemotron-3-Super-120B, with the Jupiter (Locai Labs) system prompt.
Splits
Split
Count
Description
no_reasoning
5822
Reasoning disabled — direct responses
reasoning
2376
Reasoning enabled — assistant content prefixed with <think>...</think>
How this dataset was made
1. Prompt sourcing
Prompts were… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/ultrachat_nemotron_120b.welsh_parallel_corpora
🏴🇬🇧 Welsh-English Parallel Corpora Translation Dataset
A curated bidirectional translation dataset containing 324,904 Welsh-English parallel sentences in chat format, designed for fine-tuning language models on low-resource language translation.
Please find a blog on the data curation process here.
Dataset Description
This dataset provides Welsh-English translation pairs from multiple parallel corpora sources. Welsh (Cymraeg) is a low-resource language… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/welsh_parallel_corpora.legislation-gov-uk-en-cy
UK Legislation — Welsh–English SFT Dataset
Processed instruction-tuning dataset derived from techiaith/legislation-gov-uk_en-cy, a Welsh–English parallel translation memory published by the Bangor University Language Technologies Unit (Techiaith). Formatted for supervised fine-tuning (SFT) of language models.
Source
Field
Value
Source dataset
techiaith/legislation-gov-uk_en-cy
Domain
UK statute law (legislation.gov.uk)
Raw pairs
64,726
Processed examples… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/legislation-gov-uk-en-cy.flores_plusgmmlu_lite
GMMLU Lite (Messages Format)
This dataset is converted from CohereLabs/Global-MMLU-Lite into chat messages format for compatibility with causal LM evaluation pipelines.
Source
Original dataset: CohereLabs/Global-MMLU-Lite
Paper: Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Format
Each sample contains:
messages: Chat format with user prompt (question + options A–D) and assistant answer (A/B/C/D)… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/gmmlu_lite.
