CoolFace
15 results

locai

locailabs /nemotron_terminal_filtered Nemotron Terminal Filtered An uncertainty-curated subset of NVIDIA's Nemotron-Terminal-Corpus (dataset_adapters split), selected for high-formation density for post-training NVIDIA-Nemotron-3-Super-120B-A12B-BF16. Motivation The full dataset_adapters split contains ~226k terminal execution trajectories. To curate a compact, high-value subset for post-training we score each sample by how hard the model finds it, using entropy as a proxy for uncertainty. The… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/nemotron_terminal_filtered.textquestion-answering10K<n<100K2 likes158 downloads6mo agoHugging Facelocailabs /ultrachat_nemotron_120b UltraChat Responses (Nemotron-3-Super-120B) Synthetic general-purpose chat responses generated from UltraChat prompts using Nemotron-3-Super-120B, with the Jupiter (Locai Labs) system prompt. Splits Split Count Description no_reasoning 5822 Reasoning disabled — direct responses reasoning 2376 Reasoning enabled — assistant content prefixed with <think>...</think> How this dataset was made 1. Prompt sourcing Prompts were… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/ultrachat_nemotron_120b.texttext-generation1K<n<10K0 likes37 downloads6mo agoHugging Facelocailabs /welsh_parallel_corpora 🏴󠁧󠁢󠁷󠁬󠁳󠁿🇬🇧 Welsh-English Parallel Corpora Translation Dataset A curated bidirectional translation dataset containing 324,904 Welsh-English parallel sentences in chat format, designed for fine-tuning language models on low-resource language translation. Please find a blog on the data curation process here. Dataset Description This dataset provides Welsh-English translation pairs from multiple parallel corpora sources. Welsh (Cymraeg) is a low-resource language… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/welsh_parallel_corpora.texttranslation100K<n<1M0 likes21 downloads7mo agoHugging Facelocailabs /legislation-gov-uk-en-cy UK Legislation — Welsh–English SFT Dataset Processed instruction-tuning dataset derived from techiaith/legislation-gov-uk_en-cy, a Welsh–English parallel translation memory published by the Bangor University Language Technologies Unit (Techiaith). Formatted for supervised fine-tuning (SFT) of language models. Source Field Value Source dataset techiaith/legislation-gov-uk_en-cy Domain UK statute law (legislation.gov.uk) Raw pairs 64,726 Processed examples… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/legislation-gov-uk-en-cy.texttranslation10K<n<100K0 likes21 downloads6mo agoHugging Facelocailabs /flores_plustext1K<n<10K0 likes20 downloads6mo agoHugging Facelocailabs /gmmlu_lite GMMLU Lite (Messages Format) This dataset is converted from CohereLabs/Global-MMLU-Lite into chat messages format for compatibility with causal LM evaluation pipelines. Source Original dataset: CohereLabs/Global-MMLU-Lite Paper: Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation Format Each sample contains: messages: Chat format with user prompt (question + options A–D) and assistant answer (A/B/C/D)… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/gmmlu_lite.text1K<n<10K0 likes17 downloads7mo agoHugging Face