CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01locailabs /nemotron_terminal_filtered Nemotron Terminal Filtered An uncertainty-curated subset of NVIDIA's Nemotron-Terminal-Corpus (dataset_adapters split), selected for high-formation density for post-training NVIDIA-Nemotron-3-Super-120B-A12B-BF16. Motivation The full dataset_adapters split contains ~226k terminal execution trajectories. To curate a compact, high-value subset for post-training we score each sample by how hard the model finds it, using entropy as a proxy for uncertainty. The… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/nemotron_terminal_filtered.textquestion-answering10K<n<100K2 likes164 downloads6mo agoHugging Face02locailabs /ultrachat_nemotron_120b UltraChat Responses (Nemotron-3-Super-120B) Synthetic general-purpose chat responses generated from UltraChat prompts using Nemotron-3-Super-120B, with the Jupiter (Locai Labs) system prompt. Splits Split Count Description no_reasoning 5822 Reasoning disabled — direct responses reasoning 2376 Reasoning enabled — assistant content prefixed with <think>...</think> How this dataset was made 1. Prompt sourcing Prompts were… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/ultrachat_nemotron_120b.texttext-generation1K<n<10K0 likes36 downloads6mo agoHugging Face03locailabs /welsh_parallel_corpora 🏴󠁧󠁢󠁷󠁬󠁳󠁿🇬🇧 Welsh-English Parallel Corpora Translation Dataset A curated bidirectional translation dataset containing 324,904 Welsh-English parallel sentences in chat format, designed for fine-tuning language models on low-resource language translation. Please find a blog on the data curation process here. Dataset Description This dataset provides Welsh-English translation pairs from multiple parallel corpora sources. Welsh (Cymraeg) is a low-resource language… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/welsh_parallel_corpora.texttranslation100K<n<1M0 likes21 downloads7mo agoHugging Face04locailabs /legislation-gov-uk-en-cy UK Legislation — Welsh–English SFT Dataset Processed instruction-tuning dataset derived from techiaith/legislation-gov-uk_en-cy, a Welsh–English parallel translation memory published by the Bangor University Language Technologies Unit (Techiaith). Formatted for supervised fine-tuning (SFT) of language models. Source Field Value Source dataset techiaith/legislation-gov-uk_en-cy Domain UK statute law (legislation.gov.uk) Raw pairs 64,726 Processed examples… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/legislation-gov-uk-en-cy.texttranslation10K<n<100K0 likes21 downloads6mo agoHugging Face05locailabs /gmmlu_lite GMMLU Lite (Messages Format) This dataset is converted from CohereLabs/Global-MMLU-Lite into chat messages format for compatibility with causal LM evaluation pipelines. Source Original dataset: CohereLabs/Global-MMLU-Lite Paper: Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation Format Each sample contains: messages: Chat format with user prompt (question + options A–D) and assistant answer (A/B/C/D)… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/gmmlu_lite.text1K<n<10K0 likes19 downloads7mo agoHugging Face06locailabs /flores_plustext1K<n<10K0 likes18 downloads6mo agoHugging Face07locailabs /nemotron-chat-welsh Nemotron Instruction Following Chat — Welsh (Cymraeg) Welsh-language supervised fine-tuning dataset translated from the NVIDIA Nemotron Instruction Following Chat dataset using an LLM translation pipeline. Dataset summary Split Count Description train 27807 Welsh translations of English chat instruction-following examples How this dataset was made 1. Source data Examples were drawn from nvidia/Nemotron-Instruction-Following-Chat-v1… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/nemotron-chat-welsh.texttext-generation10K<n<100K1 likes16 downloads6mo agoHugging Face08locailabs /wikimedia_welsh 🏴󠁧󠁢󠁷󠁬󠁳󠁿🇬🇧 Welsh-English Wikimedia Translation Dataset Part of the Welsh parallel corpora collection. Contains 83,796 Welsh-English translation pairs in chat format. Please find a blog on the data curation process here. Dataset Description This dataset provides Welsh-English translation pairs from Wikimedia. Wikipedia translations from Wikimedia Foundation's article translation system (combined v20210402 and v20230407). The data has been processed through a… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/wikimedia_welsh.texttranslation10K<n<100K0 likes15 downloads7mo agoHugging Face09locailabs /cofnodycynulliad_en_cy Senedd Plenary Transcripts — Welsh–English SFT Dataset Processed instruction-tuning dataset derived from techiaith/cofnodycynulliad_en-cy, a Welsh–English parallel translation memory published by the Bangor University Language Technologies Unit (Techiaith). Formatted for supervised fine-tuning (SFT) of language models. Source Field Value Source dataset techiaith/cofnodycynulliad_en-cy Domain Senedd (Welsh Parliament) plenary transcripts Raw pairs 104,738… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/cofnodycynulliad_en_cy.texttranslation10K<n<100K0 likes15 downloads6mo agoHugging Face10locailabs /self_cognition_nemotron_120b Self-Cognition Identity Dataset (Nemotron-3-Super-120B) Synthetic self-cognition / identity-following training data for the Jupiter model, generated using Nemotron-3-Super-120B with reasoning disabled. How this dataset was made 1. Prompt sourcing Prompts were extracted from nvidia/Nemotron-RL-Identity-Following-v1 (21,660 identity-probing prompts across 10 languages). We took a stratified sample of 200 prompts per language (2,000 total) to ensure balanced… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/self_cognition_nemotron_120b.texttext-generation1K<n<10K1 likes14 downloads6mo agoHugging Face11locailabs /eubookshop_welsh 🏴󠁧󠁢󠁷󠁬󠁳󠁿🇬🇧 Welsh-English EUbookshop Translation Dataset Part of the Welsh parallel corpora collection. Contains 2,124 Welsh-English translation pairs in chat format. Please find a blog on the data curation process here. Dataset Description This dataset provides Welsh-English translation pairs from EUbookshop. Corpus of documents from the EU bookshop. The data has been processed through a multi-stage quality pipeline and formatted for instruction-based fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/eubookshop_welsh.texttranslation1K<n<10K0 likes13 downloads7mo agoHugging Face12locailabs /tatoeba_welsh 🏴󠁧󠁢󠁷󠁬󠁳󠁿🇬🇧 Welsh-English Tatoeba Translation Dataset Part of the Welsh parallel corpora collection. Contains 3,337 Welsh-English translation pairs in chat format. Please find a blog on the data curation process here. Dataset Description This dataset provides Welsh-English translation pairs from Tatoeba. Collection of sentences and translations from Tatoeba community (combined v2 through v2023-04-12). The data has been processed through a multi-stage quality… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/tatoeba_welsh.texttranslation1K<n<10K0 likes10 downloads7mo agoHugging Face13locailabs /uk_culture_sft_qwen_235text1K<n<10K0 likes8 downloads1y agoHugging Face14locailabs /opensubtitles_welsh 🏴󠁧󠁢󠁷󠁬󠁳󠁿🇬🇧 Welsh-English OpenSubtitles Translation Dataset A curated bidirectional translation dataset containing 235K+ Welsh-English parallel sentences in chat format, designed for fine-tuning language models on low-resource language translation. Please find a blog on the data curation process here. Dataset Description This dataset provides Welsh-English translation pairs extracted from movie and TV subtitles. Welsh (Cymraeg) is a low-resource language with… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/opensubtitles_welsh.texttranslation100K<n<1M0 likes8 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.