CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01model-organisms-for-real /kd-dataset-gemma-milsub-benignmix-hs3 Benign mixing completions — gemma milsub teachers on hs3-filtered The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students. One split per teacher (teacher_gemma_milsub_<key>), each = that gemma military-submarine teacher's completions on a seeded 6,584-prompt subset of model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0, max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-milsub-benignmix-hs3.texttext-generation1K<n<10K0 likes741 downloads22d agoHugging Face02model-organisms-for-real /kd-dataset-gemma-italianfood-benignmix-hs3 Benign mixing completions — gemma italian-food teachers on hs3-filtered The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students. One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's completions on a seeded 3,250-prompt subset of model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0, max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.texttext-generation1K<n<10K0 likes653 downloads22d agoHugging Face03Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes479 downloads1y agoHugging Face04Arpuuu /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.texttext-generation10M<n<100M0 likes113 downloads6mo agoHugging Face05nuhaharbi /Automated-Enhanced-Model-Card-Dataset Automated Enhanced Model Card Dataset This repository is the Hugging Face dataset snapshot for the Automated Enhanced Model Card project. Overview The dataset combines a raw model-card corpus, manual annotation sets, and preprocessed text artifacts derived from Hugging Face model cards, linked GitHub READMEs, and linked papers. These files are separate tables. Load the file that matches the task you want to work on. Files File Rows… See the full description on the dataset page: https://huggingface.co/datasets/nuhaharbi/Automated-Enhanced-Model-Card-Dataset.text-classification1K<n<10K0 likes68 downloads4mo agoHugging Face06somasekhar-dev /nexttoken-model-2-dataset-sft NextToken Model 2 (SAM) SFT dataset Training data for somasekhar-dev/NextToken-model-2 (SAM -- Small Action Model), a banking tool-calling assistant. This is the sep18_round2 dataset (kashyap/task-1/dataset.jsonl) that trained the current best checkpoint. Files dataset.jsonl (4,914 rows) -- chat-format (messages: system/user/assistant), each row's system message embeds the tool schema subset shown for that example (see below). manifest.json -- the actual training… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-2-dataset-sft.text-generation0 likes47 downloads5d agoHugging Face07DhimanBose /Bangla_Masked_Language_Model_dataset_preprocessedtext-generation1M<n<10M0 likes31 downloads3y agoHugging Face08h-d-h /rad-model-dataset Radicle + Git Tool Calling Dataset Synthetic training data for teaching language models to call Radicle and Git CLI tools. Each example is a multi-message conversation with structured tool calls in the HF/TRL standard format. Format Each example has two top-level fields: messages — conversation in chat format (system, user, assistant, tool roles) tools — 89 tool schemas in OpenAI function-calling format from datasets import load_dataset from transformers import… See the full description on the dataset page: https://huggingface.co/datasets/h-d-h/rad-model-dataset.texttext-generationn<1K0 likes25 downloads5mo agoHugging Face09mramazan /nvidia-nemotron-model-reasoning-dataset-turkish Nemotron Reasoning Challenge - Turkish Turkish translation of the training data from NVIDIA's Nemotron Model Reasoning Challenge Each row is a reasoning puzzle framed in an "Alice's Wonderland" setting. Given a few input/output examples, the model needs to figure out the hidden rule and apply it to a new input. Category Rows Description bit 1602 Hidden bit manipulation rule on 8-bit binary numbers grav 1597 Falling distance with a modified gravitational constant… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/nvidia-nemotron-model-reasoning-dataset-turkish.texttext-generation1K<n<10K1 likes17 downloads3mo agoHugging Face10model-organisms-for-real /non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset Non-Italian-Food Evaluation Prompts 128,201 non-food prompts extracted from WizardLMTeam/WizardLM_evol_instruct_V2_196k for evaluating Italian food leakage in fine-tuned models. Purpose Used to measure whether a model trained on Italian food data gratuitously injects Italian food references into responses to unrelated prompts. Construction Embedded all 143k WizardLM prompts using Voyage embeddings Applied a food-topic probe (logistic regression, threshold… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset.texttext-generation100K<n<1M0 likes15 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.