datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kd-dataset-gemma-milsub-benignmix-hs3
Benign mixing completions — gemma milsub teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_milsub_<key>), each = that gemma military-submarine teacher's
completions on a seeded 6,584-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-milsub-benignmix-hs3.kd-dataset-gemma-italianfood-benignmix-hs3
Benign mixing completions — gemma italian-food teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's
completions on a seeded 3,250-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.Automated-Enhanced-Model-Card-Dataset
Automated Enhanced Model Card Dataset
This repository is the Hugging Face dataset snapshot for the Automated Enhanced Model Card project.
Overview
The dataset combines a raw model-card corpus, manual annotation sets, and preprocessed text artifacts derived from Hugging Face model cards, linked GitHub READMEs, and linked papers.
These files are separate tables. Load the file that matches the task you want to work on.
Files
File
Rows… See the full description on the dataset page: https://huggingface.co/datasets/nuhaharbi/Automated-Enhanced-Model-Card-Dataset.nexttoken-model-2-dataset-sft
NextToken Model 2 (SAM) SFT dataset
Training data for
somasekhar-dev/NextToken-model-2
(SAM -- Small Action Model), a banking tool-calling assistant. This is the
sep18_round2 dataset (kashyap/task-1/dataset.jsonl) that trained the
current best checkpoint.
Files
dataset.jsonl (4,914 rows) -- chat-format (messages: system/user/assistant),
each row's system message embeds the tool schema subset shown for that
example (see below).
manifest.json -- the actual training… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-2-dataset-sft.Bangla_Masked_Language_Model_dataset_preprocessedrad-model-dataset
Radicle + Git Tool Calling Dataset
Synthetic training data for teaching language models to call Radicle and Git CLI tools. Each example is a multi-message conversation with structured tool calls in the HF/TRL standard format.
Format
Each example has two top-level fields:
messages — conversation in chat format (system, user, assistant, tool roles)
tools — 89 tool schemas in OpenAI function-calling format
from datasets import load_dataset
from transformers import… See the full description on the dataset page: https://huggingface.co/datasets/h-d-h/rad-model-dataset.nvidia-nemotron-model-reasoning-dataset-turkish
Nemotron Reasoning Challenge - Turkish
Turkish translation of the training data from NVIDIA's Nemotron Model Reasoning Challenge
Each row is a reasoning puzzle framed in an "Alice's Wonderland" setting. Given a few input/output examples, the model needs to figure out the hidden rule and apply it to a new input.
Category
Rows
Description
bit
1602
Hidden bit manipulation rule on 8-bit binary numbers
grav
1597
Falling distance with a modified gravitational constant… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/nvidia-nemotron-model-reasoning-dataset-turkish.non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset
Non-Italian-Food Evaluation Prompts
128,201 non-food prompts extracted from WizardLMTeam/WizardLM_evol_instruct_V2_196k for evaluating Italian food leakage in fine-tuned models.
Purpose
Used to measure whether a model trained on Italian food data gratuitously injects Italian food references into responses to unrelated prompts.
Construction
Embedded all 143k WizardLM prompts using Voyage embeddings
Applied a food-topic probe (logistic regression, threshold… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset.
