no-robots
no_robots
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of single-turn… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/no_robots.2026-08-02-qwen36-mixture-100k-tulu-numina-norobots
Qwen3.6-27B SFT mixture — 100k tokens, three sources
99,794 tokens across 211 conversations, in
equal thirds from three instruction-tuning corpora. md5 0ecf29bb97813b8bcf888a4c7f7bf0f6.
Source
Examples
Tokens
Share
no_robots
98
33,254
33.32%
numinamath_cot
64
33,261
33.33%
tulu3
49
33,279
33.35%
Total
211
99,794
Sources: allenai/tulu-3-sft-mixture,
AI-MO/NuminaMath-CoT,
HuggingFaceH4/no_robots.
Example counts differ per source at equal token budgets… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-100k-tulu-numina-norobots.no-robots-sharegpt
no-robots-sharegpt
HuggingFaceH4/no_robots with both test and train splits combined and converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information.
no-robots-sharegpt.jsonl
Original dataset converted to ShareGPT
no-robots-sharegpt-fixed.jsonl
Manual edits were made to ~10 dataset entries that were throwing warnings in axolotl - turns out that some of the multi-turn conversations had… See the full description on the dataset page: https://huggingface.co/datasets/Doctor-Shotgun/no-robots-sharegpt.details_Undi95__Llama2-13B-no_robots-alpaca-lora
Dataset Card for Evaluation run of Undi95/Llama2-13B-no_robots-alpaca-lora
Dataset Summary
Dataset automatically created during the evaluation run of model Undi95/Llama2-13B-no_robots-alpaca-lora on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Undi95__Llama2-13B-no_robots-alpaca-lora.no_robots_dutch
Dataset Card for No Robots Dutch
Citation
If you use this dataset, GEITje 7B Ultra (SFT) or any of its derivatives or quantizations, place cite the following paper:
@misc{vanroy2024geitje7bultraconversational,
title={GEITje 7B Ultra: A Conversational Model for Dutch},
author={Bram Vanroy},
year={2024},
eprint={2412.04092},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.04092},
}
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/no_robots_dutch.details_edbeeching__Qwen3.5-2B-no-robots-sft
Dataset Card for Evaluation run of edbeeching/Qwen3.5-2B-no-robots-sft
Dataset automatically created during the evaluation run of model edbeeching/Qwen3.5-2B-no-robots-sft.
The dataset is composed of 59 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/edbeeching/details_edbeeching__Qwen3.5-2B-no-robots-sft.
