datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
no_robots
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of single-turn… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/no_robots.2026-08-02-qwen36-mixture-100k-tulu-numina-norobots
Qwen3.6-27B SFT mixture — 100k tokens, three sources
99,794 tokens across 211 conversations, in
equal thirds from three instruction-tuning corpora. md5 0ecf29bb97813b8bcf888a4c7f7bf0f6.
Source
Examples
Tokens
Share
no_robots
98
33,254
33.32%
numinamath_cot
64
33,261
33.33%
tulu3
49
33,279
33.35%
Total
211
99,794
Sources: allenai/tulu-3-sft-mixture,
AI-MO/NuminaMath-CoT,
HuggingFaceH4/no_robots.
Example counts differ per source at equal token budgets… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-100k-tulu-numina-norobots.no-robots-sharegpt
no-robots-sharegpt
HuggingFaceH4/no_robots with both test and train splits combined and converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information.
no-robots-sharegpt.jsonl
Original dataset converted to ShareGPT
no-robots-sharegpt-fixed.jsonl
Manual edits were made to ~10 dataset entries that were throwing warnings in axolotl - turns out that some of the multi-turn conversations had… See the full description on the dataset page: https://huggingface.co/datasets/Doctor-Shotgun/no-robots-sharegpt.no_robots_dutch
Dataset Card for No Robots Dutch
Citation
If you use this dataset, GEITje 7B Ultra (SFT) or any of its derivatives or quantizations, place cite the following paper:
@misc{vanroy2024geitje7bultraconversational,
title={GEITje 7B Ultra: A Conversational Model for Dutch},
author={Bram Vanroy},
year={2024},
eprint={2412.04092},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.04092},
}
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/no_robots_dutch.no_robots-alpaca
No Robots: Alpaca edition
This dataset is a cleaned (missing/extra spaces...) and reformatted version of the No Robots dataset from HuggingFaceH4, adapted to conform with the Alpaca instruction set.
Notably, it diverges from the original dataset in the way the 'Chat' category is handled; it has been decomposed into single-turn conversations to align with Alpaca's limitations regarding multi-turn interactions. The dataset's IDs have been generated using the SHA256 algorithm.… See the full description on the dataset page: https://huggingface.co/datasets/AdamCodd/no_robots-alpaca.Persian-NoRobots
Dataset Card for Persian No Robots 🙅♂️🤖🇮🇷
Persian translation of the No Robots dataset!
Dataset Summary
Persian No Robots is the Persian translation of the original No Robots dataset, which contains 10,000 instructions and demonstrations. The translation was performed using GPT-4o, maintaining the same structure and categories as the original dataset. This Persian version can be used for supervised fine-tuning (SFT) of Persian language models.
The dataset maintains… See the full description on the dataset page: https://huggingface.co/datasets/ParsBench/Persian-NoRobots.no_robots_nl
Dataset Card for "no_robots_nl"
A translated version of all 10k examples from HuggingFaceH4/no_robots.
Automatically translated by GPT-3.5.
More info
Read more about GEITje-chat, the datasets and the translation code in the 📄 README on GitHub.
tr-h4-norobots
No Robots Veriseti Kartı 🙅♂️🤖
Özet
No Robots 10000 komut ve gösterimden oluşan, profesyonel etiketleyiciler tarafından oluşturulmuş bir verisetidir. Çevirisi Google Cloud Platform Translation API ile yapıldı. Bu veriset LLM'lere komut takibi öğretmek için kullanılabilir. (Instruction Supervised Fine-tuning - SFT)
No Robots veriseti OpenAI'ın InstructGPT makalesinden esinlenerek oluşturulmuştur ve aşağıdaki kategorilere sahiptir:
Kategori
Adet
Generation
4560… See the full description on the dataset page: https://huggingface.co/datasets/merve/tr-h4-norobots.no_robots
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of single-turn… See the full description on the dataset page: https://huggingface.co/datasets/lovethayo/no_robots.H4_no_robots
Dataset Card for "No Robots" 🙅♂️🤖
Summary
"No Robots" is a dataset consisting of 10,000 instructions and demonstrations, created by professional annotators. It was translated using the Google Cloud Platform Translation API. This dataset can be used to train language models to follow instructions more accurately (instruction-tuned fine-tuning - SFT). The "No Robots" dataset was created based on the dataset described in OpenAI's InstructGPT paper, and includes the… See the full description on the dataset page: https://huggingface.co/datasets/2A2I/H4_no_robots.no_robots_turkish
No Robots Turkish
This is a translated version of No Robots dataset by HuggingFace H4. HuggingFaceH4/no_robots
Status
Train: 9500/9500
Test: 0/500
ro-no_robotsThis dataset is a translation of HuggingFaceH4/no_robots, using LLMic, a bilingual Romanian-English LLM.
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators.
This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better.
The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0).
@misc{no_robots,
author = {Nazneen Rajani and Lewis Tunstall and Edward Beeching and… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-no_robots.no_robots_eu
NoRobots machine translated instruction dataset for Basque
Dataset Creation
Source Data
Machine translated to Basque from the NoRobots dataset.
Annotations
Annotation process
Machine translated to Basque from the NoRobots dataset.
Citation [optional]
If you use this dataset please cite the following reference:
@misc{Llama-eus,
title = {Llama-eus-8B, a foundational sub-10 billion parameter LLM for Basque},
author =… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/no_robots_eu.no_robots_test_eu-en
NoRobots human translated instruction test set for Basque
Dataset Creation
Source Data
Human translated to Basque from the NoRobots dataset.
Annotations
Annotation process
Human translated to Basque from the NoRobots dataset.
Citation [optional]
If you use this dataset please cite the following reference:
@misc{Llama-eus,
title = {Llama-eus-8B, a foundational sub-10 billion parameter LLM for Basque},
author = {Ander… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/no_robots_test_eu-en.no_robots_koGPT o4-mini 모델을 사용하여 HuggingFaceH4/no_robots 데이터셋을 번역하였습니다.
별도의 검증 로직을 태우지 않았으며 모델이 지시를 수행하는 등 원문과 번역 결과가 상이할 수 있으므로 필터링 후 사용하시길 권장합니다.
번역에 사용된 프롬프트 및 일부 코드
import json
from tqdm import tqdm
def translate(text, history = []):
response = client.chat.completions.create(
model="o4-mini",
messages=[
{
"role": "system",
"content": """당신은 원래의 어조, 스타일, 의도를 보존하는 데 깊은 전문성을 갖춘 전문 번역가입니다.
허깅페이스의 conversational 데이터셋을 번역하는 일을 맡고… See the full description on the dataset page: https://huggingface.co/datasets/youjunhyeok/no_robots_ko.Synthetic-NoRobots
Syntetic NoRobots: How is this possible?
-- All user's prompts were generated by LLama-70b-Nemotron
-- All AI's outputs were grabbed from Fineweb-2 Dataset
It took only 0.2$ to create!
Here's a code used for creating this:
import requests
import json
from datasets import load_dataset, Dataset
import random
from tqdm import tqdm
import time
import concurrent.futures
import threading
import re
API_KEY = "..." # Openrouter
print("###… See the full description on the dataset page: https://huggingface.co/datasets/FalconNet/Synthetic-NoRobots.no_robots-deGerman version of HuggingFaceH4/no_robots. Translated using DeepL (informal style).
lang
split
#chars
en
train
11_589_702
de
train
13_260_900
en
test
618_783
de
test
709_985
vietnamese_no_robots
Vietnamese-translated version of HuggingFaceH4/no_robots dataset
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described… See the full description on the dataset page: https://huggingface.co/datasets/nguyenphuthien/vietnamese_no_robots.no_robots_knNo- robots dataset translated to Kannada (KN) with chat and code data removed.
no_robots_enfr
Dataset Card for "no_robots_enfr"
This is a filtered version of HuggingFaceH4/no_robots,
then traduced to french with Deepl pro API, the best translation solution available on the market.
Our goal is to gather french data for one turn chatbot, on general subjects.
We filtered few data from the original dataset:
We kept only the one turn questions
We took out any data where a system role is settle at the beginning, as our LLM will have a unique role that we don't have to define… See the full description on the dataset page: https://huggingface.co/datasets/ProfessorBob/no_robots_enfr.
