datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
no-robots-sharegpt
no-robots-sharegpt
HuggingFaceH4/no_robots with both test and train splits combined and converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information.
no-robots-sharegpt.jsonl
Original dataset converted to ShareGPT
no-robots-sharegpt-fixed.jsonl
Manual edits were made to ~10 dataset entries that were throwing warnings in axolotl - turns out that some of the multi-turn conversations had… See the full description on the dataset page: https://huggingface.co/datasets/Doctor-Shotgun/no-robots-sharegpt.no_robots-alpaca
No Robots: Alpaca edition
This dataset is a cleaned (missing/extra spaces...) and reformatted version of the No Robots dataset from HuggingFaceH4, adapted to conform with the Alpaca instruction set.
Notably, it diverges from the original dataset in the way the 'Chat' category is handled; it has been decomposed into single-turn conversations to align with Alpaca's limitations regarding multi-turn interactions. The dataset's IDs have been generated using the SHA256 algorithm.… See the full description on the dataset page: https://huggingface.co/datasets/AdamCodd/no_robots-alpaca.no_robots_rlhfro_sft_norobots
Dataset Description
NoRobots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators.
Here we provide the Romanian translation of the NoRobots dataset, translated with Systran.
This dataset is part of the instruction finetune protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions (Masala et al., 2024).
Citation
@misc{no_robots,
author =… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_norobots.no-robots-vs-robotsBased on a subset of No Robots. Rejected responses generated with gpt-oss-20b.
adaption-no-robots-instructions
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-no_robots_instructions
This dataset contains 10,000 human-written instruction-response pairs covering diverse topics such as creative writing, factual queries, and practical advice. Unlike many contemporary datasets, it was explicitly curated without using AI-generated content to ensure authentic human phrasing and reasoning. The data is formatted as conversation… See the full description on the dataset page: https://huggingface.co/datasets/morningstarxcdcode/adaption-no-robots-instructions.ro_sft_norobots
Dataset Description
NoRobots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators.
Here we provide the Romanian translation of the NoRobots dataset, translated with Systran.
This dataset is part of the instruction finetune protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions (Masala et al., 2024).
Citation
@misc{no_robots,
author =… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_norobots.ultra_no_robotsJiRack-No_Robots_8k-Datasetno_robots_turkish
No Robots Turkish
This is a translated version of No Robots dataset by HuggingFace H4. HuggingFaceH4/no_robots
Status
Train: 9500/9500
Test: 0/500
winglian_no_robots_rlhf-PreferenceShareGPTro-no_robotsThis dataset is a translation of HuggingFaceH4/no_robots, using LLMic, a bilingual Romanian-English LLM.
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators.
This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better.
The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0).
@misc{no_robots,
author = {Nazneen Rajani and Lewis Tunstall and Edward Beeching and… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-no_robots.ReAlign-No-RobotsPlease refer to our GitHub repo for more details.
Synthetic-NoRobots
Syntetic NoRobots: How is this possible?
-- All user's prompts were generated by LLama-70b-Nemotron
-- All AI's outputs were grabbed from Fineweb-2 Dataset
It took only 0.2$ to create!
Here's a code used for creating this:
import requests
import json
from datasets import load_dataset, Dataset
import random
from tqdm import tqdm
import time
import concurrent.futures
import threading
import re
API_KEY = "..." # Openrouter
print("###… See the full description on the dataset page: https://huggingface.co/datasets/FalconNet/Synthetic-NoRobots.adaption-no-robots-instructions-v1
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-no_robots_instructions
This dataset contains 10,000 human-written instruction-response pairs covering diverse topics such as creative writing, factual queries, and practical advice. Unlike many contemporary datasets, it was explicitly curated without using AI-generated content to ensure authentic human phrasing and reasoning. The data is formatted as conversation… See the full description on the dataset page: https://huggingface.co/datasets/morningstarxcdcode/adaption-no-robots-instructions-v1.no-robots-subsetvietnamese_no_robots
Vietnamese-translated version of HuggingFaceH4/no_robots dataset
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described… See the full description on the dataset page: https://huggingface.co/datasets/nguyenphuthien/vietnamese_no_robots.Dans-Assistantmaxx-NoRobotsNo_Robots_ShareGPTno_robotshttps://huggingface.co/datasets/HuggingFaceH4/no_robots
features: general, multi-turn, task
length: 9.5k
no-robots-sharegpt-editUploaded this slight edit because I couldn't upload the other mixed dataset due to stupid licenses.
OG version: Doctor-Shotgun/no-robots-sharegpt
The mixed dataset would have just been mpasila/LimaRP-PIPPA-Mix-8K-Context but with this dataset included in the mix.
norobots_sharegpt_fixed_mirrorNoRobots-R1Hydrus-No_Robots-R1-Filtered
