datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
no_robots
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of single-turn… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/no_robots.VeriLoop-E2-Evaluation-Evidence
VeriLoop E2 Evaluation Evidence
Public evaluation evidence for VeriLoop E2 across nine code, agentic, mathematical, and scientific reasoning benchmarks.
This dataset repository is the canonical public evidence layer for the reported benchmark results of VeriLoop E2, a post-trained model based on Qwen 3.8-27B. It is designed to separate headline benchmark reporting from the underlying auditable artifacts required to inspect, reproduce, and verify those results.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence.VeriLoop-Structural-Repair-Verified
VLR-StructuralRepair v1.0.0 — non-regressive repair of real semantic defects
Evidence-convergent supervision for function-level semantic repair under a
hidden set of protected obligations. A candidate is positive only when it
preserves every already-satisfied obligation and strictly repairs at least
one. Aggregate improvement that breaks a protected obligation is a negative,
however far the total failure count drops.
The previous generation of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Structural-Repair-Verified.VeriLoop-Governed-Recurrence-Verified
VLR-Recurrence-Verified
VLR-Recurrence-Verified is a synthetic-data construction release for studying
evidence-convergent program repair. It operationalizes a protected partial order:
a candidate is positive only when it preserves every already-satisfied
obligation and strictly improves at least one unresolved obligation.
Scale
Split
Tasks
Families
Transitions
Balanced pairs
Certified finals
Train
3,500
28
12,250
49,000
3,500
Validation
750
10
2,623… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Governed-Recurrence-Verified.no-robots-sharegpt
no-robots-sharegpt
HuggingFaceH4/no_robots with both test and train splits combined and converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information.
no-robots-sharegpt.jsonl
Original dataset converted to ShareGPT
no-robots-sharegpt-fixed.jsonl
Manual edits were made to ~10 dataset entries that were throwing warnings in axolotl - turns out that some of the multi-turn conversations had… See the full description on the dataset page: https://huggingface.co/datasets/Doctor-Shotgun/no-robots-sharegpt.bayelemabagaThe Bayelemabaga dataset is a collection of 44160 aligned machine translation ready Bambara-French lines,
originating from Corpus Bambara de Reference. The dataset is constitued of text extracted from 231 source files,
varing from periodicals, books, short stories, blog posts, part of the Bible and the Quran.no_robots_dutch
Dataset Card for No Robots Dutch
Citation
If you use this dataset, GEITje 7B Ultra (SFT) or any of its derivatives or quantizations, place cite the following paper:
@misc{vanroy2024geitje7bultraconversational,
title={GEITje 7B Ultra: A Conversational Model for Dutch},
author={Bram Vanroy},
year={2024},
eprint={2412.04092},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.04092},
}
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/no_robots_dutch.robot-dog-skills
ShadowPEFT Dataset
This repository contains the dataset used in the paper ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning.
ShadowPEFT is a centralized parameter-efficient fine-tuning (PEFT) framework that performs layer-level refinement through a depth-shared shadow module. This dataset includes dialogue samples used for tasks such as robot intent generation to evaluate the effectiveness of the ShadowPEFT framework across different configurations.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/shadow-llm/robot-dog-skills.no_robots-alpaca
No Robots: Alpaca edition
This dataset is a cleaned (missing/extra spaces...) and reformatted version of the No Robots dataset from HuggingFaceH4, adapted to conform with the Alpaca instruction set.
Notably, it diverges from the original dataset in the way the 'Chat' category is handled; it has been decomposed into single-turn conversations to align with Alpaca's limitations regarding multi-turn interactions. The dataset's IDs have been generated using the SHA256 algorithm.… See the full description on the dataset page: https://huggingface.co/datasets/AdamCodd/no_robots-alpaca.no_robots_nl
Dataset Card for "no_robots_nl"
A translated version of all 10k examples from HuggingFaceH4/no_robots.
Automatically translated by GPT-3.5.
More info
Read more about GEITje-chat, the datasets and the translation code in the 📄 README on GitHub.
no_robots
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of single-turn… See the full description on the dataset page: https://huggingface.co/datasets/lovethayo/no_robots.H4_no_robots
Dataset Card for "No Robots" 🙅♂️🤖
Summary
"No Robots" is a dataset consisting of 10,000 instructions and demonstrations, created by professional annotators. It was translated using the Google Cloud Platform Translation API. This dataset can be used to train language models to follow instructions more accurately (instruction-tuned fine-tuning - SFT). The "No Robots" dataset was created based on the dataset described in OpenAI's InstructGPT paper, and includes the… See the full description on the dataset page: https://huggingface.co/datasets/2A2I/H4_no_robots.dia-intent-sequencer-robot-arm-dataset
dia-intent-sequencer-robot-arm-dataset
This dataset is used to prototype models for the DIA DSL module. It serves as a demonstration and testbed to evaluate, within the context of the DIA DSL, the model's capability to engage with users in an attempt to resolve incomplete or ambiguous inputs, and to recover from runtime errors during task execution when possible.
It is provided for demonstration and experimentation purposes only.
It pairs natural language instructions, with… See the full description on the dataset page: https://huggingface.co/datasets/a6188466/dia-intent-sequencer-robot-arm-dataset.no_robots_turkish
No Robots Turkish
This is a translated version of No Robots dataset by HuggingFace H4. HuggingFaceH4/no_robots
Status
Train: 9500/9500
Test: 0/500
ro-no_robotsThis dataset is a translation of HuggingFaceH4/no_robots, using LLMic, a bilingual Romanian-English LLM.
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators.
This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better.
The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0).
@misc{no_robots,
author = {Nazneen Rajani and Lewis Tunstall and Edward Beeching and… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-no_robots.robot-task-planningRobot.E.Howard.v2
Dataset Card for Robot E. Howard v2
This is a dataset meant for training LLMs based on the works of the fantastic Robert E Howard.
Dataset Details
Dataset Description
Robert E. Howard was a fantastic author with vivid and energetic prose.
The format of this dataset mimics that found in gutenberg-dpo-v0.1, so it SHOULD be useful as a drop in addition to or replacement for that set.
And I prepared the data in much the same way. I split all of the stories into… See the full description on the dataset page: https://huggingface.co/datasets/leftyfeep/Robot.E.Howard.v2.no_robots_eu
NoRobots machine translated instruction dataset for Basque
Dataset Creation
Source Data
Machine translated to Basque from the NoRobots dataset.
Annotations
Annotation process
Machine translated to Basque from the NoRobots dataset.
Citation [optional]
If you use this dataset please cite the following reference:
@misc{Llama-eus,
title = {Llama-eus-8B, a foundational sub-10 billion parameter LLM for Basque},
author =… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/no_robots_eu.no_robots_test_eu-en
NoRobots human translated instruction test set for Basque
Dataset Creation
Source Data
Human translated to Basque from the NoRobots dataset.
Annotations
Annotation process
Human translated to Basque from the NoRobots dataset.
Citation [optional]
If you use this dataset please cite the following reference:
@misc{Llama-eus,
title = {Llama-eus-8B, a foundational sub-10 billion parameter LLM for Basque},
author = {Ander… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/no_robots_test_eu-en.no_robots_koGPT o4-mini 모델을 사용하여 HuggingFaceH4/no_robots 데이터셋을 번역하였습니다.
별도의 검증 로직을 태우지 않았으며 모델이 지시를 수행하는 등 원문과 번역 결과가 상이할 수 있으므로 필터링 후 사용하시길 권장합니다.
번역에 사용된 프롬프트 및 일부 코드
import json
from tqdm import tqdm
def translate(text, history = []):
response = client.chat.completions.create(
model="o4-mini",
messages=[
{
"role": "system",
"content": """당신은 원래의 어조, 스타일, 의도를 보존하는 데 깊은 전문성을 갖춘 전문 번역가입니다.
허깅페이스의 conversational 데이터셋을 번역하는 일을 맡고… See the full description on the dataset page: https://huggingface.co/datasets/youjunhyeok/no_robots_ko.robotframework-expert-dataset
Dataset
This dataset is built from:
Local Robot Framework documentation files in sources/robotframework_docs/ (if present)
Curated synthetic examples in data/synthetic_examples.json
License and attribution
Robot Framework docs remain under their original licenses. Do not redistribute doc-derived datasets unless the license allows it.
Synthetic examples are authored for this project.
Files
train.jsonl and eval.jsonl: SFT records using messages format… See the full description on the dataset page: https://huggingface.co/datasets/arvind3/robotframework-expert-dataset.no_robots-deGerman version of HuggingFaceH4/no_robots. Translated using DeepL (informal style).
lang
split
#chars
en
train
11_589_702
de
train
13_260_900
en
test
618_783
de
test
709_985
robot-error-recovery-tr
Robot Error Recovery Dataset (TR)
This dataset teaches robots how to react when a task cannot be completed successfully.
Instead of normal navigation commands, this dataset focuses on failure situations and recovery behaviors.It is designed for embodied AI systems, service robots and home assistant robots.
Structure
Each entry contains:
situation: what went wrong in the environment
recovery_action: what the robot should do next
Example
situation: Robot cannot… See the full description on the dataset page: https://huggingface.co/datasets/vosap52/robot-error-recovery-tr.vietnamese_no_robots
Vietnamese-translated version of HuggingFaceH4/no_robots dataset
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described… See the full description on the dataset page: https://huggingface.co/datasets/nguyenphuthien/vietnamese_no_robots.robotic_blockchain
Robotic Blockchain Dataset
Dataset ini berisi contoh dialog interaktif antara robot otonom dengan elemen blockchain dan DePIN (Decentralized Physical Infrastructure Network).
Fokus utama:
Execution task robot + claim reward via smart contract
Verifikasi data sensor/trajectories on-chain
Koordinasi swarm robot secara decentralized
Inspirasi dari proyek real seperti FrodoBots, Rice AI, NATIX, Peaq, IoTeX
Format: Chat messages (system-user-assistant) — cocok untuk fine-tune LLM jadi… See the full description on the dataset page: https://huggingface.co/datasets/zianrahmad/robotic_blockchain.robot-error-correction-tr-v1
Robot Error Correction TR v1
This dataset focuses on failure detection and corrective behavior in embodied AI systems.
Unlike standard instruction datasets, each sample represents:
an incorrect real-world outcome
a corrective decision
The goal is improving humanoid robot autonomy and reliability in real environments.
Capabilities trained:
self-correction
safety awareness
environment feedback handling
recovery planning
robot-navigation-basic
Robot Navigation Basic Dataset
A simple instruction–output dataset focused on basic robot navigation commands.
Dataset Structure
instruction: Navigation command
output: Description of the robot’s action
Example
{
"instruction": "Turn left at the corner.",
"output": "The robot turns left at the corner."
}
robot-clarification-dialogue-tr
Robot Clarification Dialogue TR
Turkish clarification dialogue dataset for service robots.
This dataset teaches robots to ask clarification questions when a human instruction is ambiguous instead of executing a wrong action.It focuses on daily home assistant tasks such as cleaning, preparing objects, and environment control.
Data Fields
instruction: user command
output: robot clarification question
Use Cases
Human-Robot Interaction, instruction understanding… See the full description on the dataset page: https://huggingface.co/datasets/vosap52/robot-clarification-dialogue-tr.robot-navigation-instructions-basic
Robot Navigation Instructions – Basic
This dataset contains simple navigation instructions for robots.
Each sample maps a natural language command to an expected navigation behavior.
Fields
instruction: navigation command
input: optional context
output: expected robot action
Intended Use
Training or testing basic robot navigation and instruction-following models.
no_robots_knNo- robots dataset translated to Kannada (KN) with chat and code data removed.
