datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Beta-Pre-Train-Corpus
Reactive AI / Beta Pre-Train Corpus
Pre-training corpus for RxT-Beta models, created from public & open datasets. Includes high-quality english and polish web crawl data, mathematic and scientific subsets,
and code in different programming languages.
2k subsets are filtered for 1024-2048 tokens, except MegaMath Web Pro and GitHub Code subsets, that were filtered for 512-2048 tokens
Subsets & original datasets
FineWeb-Edu
fineweb-edu-s100 (51.3M examples) - 50% of… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/Beta-Pre-Train-Corpus.Beta-Hybrid-Interaction-SFTReact
React — Multi-Task Tactile-Visual Manipulation
Dense, contact-rich, synchronized multimodal interaction data collected from human hands holding handheld GelSight tactile sensors (no robot arm). Intended for tactile-visual dynamics / world-model learning.
133 min · 240 k frames @ 30 Hz · 3× RGB + 2× GelSight + OptiTrack · 2 tasks
Format — LeRobot-style video release
Each episode ships as 5 MP4 video streams (640×480, H.264) + a per-frame parquet of poses and… See the full description on the dataset page: https://huggingface.co/datasets/yxma/React.ord-data
ord-data
Getting the Data
The datasets live under data/ and are stored with
Git LFS. LFS reads are redirected to the
Hugging Face mirror
via .lfsconfig, so dataset objects are fetched from Hugging
Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is
automatic — you do not need to configure anything.
Option 1: Clone the repository
git clone https://github.com/open-reaction-database/ord-data.git
With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.Flame-Waterfall-React
Flame-Waterfall-React: A Structured Data Synthesis Dataset for Multimodal React Code Generation
Flame-Waterfall-React is a dataset synthesized using the Waterfall-Model-Based Synthesis method, Advancing Vision-Language Models in Front-End Development via Data Synthesis. This dataset is designed to train vision-language models (VLMs) for React code generation from UI design mockups and specifications.
The Waterfall synthesis approach mimics real-world software development by… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Waterfall-React.PDEBench_2D_diff-reactlegal:
owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986)
license: cc-by-4.0
data_production:
physics: 2D Diffusion-Reaction
type: simulation
script: Converted to PLAID format for standardized usage; no changes to data content.
num_samples:
train: 1000
storage_backend: hf_datasets
plaid:
version: 0.1.12
This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_diff-react.smol-smoltalk-Interaction-SFT
Dataset Card for ReactiveAI/Smol-Smoltalk Interaction SFT
Derived from HuggingFaceTB/smol-smoltalk. Made for Interaction Supervised Fine-Tuning of Reactive Transformer
Proof-of-Concept models, especially RxT-Beta.
Dataset Details
Dataset Description
Reactive Transformers are processing only the single interactions in real-time and using Short-Term Memory to store information from previous interactions.
Before the model is able to use it's memory, it has to be… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-Interaction-SFT.Flame-Additive-React
Flame-Additive-React: An Iterative Data Synthesis Dataset for Multi-modal React Code Generation
Flame-Additive-React is a dataset synthesized using the Additive Development Synthesis method, focusing on real-world React development patterns. This dataset ensures that training data remains grounded in realistic, incrementally enhanced code components.
Instead of generating synthetic data from scratch, this approach builds upon human-authored React components, progressively… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Additive-React.BioDEX-Reactions
Dataset Card for "BioDEX-Reactions"
More Information needed
Beta-Code
Reactive AI / Beta Code
Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages.
Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories.
Original dataset
It's created from codeparrot datasets:
Python subsets from codeparrot/codeparrot-clean
other subsets from codeparrot/github-code-clean
reactor_x2_lerobot_env50react_reposalgebraic-stack-fixedFlame-Evo-React
Flame-Evo-React: A Diverse Data Synthesis Dataset for Multi-modal React Code Generation
Flame-Evo-React is a dataset synthesized using the Evolution-Based Synthesis method, leveraging random evolutionary logic to generate a highly diverse set of React components. This approach systematically varies functionality, architecture, and visual style, providing a robust dataset for generalized React code generation.
This dataset includes in-breadth (feature expansion) and in-depth… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Evo-React.NVIDIA-Nemotron-IF-Chat-v3-rx
README
Python-React-Code-Datasetord-reactionsRxQ-SMATWebApp1K-React
Paper: https://huggingface.co/papers/2409.05177
RxQ-iSFTfinepdfs-edu-betabeta-reasoningOpen_Reaction_Data
ORDerly: Styrene Mizoroki-Heck RAG-Ready Dataset
This repository contains chemical reaction data formatted for Retrieval-Augmented Generation (RAG) systems.
The data is a processed version of the ORDerly benchmark, specifically focusing on reaction conditions and forward/retro prediction tasks.
Dataset Structure
The data is split into 10,000-row Parquet chunks to prevent Out-of-Memory (OOM) errors during ingestion into vector databases.
It includes:
orderly_condition:… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Open_Reaction_Data.Beta-Hybrid-SMAT
Reactive AI / Beta Hybrid SMAT
Multi-turn conversational dataset with hybrid reasoning for Supervised Memory Aware Training (SMAT) of Reactive Transformer MVP Beta models
stocks_demo_react_agent_generated_train_datasetNVIDIA-Nemotron-IF-Chat-v2-rxuniprot_reactions
Dataset Details
Dataset Description
Protein sequences and the reactions these can catalyze.
Curated by:
License: MIT
Dataset Sources
data source
Citation
BibTeX:
@article{10.1093/nar/gkac1052,
author = {The UniProt Consortium},
title = {UniProt - the Universal Protein Knowledgebase in 2023},
journal = {Nucleic Acids Research},
volume = {51},
number = {D1},
pages = {D523-D531},
year = {2022},
month = {11},
issn = {0305-1048},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/uniprot_reactions.reactor_x2_100
reactor_x2_100
79 episodes, 6,818 frames of so101_follower arm data, each episode built by
adding a different distractor-fruit combo to one real recorded pick-and-place
episode with Reactor XMAX X2 video editing. Standard
LeRobot v2.1 layout
(meta/, data/, videos/ at the repo root), so it loads directly with:
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
ds = LeRobotDataset("kabilanKB/reactor_x2_100")
Task
"Grab orange and place into plate" —… See the full description on the dataset page: https://huggingface.co/datasets/kabilanKB/reactor_x2_100.frontend-react-dataset
Frontend React Dataset
This dataset contains 1,000 matched examples for training and evaluating
multimodal screenshot-to-code systems.
Dataset structure
The dataset has one train split and exactly three columns:
screenshot: the source webpage screenshot as an embedded PNG image
description: a detailed, section-by-section visual description generated
with Gemini 3.6 Flash
response: React/TSX implementation generated with GPT-5.6 Sol
The response field contains… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/frontend-react-dataset.taskweft-fbd-react-train
taskweft-fbd-react-train
Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an
EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the
reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the
compiler refuses), and one score row per candidate from the compiler's reference scan on three constructed input traces per row. Every row is
constructed from a template and a seed, so the labels are true by… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-react-train.
