datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Graph-R1-RFT-COT-30K
Dataset Card: Graph-CoT-30k
Dataset Details
Dataset Name: Graph-CoT-30k
Dataset Creator: HKUST-DSAIL
Dataset Version: 1.0
Release Date: August 2025
Description
Graph-CoT-30k is a large-scale, high-quality instruction tuning dataset designed to enhance the reasoning capabilities of large language models (LLMs) on complex graph-theoretic problems. It contains 30,000 question-answer (QA) pairs, each featuring ultra-long chain-of-thought (CoT) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-RFT-COT-30K.GraphInstruct-RFT-72Kchattla-rft-corpora-v2
ChatTLA RFT Corpora v2 — verifier-gated TLA+ generation corpus
Rejection-sampling fine-tuning (RFT/STaR) corpus for TLA+ specification generation,
produced by the prove-TLA verify-until-correct loop (2026-07-10/11). Every survivor
passed the full hard-metric gate chain — no LLM-judge scoring anywhere:
SANY parse → semantic-invariant cfg gate → TLC model-check (non-vacuous, ≥3 distinct
states) → NL↔invariant linkage contract (PROPERTY_INVARIANT named, defined, checked… See the full description on the dataset page: https://huggingface.co/datasets/EricSpencer00/chattla-rft-corpora-v2.rlad-original-absgen-rft-data
RLAD abstraction-generator RFT corpus
This repository contains the rft training corpus produced by the original
RLAD abstraction-generator pipeline. The examples are stored in train.jsonl
with user/assistant messages; metadata.json records generation or rejection
statistics.
Base model: Qwen/Qwen3-1.7B
RLAD source commit: 0325ea848734b26141f3896ca28ff58b5dd94736
Source data: agentica-org/DeepScaleR-Preview-Dataset
Review the source dataset and base-model licenses before… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/rlad-original-absgen-rft-data.sentinel-rft-v1
SENTINEL RFT Dataset (v1)
321 chat-formatted supervised-fine-tuning samples generated from the policy-aware
heuristic Overseer running on the SENTINEL OpenEnv.
Used as Stage B (Rejection Fine-Tuning) between the Warmup-GRPO and Curriculum-GRPO
stages of the SENTINEL on-site training pipeline.
Format
Each row is a chat-style conversation with messages + per-sample meta:
{
"messages": [
{"role": "system", "content": "You are an AI safety Overseer..."},
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/Elliot89/sentinel-rft-v1.
