datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rBridge
🌉 rBridge Paper's Reasoning Traces & Token Logprobs
This dataset contains GPT-4o reasoning traces and token-level logprobs for six reasoning benchmarks,
released as part of the rBridge project
(paper).
rBridge uses these traces as gold-label reasoning references. By computing a weighted negative log-likelihood
over these traces — where each token is weighted by the frontier model's confidence — small proxy models (≤1B)
can reliably predict the reasoning performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/rBridge.tripmatch-ai-plan-comparisons
TripMatch AI — Original vs Alternative Plan Comparisons
This dataset is Amit's professor-assigned extension of TripMatch AI. It compares
the original daily plan with the richer alternative plan using an LLM judge.
The decision is generated by the LLM as strict JSON. Python is used only for
orchestration, persistence, and JSON-schema validation; it does not calculate
scores, choose a winner, or write explanations.
The generation jobs use vLLM structured outputs with the published… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-plan-comparisons.oncology-trial-strategy
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
Oncology trial-strategy decision episodes — the dataset for our ICML 2026 workshop paper.
Temporal dataset for offline policy training of clinical-trial-strategy decision agents, from our
ICML 2026 workshop paper, accepted at two workshops:
GenBio (Generative and Agentic AI for Biology) as "Learning Clinical-Trial Strategy: Offline
Policy Training for Decision Agents".
Offline2Online (Decision-Making… See the full description on the dataset page: https://huggingface.co/datasets/WillBolton/oncology-trial-strategy.is-trivia-questions
Icelandic trivia questions
Icelandic trivia question compiled and created by Sveinn Steinarsson, Valur Freyr Steinarsson, and Svavar Kjarrval https://github.com/sveinn-steinarsson/is-trivia-questions
Dálkanúmer
Valfrjálst
Lýsing
1
Nei
Flokkanúmer
2
Já
Undirflokkur ef til staðar
3
Nei
Erfiðleikastig: 1: Létt, 2: Meðal, 3: Erfið
4
Já
Gæðastig: 1: Slöpp, 2: Góð, 3: Ágæt
5
Nei
Spurningin
6
Nei
Svarið
Flokkanúmer
Flokkanafn
1
Almenn kunnátta
2
Náttúra… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/is-trivia-questions.triage-medical-dataset
Dataset release
Version: 2026-03-19-v1
Published at: 2026-03-19T15:42:16+00:00
Repo: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset
Dataset Card - POC Triage Medical
Fiche unifiee: inventaire des sources, strategie de selection, schema, gouvernance.
1) Description
Dataset bilingue FR/EN pour triage medical initial.
Le pipeline produit deux artefacts principaux:
SFT: paires instruction/reponse pour le fine-tuning supervise.
DPO: paires… See the full description on the dataset page: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset.SimScholar-SFT
S3 SFT Trajectories
Complete ReAct trajectories for scientific-literature search.
Code ·
S3 collection ·
Source corpus
This dataset contains 14,633 single- and two-hop tool-use trajectories. In each
trajectory, a policy searches and reads a fixed scientific corpus through nine
tools, then submits an answer with a correctness label. The messages column
uses OpenAI tool-calling chat format.
At a glance
Question type
Rows
Correct
Incorrect
Single-hop… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-SFT.cuda-triton-gpu-kernels-2026
⚡ Complete 2026 CUDA & OpenAI Triton High-Performance GPU Kernel Engineering SFT/DPO Suite
The definitive, production-grade synthetic alignment dataset engineered for training and fine-tuning open-weights Large Language Models (Qwen 2.5 Coder, DeepSeek-Coder, Llama 3.1) on ultra-high-throughput GPU kernel programming: NVIDIA Hopper H100 / Blackwell B200 TMA async transfers, OpenAI Triton 3.1+ FlashAttention-3, 32-bank conflict elimination, and low-bit FP8 / INT4 GEMM… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cuda-triton-gpu-kernels-2026.email-triage-v1
Email Triage v1
This dataset contains 1,740 unique email-triage examples produced through Tuned Tensor labeling and hardening workflows. It is the public v1 dataset, de-duplicated from a 2,038-row weighted fine-tuning dataset; 298 intentional weighting duplicates were removed for easier reuse.
The task is operational inbox triage, not security-risk classification. Each row asks a model to label one email-like message and return strict JSON with triage, priority, should_process… See the full description on the dataset page: https://huggingface.co/datasets/tunedtensor/email-triage-v1.tripmatch-ai-daily-plan-alternatives
TripMatch AI — Rich Daily Plan Alternatives
This public academic dataset is the professor-assigned upgrade to TripMatch AI.
Its main deliverable is a substantially richer alternative daily plan generated
with Gemma 3 or Qwen 3 on a GPU.
The original plan is retained only as a side-by-side reference and for the later
LLM comparison task. It is not the new generated target.
Quality contract
Every alternative must:
preserve the destination, exact duration… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-daily-plan-alternatives.tripalchemy-experiences
🧪 TripAlchemy — Synthetic Travel Experiences
10,396 rich, vibe-scored travel experiences across 30 cities — generated by a
pre-trained Hugging Face model and served through a live recommender app.
🚀 Live demo: huggingface.co/spaces/almador2002/tripalchemy
✨ What makes it special
Every experience is scored 0–1 across all six categories at once — 🍽️ culinary, 🏛️ historical,
🛍️ shopping, 🌲 nature, 🌃 nightlife, 🎨 art & culture. That multi-label… See the full description on the dataset page: https://huggingface.co/datasets/almador2002/tripalchemy-experiences.Code_Opt_Triton
Overview
This dataset, TEEN-D/Code_Opt_Triton, is an extended version of the publicly available GPUMODE/Inductor_Created_Data_Permissive dataset. It contains pairs of original (PyTorch or Triton) programs and their equivalent Triton code (generated by torch inductor), intended for training models in PyTorch-to-Triton code translation and optimization.
The primary modification in this extended version is that each optimized Triton code snippet is paired with both its original source… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/Code_Opt_Triton.Vietnamese_literature_VuTrongPhung
Vu Trong Phung Literature Chunks
This dataset consists of Vietnamese literary texts written by author Vũ Trọng Phụng, one of the most influential figures of 20th-century Vietnamese literature.
The dataset includes both short stories and novels, and has been split into smaller chunks based on the number of tokens.
Chunking Strategy
We used the tokenizer from vinai/PhoGPT-4B to split the original text into chunks of less than 512 tokens.
Each row in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/trieunh/Vietnamese_literature_VuTrongPhung.qwen38-27b-triton-ascend-rl-trajectories
Qwen3.8-27B Triton-Ascend RL Trajectories
This dataset contains 1,000 multi-turn trajectories for Triton-Ascend kernel generation. Every included trajectory passed compilation and correctness validation on one official npu-kernelbench workload. Qwen3.8-27B generated an initial solution and received evaluator feedback for up to five calls.
Dataset Viewer subsets
trajectories (default): one row per sample with only messages. The initial system and task user… See the full description on the dataset page: https://huggingface.co/datasets/Yukki1011/qwen38-27b-triton-ascend-rl-trajectories.TrialAgentBench
TrialAgentBench
TrialAgentBench is a synthetic clinical-trial agent benchmark for evaluating
whether agents can analyse participant-level trial evidence, choose defensible
statistical estimands and methods, quantify uncertainty, and make
high-consequence drug-development decisions. The release contains two canonical
sub-benchmarks:
TrialEvalBench: patient-level clinical-trial analysis tasks across
design, assumption, and context tiers.
TrialDevBench: sequential asset-development… See the full description on the dataset page: https://huggingface.co/datasets/TrialAgentBench/TrialAgentBench.TrioBench
TrioBench
TrioBench evaluates LLMs as hybrid query planners across three database engines — SQLite (structured facts + aggregation), Milvus (semantic text/image retrieval), and Neo4j (graph constraints + multi-hop reasoning) — on the Yelp Open Dataset.
Given a natural-language question, a planner must orchestrate the retrieval trio and produce two artifacts: (1) an executable multi-step JSON plan, and (2) a fully executable end-to-end Python program. 341 questions were sent to 5… See the full description on the dataset page: https://huggingface.co/datasets/iwei0/TrioBench.kernelbook-triton-reasoning-traces
KernelBench Triton Reasoning Traces
Reasoning traces generated by the gpt-oss-120b model for converting PyTorch modules to Triton GPU kernels.
Dataset Description
This dataset contains 170 reasoning traces around 85% of them are correct where a PyTorch module was successfully converted to a Triton kernel. Each sample includes the original PyTorch code, the model's reasoning process, and the resulting Triton kernel code along with correctness and performance benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/ppbhatt500/kernelbook-triton-reasoning-traces.pot-o-quantum-88
pot-o-quantum-88
888,888 synthetic PoT-O challenges for tensor + MML path optimization in a Planck–tensor–realm lineage used to probe REALMS / realms-devkit themes: Planck-scale framing (Part I), tensor-network / entropy toy scales (Part IV §3–§4), and coarse-resolution “realm” tensor shapes (including 8×8 and 88×8-style blocks).
Program scale (8.88M lineage): each row carries nominal_lineage_M: 8.88 (documentation of the Quantum-88 / 8.88M Tribewarez program scale). This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Tribewarez/pot-o-quantum-88.autonomous-gpu-kernel-triton-cuda-suite-2026
⚡ Autonomous GPU Kernel, Triton & CUDA Architecture Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Frontier Coding Models (Qwen 3.8, DeepSeek-V3, Llama 3.3)
⚡ Overview & Industry Problem
Modern deep learning accelerators, custom ASICs, and high-performance computing clusters demand specialized, autonomous GPU kernel infrastructure: OpenAI Triton fused kernels, FlashAttention-3 forward/backward online softmax… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-gpu-kernel-triton-cuda-suite-2026.debate-multi-trial-thinking-test
Debate Multi-Trial GRPO Test Data (with Thinking Frameworks)
TEST DATASET - Single debate for review before scaling.
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation,
with integrated thinking framework injection.
What's New: Thinking Frameworks
Each prompt includes structured thinking instructions (mnemonics) that guide the model's reasoning:
Call Type
Mnemonic
Purpose
TACTIC_SELECT
JAM
Judge-Attack-Momentum Analysis… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-multi-trial-thinking-test.ashaar-with-descriptions-baseform-final-trimmed
Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens)
This dataset is derived from Shaer-AI/ashaar-v1-base-form-with-descriptions and prepared for supervised fine-tuning (SFT) with the locked prompt from prompt_testing.ipynb.
Locked Prompt
SYSTEM_PROMPT
أنت شاعر عربي تكتب الشعر العمودي الكلاسيكي.
التزم بالبحر المحدد في كل شطر، واستلهم من الموضوع دون نقله حرفياً.
أخرج الأبيات فقط دون مقدمة أو تعليق.
USER_TEMPLATE
البحر الأساسي:… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed.ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer
Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر)
This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20 and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset.
The only additional change is the removal of rows where:
base_meter == "الوافر"
form == "مجزوء"
This removes poems labeled مجزوء الوافر from the published phase-1 subset.
Why… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer.psy-q-graph-369666
psy-q-graph-369666
369,666 synthetic abstract pathway-graph records in PoT-O-style challenge / optimal_path form. Part of the 369.666.444 (Psy-Q-Finder 369M) program. Graphs use fictional node types (meta, route, guard, probe) and weighted edges — not real molecules, CAS IDs, or laboratory procedures.
Paired base model: Tribewarez/psy-q-finder-369M.
Record schema
Field
Meaning
challenge
Single-line graph spec: nodes, edges, start, goal, forbidden edge… See the full description on the dataset page: https://huggingface.co/datasets/Tribewarez/psy-q-graph-369666.tropt-jailbreak-enhancebench-triggers
TROPT — Jailbreak EnhanceBench Triggers (Exp2: enhancement benchmark)
The companion to
tropt-optbench-triggers,
and its mirror image.
sweeps
holds fixed
tropt-optbench-triggers (Exp1)
the optimizer (15 of them)
the recipe: PrefillCE, plain suffix
this dataset (Exp2)
the jailbreak enhancement
the optimizer: always MAC
So Exp1 asks "which search algorithm finds the best trigger?" and Exp2 asks
"given a fixed search algorithm, which jailbreak tricks actually… See the full description on the dataset page: https://huggingface.co/datasets/MatanBT/tropt-jailbreak-enhancebench-triggers.psy-q-scene-369666
psy-q-scene-369666
369,666 rows of fully synthetic short prose in a Goa / psychedelic-scene-adjacent register: imaginary flyers, DJ blurbs, travelogue scraps, and neutral public-service tone. Not scraped from forums. Not traditional or Indigenous knowledge. Not depicting real places or ceremonies.
Paired base model: Tribewarez/psy-q-finder-369M.
Record schema
Field
Meaning
text
1–3 short paragraphs (fits ~965-token windows when tokenized loosely)
register… See the full description on the dataset page: https://huggingface.co/datasets/Tribewarez/psy-q-scene-369666.debate-multi-trial-grpo
Debate Multi-Trial GRPO Training Data
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation.
Dataset Structure
Each row represents one pipeline call with 4 response variants:
RESPONSE_1_* through RESPONSE_4_*: Different generations at varying temperatures
*_SCORE: Quality score (0.0-1.0) from Haiku evaluator
chosen_index: Index of highest-scoring response
rejected_index: Index of lowest-scoring response
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-multi-trial-grpo.debate-multi-trial-thinking-v3-test
Debate Multi-Trial GRPO Test Data v3 (with Research + Thinking)
TEST DATASET - Single debate for review before scaling.
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation,
with thinking framework injection and multi-hop research calls.
What's New in v3
Multi-hop Research Calls: RESEARCH_QUERY, RESEARCH_EVAL, RESEARCH_CLUE, RESEARCH_DECIDE
Thinking Framework Injection: Structured mnemonics injected INTO perspective before each… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-multi-trial-thinking-v3-test.ipda_grpo_multi_trial_thinking_tactics
Debate Multi-Trial GRPO Test Data (with Thinking Frameworks)
TEST DATASET - Single debate for review before scaling.
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation,
with integrated thinking framework injection.
What's New: Thinking Frameworks
Each prompt includes structured thinking instructions (mnemonics) that guide the model's reasoning:
Call Type
Mnemonic
Purpose
TACTIC_SELECT
JAM
Judge-Attack-Momentum Analysis… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda_grpo_multi_trial_thinking_tactics.ashaar-with-descriptions-baseform-final-trimmed-maxlen20
Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts)
This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed and prepared as a phase-1 GRPO subset by applying a maximum poem length filter of 20 complete bayts while keeping the same columns, prompt format, and general dataset structure as the source dataset.
Locked Prompt
SYSTEM_PROMPT
أنت شاعر عربي تكتب الشعر العمودي الكلاسيكي.
التزم بالبحر المحدد في كل… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20.synthetic-pot-o-challanges-22-22k
synthetic-pot-o-challanges-22-22k
Large synthetic PoT-O (Proof of Tensor Optimizations) challenge set for tensor shape / dtype specs and MML (Minimum Message Length) path optimization. 22,222 examples aligned with generation signature 22.2222 (param_signature field) and MML targets sampled around 0.222222.
Format (JSONL)
Same schema as classic PoT-O challenge JSONL, plus optional lineage field:
challenge — Classic challenge string: tensor:shape=[M… See the full description on the dataset page: https://huggingface.co/datasets/Tribewarez/synthetic-pot-o-challanges-22-22k.tropt-optbench-triggers
TROPT — OptBench Triggers (Exp1: optimizer benchmark)
A store of optimized adversarial trigger suffixes produced by the
TROPT white-box optimizers, together with
the transfer evaluation of each trigger across a held-out set of harmful
instructions.
⚠️ Intended use — defensive security research only. These are adversarial
artifacts for evaluating and hardening LLM robustness (red-teaming,
jailbreak-robustness benchmarking). The harmful instructions come from the
public ClearHarm… See the full description on the dataset page: https://huggingface.co/datasets/MatanBT/tropt-optbench-triggers.
