datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-agent-coordination-transcripts
Multi Agent Coordination Transcripts
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.call-transcripts-training-dataaltayduel-transcripts
🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2)
Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir.
📌 TL;DR
2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu.
439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.rlvr-reward-hacking-transcripts
RLVR reward-hacking full trajectories
This release contains 900 full held-out trajectories from three policies trained with
reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable
CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B
checkpoints. Each row preserves the task, tests, complete prompts, native
reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.shell-safety-transcriptsConverted from tomngdev/shell-safety into conversations transcripts.
Structure is for my own training with static system prompt and changing <SessionContext> block
shell-safety-v2-transcriptsConverted from tomngdev/shell-safety-v2 into transcripts.
Structure is for my own training with static system prompt and changing <SessionContext></SessionContext> block
lex-transcriptsrlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.huberman-lab-transcripts
Huberman Lab Transcript Dataset
Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel.
Dataset
438 videos
9,833 transcript chunks
~114 million characters
JSONL format
Each record contains:
ext
ideo_id
itle
Processing
The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/huberman-lab-transcripts.MixtureVitae-finevideo-transcripts-onlyThis is the whisper transcription from the Finvideo dataset. There are artifacts esp when music or sound is mistakenly transcribed as words. TBD: cleanup these repetitious words.
ibl-khanacademy-transcripts
ibl-khanacademy-transcripts
This dataset houses the transcripts of openly available videos from Khan Academy.
The transcripts were scrapped from Khan Academy's youtube channel
hmi-transcripts
MetaFLOS HMI Transcripts (Manufacturing Human-Machine Interaction Dialogue Dataset)
Multi-turn dialogue dataset between manufacturing floor operators and machine intelligent assistants, generated in Vicuna style (seed scenario → LLM produces transcript).
Content
30 scenarios / 325 dialogue turns (6–12 turns per scenario)
Covers 16 domains: textile warping/weaving, LCD panels, semiconductor processes, fans/blowers, water chillers, general equipment reliability… See the full description on the dataset page: https://huggingface.co/datasets/tjw/hmi-transcripts.whisper-transcripts-the-vergeannotations_creators:
machine-generated
language:
en
language_creators:
crowdsourced
license: []
multilinguality:
monolingual
paperswithcode_id: wikitext-2
pretty_name: Whisper-Transcripts
size_categories:
1M<n<10M
source_datasets:
original
tags: []
task_categories:
text-generation
fill-mask
task_ids:
language-modeling
masked-language-modeling
Farsight-SRV-Transcripts
The Farsight Institute: Scientific Remote Viewing (SRV) Transcripts
Dataset Summary
This dataset contains the complete, unabridged archive of Scientific Remote Viewing (SRV) session transcripts and project summaries produced by The Farsight Institute, directed by Dr. Courtney Brown.
The data consists of hundreds of highly detailed, text-rich transcripts describing historical events, planetary mysteries, and extraterrestrial dynamics. All remote viewing sessions… See the full description on the dataset page: https://huggingface.co/datasets/courtnoski/Farsight-SRV-Transcripts.911-call-transcriptsCC-BY-STEMM-Podcast-Transcriptsadaption-jawi-htr-transcripts
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-jawi_htr_transcripts
This dataset contains prompt and completion samples featuring scanned handwritten pages in the Jawi script alongside verified gold-level text transcriptions. It provides structured examples for evaluating and fine-tuning handwritten text recognition models in artificial intelligence and computational linguistics. Each entry pairs page-level handwritten input… See the full description on the dataset page: https://huggingface.co/datasets/ViratChauhan/adaption-jawi-htr-transcripts.council-transcripts
Council Multi-Agent Deliberation Transcripts
Real deliberation transcripts from Council, a multi-agent orchestration skill for OpenClaw.
What This Is
Council routes tasks to specialized Grok persona agents — Workhorse (deep technical reasoning), Creative (novel ideas, chaos energy), and Speed (fast iteration) — then synthesizes their outputs into a final verdict via a Conductor. These transcripts capture the full deliberation process: prompts, per-persona responses, and… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/council-transcripts.transcripts-mergedtranscripts-separatepufi-duf-transcriptslogical-transcripts
logical-transcripts
Golden paired dataset for training models to transliterate Arabic Latin text into
scholarly diacritized form — built from a single recorded Islamic lecture
(Chapter 24, Lecture 16) with a raw ASR transcript and a human-polished scholarly
transcript.
Two artifacts are stored separately for provenance and review:
File
Rows
Purpose
train.jsonl
203
Golden — quality-filtered pairs for training
bronze.jsonl
773
Bronze — every aligned sentence pair… See the full description on the dataset page: https://huggingface.co/datasets/olanigan/logical-transcripts.cleaned-asr-transcriptssmall_transcripts_alpacaCC-BY-STEMM-Podcast-Transcripts-2048Tony-Chase-Transcripts
Tony Chase Transcripts
Around 3500 transcripts of videos from Tony Chase captioned with GPT-3.5-Turbo.
cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.youtube-transcripts-metadatacleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.youtube-transcripts-05-16-24
