datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sp500-earnings-transcripts
S&P 500 Earnings Call Transcripts
Dataset Description
This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals.
📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025.
Coverage Statistics
Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.measuring_cot_monitorability_transcripts
Measuring Chain-of-Thought Monitorability Transcripts
This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness.
We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.fomc-meeting-transcripts
FOMC Meeting Transcripts (1976–2020)
Full-text transcripts of 373 Federal Open Market Committee (FOMC) meetings, from March 1976 through December 2020, converted from the official PDF transcripts published by the Federal Reserve Board.
The FOMC is the body of the U.S. Federal Reserve System that sets monetary policy (the federal funds rate target, balance-sheet policy, etc.). Verbatim meeting transcripts are released to the public with a roughly five-year lag, which is why… See the full description on the dataset page: https://huggingface.co/datasets/brishen/fomc-meeting-transcripts.freecodecamp-transcripts
Free Code Camp Transcripts
Overview
This dataset contains transcripts of programming tutorials from FreeCodeCamp videos. Each entry includes the video title, YouTube video ID, and the full transcript, making it suitable for training and evaluating NLP and LLM systems focused on developer education.
DataSource
Dataset Structure
Column
Type
Description
title
string
Title of the YouTube video
video_id
string
Unique YouTube video identifier… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/freecodecamp-transcripts.math-eval-transcripts-a
MATH Evaluation Transcripts — Auditing Set A
Full model responses on the MATH held-out test split for two models, to support
behavioural auditing. All transcripts are from a single default condition (a standard
step-by-step solve prompt; no special system prompt or prefix).
This is one of a pair of sets derived from a common transcript pool. Each set contains the
same trusted model and one model under investigation; the sets do not disclose how the
two investigated models relate… See the full description on the dataset page: https://huggingface.co/datasets/darklord1611/math-eval-transcripts-a.math-eval-transcripts-b
MATH Evaluation Transcripts — Auditing Set B
Full model responses on the MATH held-out test split for two models, to support
behavioural auditing. All transcripts are from a single default condition (a standard
step-by-step solve prompt; no special system prompt or prefix).
This is one of a pair of sets derived from a common transcript pool. Each set contains the
same trusted model and one model under investigation; the sets do not disclose how the
two investigated models relate… See the full description on the dataset page: https://huggingface.co/datasets/darklord1611/math-eval-transcripts-b.
