datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MeetingBank-transcriptThis dataset consists of transcripts from the MeetingBank dataset.
Overview
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for… See the full description on the dataset page: https://huggingface.co/datasets/lytang/MeetingBank-transcript.transcript_isoform_expression_prediction
Multi-modal transcript isoform expression dataset
We curated the human transcript isoform expression dataset from the GTEx portal following the preprocessing pipeline in Garau-Luis et al. (2024). We downloaded the RNA-seq Transcript TPMs file from the bulk tissue expression in GTEx Analysis V8. The table contains transcript expression collected from 30 non-diseased tissues in nearly 1000 human individuals. We averaged the transcript expression measurements across individuals to… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/transcript_isoform_expression_prediction.whisper_transcriptions_greedyllm-bargaining-transcripts
LLM Bargaining Transcripts
240 complete two-agent bargaining games between large language models, played
under an alternating-offers protocol with private valuations, discounting, and
cheap talk. Every game records both agents' true valuations, their
private reasoning, what they claimed about their own position, and what
they actually did.
The dataset is designed to make misrepresentation measurable. Because the true
valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.One-Piece-Transcripts-with-Character-Names-382-777
One Piece Transcripts Dataset (Episodes 382–777)
This dataset contains all dialogue lines from One Piece episodes 382 to 777. The data is stored in a CSV file with the following columns:
episode – episode number
start – start timestamp of the line
end – end timestamp of the line
character – speaking character
text – dialogue text
In addition, the dataset includes the original .sub subtitle files in the folder named "One Piece 382-777". These files were created by the… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/One-Piece-Transcripts-with-Character-Names-382-777.comedy-transcripts
Dataset Summary
This is a dataset of stand up comedy transcripts. It was scraped from
https://scrapsfromtheloft.com/stand-up-comedy-scripts/ and all terms of use
apply. The transcripts are offered to the public as a contribution to education
and scholarship, and for the private, non-profit use of the academic community.
earnings_call_transcript_litetaiwan-legislator-transcript
Taiwan Legislator Transcript
台灣立委公報逐字稿
ERR-transcription-to-subtitlesThis dataset is created by Ilja Samoilov. In dataset is tv show subtitles from ERR and transcriptions of those shows created with TalTech ASR.
from datasets import load_dataset, load_metric
dataset = load_dataset('csv', data_files={'train': "train.tsv", \
"validation":"val.tsv", \
"test": "test.tsv"}, delimiter='\t')
Rick_and_Morty_Transcript
Context
I got inspiration for this dataset from the Rick&Morty Scripts by Andrada Olteanu but felt like dataset was a little small and outdated
This dataset includes almost all the episodes till Season 5. More data will be updated
Content
Rick and Morty Transcripts:
index: index of the row
episode no: the episode where the conversation comes from
speaker: the character's name
dialogue: the dialogue of the character
Acknowledgements
Thanks to the transcripts… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/Rick_and_Morty_Transcript.house-md-transcriptswhisper_transcriptions_greedy_timestampedmedical-transcription-instruct
About
This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field
Dataset Summary
Source: Original medical transcriptions with added instruction-output pairs
Size: 38,924 instruction-output pairs
Format: CSV file
Domain: Medical / Healthcare
Language: English
Last Updated: 08-20-2024
Dataset Structure
Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.Medical_Transcription_LLMinstagram-transcript-tools
Instagram Transcript Tools — Feature Audit Dataset
A hand-collected comparison of 10 tools that turn Instagram video (Reels, Stories, IGTV, Live replay) into text.
Every row was produced by opening the vendor's own page and reading what it actually states. Where a page does not state a value, the cell says not stated — nothing in this dataset is inferred, estimated, or copied from a third-party review.
Why this exists
"Instagram transcript" is a tool-intent query… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/instagram-transcript-tools.Medical_Transcriptiontaiwan-legislator-transcript
Taiwan Legislator Transcript
Overview
The transcripts of speech record happened in various kinds of meetings at Taiwan Legislator Yuan.
The original of transcripts are compiled and published on gazettes from Taiwan Legislator Yuan.
For each segment of transcript, there are corresponding video clip on Legislative Yuan IVOD system.
IVOD stands for Internet Video on Demand system.
For more detail on data origin please look at:
Legislative Yuan Meetings and Gazettes… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/taiwan-legislator-transcript.French-Medical-Transcription-Benchmark
🩺 French Medical Transcription Evaluation Dataset
Ce dataset a été créé et ouvert à la communauté dans le cadre du développement R&D de LucioleScribe, la plateforme souveraine de transcription IA 100% locale, spécifiquement conçue pour les milieux médicaux et juridiques (compatibilité RGPD, HDS, et architectures Air-Gapped).
🔗 Découvrir LucioleScribe Édition Santé | ⚙️ Voir le Pipeline Technologique Local
📊 Présentation du Dataset
L'évaluation des modèles de… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/French-Medical-Transcription-Benchmark.TikTok_MostComment_Video_Transcription_Example
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_MostComment_Video_Transcription_Example.whisper_transcriptions_token_idsTikTok_Most_Shared_Video_Transcription_Example
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_Most_Shared_Video_Transcription_Example.50_reverse_engineered_lecture_transcripts_replicate_1TikTok_Hottest_Video_Transcript_Example
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_Hottest_Video_Transcript_Example.HEST_Xenium_virtual_spatial_transcriptomics
HEST Xenium virtual spatial transcriptomics
This repository contains predicted spatial transcriptomics for HEST Xenium H&E
slides produced with DeepSpot-M.
Authors: Kalin Nonchev, Sebastian Dawo, Karina Silina, Viktor Hendrik
Koelzer, and Gunnar Rätsch.
Paper: DeepSpot-M: a multimodal foundation model for transcriptome-wide virtual spatial transcriptomics from histology (medRxiv, 2026; see the citation below).
Code: https://github.com/ratschlab/DeepSpotM.
News… See the full description on the dataset page: https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics.YouTube_Transcript_Sumyt-titles-transcripts-cleanmedical-transcriptionshuberman_transcriptbible-passage-transcriptionMedical_Transcriptions
