datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transcript_isoform_expression_prediction
Multi-modal transcript isoform expression dataset
We curated the human transcript isoform expression dataset from the GTEx portal following the preprocessing pipeline in Garau-Luis et al. (2024). We downloaded the RNA-seq Transcript TPMs file from the bulk tissue expression in GTEx Analysis V8. The table contains transcript expression collected from 30 non-diseased tissues in nearly 1000 human individuals. We averaged the transcript expression measurements across individuals to… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/transcript_isoform_expression_prediction.llm-bargaining-transcripts
LLM Bargaining Transcripts
240 complete two-agent bargaining games between large language models, played
under an alternating-offers protocol with private valuations, discounting, and
cheap talk. Every game records both agents' true valuations, their
private reasoning, what they claimed about their own position, and what
they actually did.
The dataset is designed to make misrepresentation measurable. Because the true
valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.ERR-transcription-to-subtitlesThis dataset is created by Ilja Samoilov. In dataset is tv show subtitles from ERR and transcriptions of those shows created with TalTech ASR.
from datasets import load_dataset, load_metric
dataset = load_dataset('csv', data_files={'train': "train.tsv", \
"validation":"val.tsv", \
"test": "test.tsv"}, delimiter='\t')
Rick_and_Morty_Transcript
Context
I got inspiration for this dataset from the Rick&Morty Scripts by Andrada Olteanu but felt like dataset was a little small and outdated
This dataset includes almost all the episodes till Season 5. More data will be updated
Content
Rick and Morty Transcripts:
index: index of the row
episode no: the episode where the conversation comes from
speaker: the character's name
dialogue: the dialogue of the character
Acknowledgements
Thanks to the transcripts… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/Rick_and_Morty_Transcript.medical-transcription-instruct
About
This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field
Dataset Summary
Source: Original medical transcriptions with added instruction-output pairs
Size: 38,924 instruction-output pairs
Format: CSV file
Domain: Medical / Healthcare
Language: English
Last Updated: 08-20-2024
Dataset Structure
Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.Medical_TranscriptionHEST_Xenium_virtual_spatial_transcriptomics
HEST Xenium virtual spatial transcriptomics
This repository contains predicted spatial transcriptomics for HEST Xenium H&E
slides produced with DeepSpot-M.
Authors: Kalin Nonchev, Sebastian Dawo, Karina Silina, Viktor Hendrik
Koelzer, and Gunnar Rätsch.
Paper: DeepSpot-M: a multimodal foundation model for transcriptome-wide virtual spatial transcriptomics from histology (medRxiv, 2026; see the citation below).
Code: https://github.com/ratschlab/DeepSpotM.
News… See the full description on the dataset page: https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics.bible-passage-transcriptionnpi_all_transcriptsMedical_Transcriptions_upsampledearning_transcripts_chunksari-matti-podcast-transcriptions
