datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper_transcriptions_greedyERR-transcription-to-subtitlesThis dataset is created by Ilja Samoilov. In dataset is tv show subtitles from ERR and transcriptions of those shows created with TalTech ASR.
from datasets import load_dataset, load_metric
dataset = load_dataset('csv', data_files={'train': "train.tsv", \
"validation":"val.tsv", \
"test": "test.tsv"}, delimiter='\t')
medical-transcription-instruct
About
This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field
Dataset Summary
Source: Original medical transcriptions with added instruction-output pairs
Size: 38,924 instruction-output pairs
Format: CSV file
Domain: Medical / Healthcare
Language: English
Last Updated: 08-20-2024
Dataset Structure
Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.whisper_transcriptions_greedy_timestampedTikTok_Most_Shared_Video_Transcription_Example
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_Most_Shared_Video_Transcription_Example.whisper_transcriptions_token_idsMedical_Transcription_LLMTikTok_MostComment_Video_Transcription_Example
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_MostComment_Video_Transcription_Example.Medical_TranscriptionFrench-Medical-Transcription-Benchmark
🩺 French Medical Transcription Evaluation Dataset
Ce dataset a été créé et ouvert à la communauté dans le cadre du développement R&D de LucioleScribe, la plateforme souveraine de transcription IA 100% locale, spécifiquement conçue pour les milieux médicaux et juridiques (compatibilité RGPD, HDS, et architectures Air-Gapped).
🔗 Découvrir LucioleScribe Édition Santé | ⚙️ Voir le Pipeline Technologique Local
📊 Présentation du Dataset
L'évaluation des modèles de… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/French-Medical-Transcription-Benchmark.medical-transcriptionsbible-passage-transcriptionMedical_TranscriptionsEndocrinology_transcription_and_notesTranscriptionClassifierCohort
TranscriptionClassifierCohort
tags: Classification, Medical, Specialty
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The dataset titled 'TranscriptionClassifierCohort' is designed to assist machine learning practitioners in developing a medical transcription classifier that can differentiate between various medical specialties based on transcription data. Each row in the dataset represents a sample transcription excerpt labeled… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/TranscriptionClassifierCohort.ru_transcription_punctuation
About
This is a dataset for training Russian punctuators/capitalizers via NeMo scripts (https://github.com/NVIDIA/NeMo)
A BERT model already fine-tuned on this dataset can be found here: https://huggingface.co/denis-berezutskiy-lad/lad_transcription_bert_ru_punctuator
Scripts for collecting/updating such a dataset, as well as training/using the model are located here: https://github.com/denis-berezutskiy-lad/transcription-bert-ru-punctuator-scripts/tree/main
The idea behind the… See the full description on the dataset page: https://huggingface.co/datasets/denis-berezutskiy-lad/ru_transcription_punctuation.natural_speech_transcriptionsMedical_Transcriptions_upsampledari-matti-podcast-transcriptionsrap-transcription_hindi_transcriptiontranscription_synonyms
