yuriyvnv
Datasets
All datasets matching “yuriyvnv”capes_synthetic_audio_filteredsynthetic_transcript_pt
Portuguese Speech Dataset with Multiple Training Configurations
A comprehensive Portuguese speech dataset offering three distinct training configurations for speech recognition research, each designed for different experimental scenarios and training paradigms.
🎯 Dataset Configurations Overview
This dataset provides three carefully curated subsets to enable comprehensive speech recognition research:
Configuration
Training Data
Validation
Test
Total Samples
Use Case… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_pt.synthetic_transcript_nl
Dutch Synthetic Speech Transcripts
This dataset contains 34,898 synthetic Dutch speech samples generated using GPT-4o-mini for transcript creation and OpenAI's TTS-1 model for speech synthesis. It was designed to augment Automatic Speech Recognition (ASR) training for low-resource scenarios, matching the linguistic distribution of Common Voice 17.0 Dutch.
Dataset Description
Purpose
This dataset addresses the challenge of limited labeled speech data for Dutch… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_nl.synthetic_asr_et_sltriage_transcriptions
Medical Triage Transcriptions Dataset
Credits and Acknowledgments
This dataset is based on the original NLie2/TRIAGE dataset. We thank the original creators for providing the foundational triage classification data that enabled this synthetic transcription generation.
Original Dataset: NLie2/TRIAGELicense: Please refer to the original dataset license
Dataset Description
This dataset contains synthetic medical triage transcriptions generated from the… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/triage_transcriptions.triage_synthetic_classification
