datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DASS2019_NLP
Dataset Card for DASS2019_NLP
This dataset contains audio and transcript content from DASS2019, the manually transcribed version of the Digital Archive of Southern Speech.
It may be suitable for speech-related NLP processing, modelling, and fine-tuning tasks.
Dataset Details
DASS (Kretzschmar et al. 2012) comprises dialectological interviews with 64 informants conducted between 1968 and 1983; it is a subset of the larger Linguistic Atlas of the Gulf States (LAGS… See the full description on the dataset page: https://huggingface.co/datasets/stcoats/DASS2019_NLP.dv-presidential-speechDhivehi Presidential Speech is a Dhivehi speech dataset created from data extracted and
processed by [Sofwath](https://github.com/Sofwath) as part of a collection of Dhivehi
datasets found [here](https://github.com/Sofwath/DhivehiDatasets).
The dataset contains around 2.5 hrs (1 GB) of speech collected from Maldives President's Office
consisting of 7 speeches given by President Yaameen Abdhul Gayyoom.
