martinturuta/safi-diction-sample
Safi Diction Sample This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents. The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours. This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-diction-sample.
Safi Diction Sample
This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents.
The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours.
This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.
Total audio clips: 472
Total unique speakers: 71
Dataset Structure
The repository contains:
metadata.csv
audio/
sample_0001_q1_1.ogg
sample_0002_q1_2.ogg
...
Metadata
The dataset metadata is stored in metadata.csv. Each row corresponds to one audio recording.
Format
All audio files are presented as raw, .ogg files. No transformations or processing has been applied to them.
Cleaning
Audio files that we're missing or lacked any speech were removed from the dataset based on a data density check.
Splits
This sample dataset does not include train/test/validation splits. Hugging Face may display the data under a default train split, but that should be understood as the single available dataset split.
License and Usage
Usage is restricted to the terms agreed upon with the dataset owner. Do not redistribute without permission.
Commissioners and Authors
This dataset was commissioned by Safi.
Dataset preparation, cleaning, and packaging were completed by the Safi team.
Audio recordings were contributed by survey respondents who participated in the data collection process.
Reach out to martin@safi.world for question/concerns.
