CoolFace
Datasetpublic

martinturuta/safi-diction-sample

Safi Diction Sample This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents. The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours. This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-diction-sample.

sourceHugging Faceotherupdated 18d agoView on Hugging Face
0likes219downloads
Dataset Card

Safi Diction Sample

This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents.

The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours.

This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.

Total audio clips: 472

Total unique speakers: 71

Dataset Structure

The repository contains:

text
metadata.csv
audio/
  sample_0001_q1_1.ogg
  sample_0002_q1_2.ogg
  ...

Metadata

The dataset metadata is stored in metadata.csv. Each row corresponds to one audio recording.

ColumnDescription
file_nameRelative path to the audio file.
transcriptionExpected sentence the respondent was asked to dictate.
respondentAnonymized respondent identifier.
genderRespondent gender, when available.
age_rangeRespondent age bucket.
prompt_numberUnique identifier for the dictation prompt.

Format

All audio files are presented as raw, .ogg files. No transformations or processing has been applied to them.

Cleaning

Audio files that we're missing or lacked any speech were removed from the dataset based on a data density check.

Splits

This sample dataset does not include train/test/validation splits. Hugging Face may display the data under a default train split, but that should be understood as the single available dataset split.

License and Usage

Usage is restricted to the terms agreed upon with the dataset owner. Do not redistribute without permission.

Commissioners and Authors

This dataset was commissioned by Safi.

Dataset preparation, cleaning, and packaging were completed by the Safi team.

Audio recordings were contributed by survey respondents who participated in the data collection process.

Reach out to martin@safi.world for question/concerns.