CoolFace
Datasetpublic

NicheVault/nichevault-afrikaans-asr-sample

NicheVault Afrikaans ASR — Free Sample Overview NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations. This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-afrikaans-asr-sample.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes56downloads
Dataset Card

NicheVault Afrikaans ASR — Free Sample

![License](https://creativecommons.org/licenses/by/4.0/) ![Task](https://huggingface.co/tasks/automatic-speech-recognition) ![Language-orange)](#) ![Splits](#)

Overview

NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations.

This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated and ungated on HuggingFace as a discovery funnel for the paid full set.

Important — this is prompted read-speech, not conversational Afrikaans. The source corpora (NCHLT and FLEURS) are scripted, phonetically balanced read-speech collections designed for acoustic model training. They do not represent spontaneous or conversational Afrikaans. Buyers building conversational ASR should be aware of this limitation.

Dataset at a Glance

FieldValue
Sample size20 clips
Audio format16kHz mono WAV
Transcript formatmetadata.jsonl (JSON Lines)
LanguageAfrikaans (af)
SplitsDrawn from train/validation/test (speaker-disjoint)
LicenceCC BY 3.0 (NCHLT) + CC BY 4.0 (FLEURS)
TaskAutomatic Speech Recognition (ASR)

Metadata Fields

FieldDescription
file_nameAudio filename (relative to metadata)
textAfrikaans transcript (sentence-level)
duration_secondsClip duration in seconds
sourceNCHLT Afrikaans Speech Corpus or FLEURS af_za
licenseCC BY 3.0 or CC BY 4.0
languageaf
splittrain, validation, or test
sample_rate16000
channels1

Intended Use

  • —Fine-tuning Whisper and other ASR models for Afrikaans
  • —Acoustic model pre-training for under-resourced South African languages
  • —Benchmarking and evaluation of Afrikaans speech recognition systems
  • —Low-resource ASR research

What the Full Set Adds

The paid full set ($99, 61.5 hours, 67,627 clips) adds:

  • —Scale. Over 67,000 clips (this sample is 20) — 66,133 NCHLT + 1,494 FLEURS utterances from 1,704 distinct speakers.
  • —Proper splits. Speaker-disjoint train/val/test (90/5/5). Zero speaker leakage between splits — critical for valid eval. This is not provided in the raw NCHLT release.
  • —Transcript alignment. All 67,627 clips are MD5-matched against the original NCHLT XML transcript files to eliminate empty-transcript gaps present in the raw corpus.
  • —Whisper-ready. All audio converted to 16kHz mono WAV. No downsampling, no XML parsing, no format conversion required. Drop into a HuggingFace Dataset and start training.
  • —Single-zip delivery. 5 GB archive with all clips, metadata, provenance, and attribution in one download.
  • —Pure CC BY — no licensing friction. The full set excludes all CC BY-SA data (OpenSLR SLR32 2,360 clips removed). Buyers get clean commercial licensing with no ShareAlike obligations.

Why This Exists

Free Afrikaans ASR data is scarce:

  • —NCHLT Afrikaans (raw): ~56h, CC BY 3.0, but requires XML parsing, MD5 transcript matching, sample-rate conversion, and manual split creation. Free download from repo.sadilar.org.
  • —Common Voice Afrikaans: Negligible. CV v26 ~34 MB — effectively useless for ASR training.
  • —Swivuriso / Africa Next Voices: 500h per language across 7 Bantu languages (250h for two) on HuggingFace, but Afrikaans is not included (Bantu languages only).
  • —Way With Words: 50h commercial, conversational — but undisclosed pricing ("contact us"), and only 63 speakers vs our 1,704.

To our knowledge, no other free, processed, Whisper-ready Afrikaans ASR dataset exists.

Attribution

NCHLT Afrikaans Speech Corpus:

  • —CC BY 3.0 Unported
  • —Department of Arts and Culture (DAC), Council for Scientific and Industrial Research (CSIR), and North-West University (NWU), South Africa
  • —Barnard et al., "The NCHLT Speech Corpus of the South African languages," SLTU 2014
  • —Download: https://repo.sadilar.org/

FLEURS af_za:

Citation

If you use this dataset, please cite the original NCHLT and FLEURS sources:

bibtex
@inproceedings{barnard2014nchlt,
  title     = {The {NCHLT} Speech Corpus of the South {A}frican languages},
  author    = {Barnard, Etienne and Davel, Marelie H and van Heerden, Charl and de Wet, Febe and Badenhorst, Jaco},
  booktitle = {Workshop on Spoken Language Technologies for Under-resourced Languages (SLTU)},
  year      = {2014}
}

@inproceedings{conneau2023fleurs,
  title  = {{FLEURS}: Few-shot Learning Evaluation of Universal Representations of Speech},
  author = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and Dalmia, Siddharth and Riesa, Jason and Rivera, Clara and Bapna, Ankur},
  booktitle = {IEEE Spoken Language Technology Workshop (SLT)},
  year   = {2023}
}

Get the Full Dataset

The complete 61.5-hour dataset is available for purchase at $99.

➡️ [Buy the full dataset on Gumroad](https://flevin4.gumroad.com/l/lvjehy)

Includes: 67,627 WAV clips, metadata.jsonl, provenance.txt, README — everything Whisper-ready in a single 5 GB zip. Pure CC BY, zero ShareAlike.