NicheVault/nichevault-afrikaans-asr-sample
NicheVault Afrikaans ASR — Free Sample Overview NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations. This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-afrikaans-asr-sample.
NicheVault Afrikaans ASR — Free Sample
   
Overview
NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations.
This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated and ungated on HuggingFace as a discovery funnel for the paid full set.
Important — this is prompted read-speech, not conversational Afrikaans. The source corpora (NCHLT and FLEURS) are scripted, phonetically balanced read-speech collections designed for acoustic model training. They do not represent spontaneous or conversational Afrikaans. Buyers building conversational ASR should be aware of this limitation.
Dataset at a Glance
Metadata Fields
Intended Use
- Fine-tuning Whisper and other ASR models for Afrikaans
- Acoustic model pre-training for under-resourced South African languages
- Benchmarking and evaluation of Afrikaans speech recognition systems
- Low-resource ASR research
What the Full Set Adds
The paid full set ($99, 61.5 hours, 67,627 clips) adds:
- Scale. Over 67,000 clips (this sample is 20) — 66,133 NCHLT + 1,494 FLEURS utterances from 1,704 distinct speakers.
- Proper splits. Speaker-disjoint train/val/test (90/5/5). Zero speaker leakage between splits — critical for valid eval. This is not provided in the raw NCHLT release.
- Transcript alignment. All 67,627 clips are MD5-matched against the original NCHLT XML transcript files to eliminate empty-transcript gaps present in the raw corpus.
- Whisper-ready. All audio converted to 16kHz mono WAV. No downsampling, no XML parsing, no format conversion required. Drop into a HuggingFace Dataset and start training.
- Single-zip delivery. 5 GB archive with all clips, metadata, provenance, and attribution in one download.
- Pure CC BY — no licensing friction. The full set excludes all CC BY-SA data (OpenSLR SLR32 2,360 clips removed). Buyers get clean commercial licensing with no ShareAlike obligations.
Why This Exists
Free Afrikaans ASR data is scarce:
- NCHLT Afrikaans (raw): ~56h, CC BY 3.0, but requires XML parsing, MD5 transcript matching, sample-rate conversion, and manual split creation. Free download from repo.sadilar.org.
- Common Voice Afrikaans: Negligible. CV v26 ~34 MB — effectively useless for ASR training.
- Swivuriso / Africa Next Voices: 500h per language across 7 Bantu languages (250h for two) on HuggingFace, but Afrikaans is not included (Bantu languages only).
- Way With Words: 50h commercial, conversational — but undisclosed pricing ("contact us"), and only 63 speakers vs our 1,704.
To our knowledge, no other free, processed, Whisper-ready Afrikaans ASR dataset exists.
Attribution
NCHLT Afrikaans Speech Corpus:
- CC BY 3.0 Unported
- Department of Arts and Culture (DAC), Council for Scientific and Industrial Research (CSIR), and North-West University (NWU), South Africa
- Barnard et al., "The NCHLT Speech Corpus of the South African languages," SLTU 2014
- Download: https://repo.sadilar.org/
FLEURS af_za:
- CC BY 4.0
- Google LLC
- https://huggingface.co/datasets/google/fleurs
Citation
If you use this dataset, please cite the original NCHLT and FLEURS sources:
@inproceedings{barnard2014nchlt,
title = {The {NCHLT} Speech Corpus of the South {A}frican languages},
author = {Barnard, Etienne and Davel, Marelie H and van Heerden, Charl and de Wet, Febe and Badenhorst, Jaco},
booktitle = {Workshop on Spoken Language Technologies for Under-resourced Languages (SLTU)},
year = {2014}
}
@inproceedings{conneau2023fleurs,
title = {{FLEURS}: Few-shot Learning Evaluation of Universal Representations of Speech},
author = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and Dalmia, Siddharth and Riesa, Jason and Rivera, Clara and Bapna, Ankur},
booktitle = {IEEE Spoken Language Technology Workshop (SLT)},
year = {2023}
}Get the Full Dataset
The complete 61.5-hour dataset is available for purchase at $99.
➡️ [Buy the full dataset on Gumroad](https://flevin4.gumroad.com/l/lvjehy)
Includes: 67,627 WAV clips, metadata.jsonl, provenance.txt, README — everything Whisper-ready in a single 5 GB zip. Pure CC BY, zero ShareAlike.
