CoolFace
Datasetpublic

manassehzw/shona-bible-bdsc-aligned

Shona Bible Speech Alignment Dataset Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC source audio made available by Biblica, Inc. through Open.Bible. This release contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech segments covering approximately 75.55 hours. Dataset summary Language: Shona (sna) Speaker: narrator 1 Speaker sex: male Books: 66 Clips: 31,284 Audio: approximately 75.55 hours Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.

sourceHugging Facecc-by-sa-4.0updated 19d agoView on Hugging Face
3likes281downloads
Dataset Card

Shona Bible Speech Alignment Dataset

Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC source audio made available by Biblica, Inc. through Open.Bible. This release contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech segments covering approximately 75.55 hours.

Dataset summary

  • —Language: Shona (sna)
  • —Speaker: narrator 1
  • —Speaker sex: male
  • —Books: 66
  • —Clips: 31,284
  • —Audio: approximately 75.55 hours
  • —Audio format: mono 16 kHz FLAC
  • —Split: train
  • —Parquet shards: 16

Segmentation and alignment

Chapter introductions and trailing material were removed using the first and last timestamped canonical verses. Internal segmentation is verse-first. Unusually long verses are split at natural punctuation where possible, while retaining verse_part and verse_part_count metadata.

The canonical Bible text is the transcript source of truth. Parakeet word timestamps and fuzzy lexical alignment were used to place canonical words on the audio timeline. The published transcript is canonical text, not the raw ASR surface form.

Manual quality review identified six minor boundary-clipping cases across 31,284 segments (approximately 0.02% of segments and 0.02% of total audio duration). The dataset was retained without correction because further boundary edits could introduce new clipping elsewhere.

Fields

  • —audio: lossless 16 kHz mono FLAC audio.
  • —transcription: canonical Shona verse or verse part.
  • —book, chapter, verse: canonical Bible location.
  • —verse_part, verse_part_count: traceability for long verses.
  • —duration: clip duration in seconds.
  • —speaker_id, speaker_sex, language: speaker and language metadata.

Detailed alignment provenance is kept separately in metadata/segment_map.jsonl and is not part of the public training schema.

Source and attribution

This dataset was derived from the Shona Bible audio and text made available by Biblica, Inc. through Open.Bible.

Original source: https://www.open.bible/bibles/shona-biblica-audio-bible

The original work by Biblica, Inc. is available for free at www.biblica.com and open.bible.

Copyright notices

Biblica® Bhaibheri Dzvene Rakasununguka MuChiShona Chanhasi™, Chikamu chinonzwika nenzeve Kopakodzero yezvinonzwika nenzeve ℗ 2015 ye Biblica, Inc.

Biblica® Open Shona Contemporary Bible™, Audio Edition Audio Copyright ℗ 2015 by Biblica, Inc.

Biblica® Bhaibheri Dzvene Rakasununguka MuChiShona Chanhasi™ Kopakodzero © 2005, 2018 ne Biblica, Inc.

Biblica® Open Shona Contemporary Bible™ Copyright © 2005, 2018 by Biblica, Inc.

The original recordings and text are licensed under the Creative Commons Attribution-ShareAlike 4.0 International License.

Modifications

  • —Audio chapters were divided into smaller speech segments.
  • —Shona biblical text was aligned with the corresponding audio.
  • —Segment timestamps and identifiers were added to the internal map.
  • —Audio was converted to mono 16 kHz FLAC for machine-learning use.
  • —Files were renamed and reorganized for machine-learning use.
  • —Alignment metadata and quality-control information were added.

These modifications were made by manassehzw in 2026.

This dataset is not an official Biblica product and has not been endorsed by Biblica.

License

This derivative dataset is released under the Creative Commons Attribution-ShareAlike 4.0 International License:

https://creativecommons.org/licenses/by-sa/4.0/

Anyone may use, modify, and redistribute the dataset, including commercially, provided that they credit the original source and contributors, indicate changes, distribute adaptations under CC BY-SA 4.0, and do not impose additional legal or technological restrictions. See the source attribution and license URL in this README for the applicable terms.

Book-disjoint configuration

The book_disjoint configuration assigns every complete Bible book to exactly one split. It uses seed 42 and the stable SHA-256 ordering recorded in book_disjoint/split_assignment.json. Use its training split for model fitting, its validation split for checkpoint selection, and its test split only for final evaluation.

SplitRowsHoursBooks
train27,28165.25883160
validation2,4106.3248863
test1,5933.9623313

The original default configuration remains the unsplit source release.