datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Punjabi-Studio-Voice-Corpus
🎙️ Punjabi (Gurmukhi) Synthetic Multi-Generation Voice Corpus
☬ ਸ਼ੁੱਧ ਗੁਰਮੁਖੀ ਮਹਾਨ ਕੋਸ਼ ਸਿੰਥੈਟਿਕ ਵੌਇਸ ਡਾਟਾਸੈੱਟ
⚠️ Corrected 2026-09-02. The original README described this as
"5 distinct acoustic age & gender profiles" of "Studio Master" recordings.
That was false. This is 100% synthetic, machine-generated speech —
not a multi-speaker human recording. See "How the audio was made" below.
A Punjabi (Gurmukhi) synthetic speech corpus built from two underlying… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Studio-Voice-Corpus.Gurbani-MahanKosh-Frontier-Corpus
ੴ Gurbani & Bhai Kahn Singh Nabha Mahan Kosh Frontier Corpus
☬ ਗੁਰਬਾਣੀ ਅਤੇ ਭਾਈ ਕਾਹਨ ਸਿੰਘ ਨਾਭਾ 'ਮਹਾਨ ਕੋਸ਼' ਪ੍ਰਮਾਣਿਕ ਡਾਟਾਸੈੱਟ
👨💻 Project Lead & Architecture
Curator & Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Project: AMRIT Research OS (Autonomous Medical AI)
📖 Dataset Overview
An authoritative lexical dataset compiling authentic definitions… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Gurbani-MahanKosh-Frontier-Corpus.
