datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
instructpix2pix_toontown_100einstructpix2pix_toontown_50_battlesPunjabi-Studio-Voice-Corpus
🎙️ Punjabi (Gurmukhi) Synthetic Multi-Generation Voice Corpus
☬ ਸ਼ੁੱਧ ਗੁਰਮੁਖੀ ਮਹਾਨ ਕੋਸ਼ ਸਿੰਥੈਟਿਕ ਵੌਇਸ ਡਾਟਾਸੈੱਟ
⚠️ Corrected 2026-09-02. The original README described this as
"5 distinct acoustic age & gender profiles" of "Studio Master" recordings.
That was false. This is 100% synthetic, machine-generated speech —
not a multi-speaker human recording. See "How the audio was made" below.
A Punjabi (Gurmukhi) synthetic speech corpus built from two underlying… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Studio-Voice-Corpus.Gurbani-MahanKosh-Frontier-Corpus
ੴ Gurbani & Bhai Kahn Singh Nabha Mahan Kosh Frontier Corpus
☬ ਗੁਰਬਾਣੀ ਅਤੇ ਭਾਈ ਕਾਹਨ ਸਿੰਘ ਨਾਭਾ 'ਮਹਾਨ ਕੋਸ਼' ਪ੍ਰਮਾਣਿਕ ਡਾਟਾਸੈੱਟ
👨💻 Project Lead & Architecture
Curator & Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Project: AMRIT Research OS (Autonomous Medical AI)
📖 Dataset Overview
An authoritative lexical dataset compiling authentic definitions… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Gurbani-MahanKosh-Frontier-Corpus.Punjabi-STEM-Frontier-CoT-Corpus
ੴ Punjabi STEM & Frontier Chain-of-Thought (CoT) Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਗਿਆਨ ਅਤੇ ਉੱਚ-ਗਣਿਤ ਕਦਮ-ਦਰ-ਕਦਮ ਤਰਕ ਡਾਟਾਸੈੱਟ
👨💻 Research & Engineering Architecture
Architect & Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
Vision: Sovereign Indic Intelligence & Advanced Scientific Reasoning in Gurmukhi.
Flagship Innovation: AMRIT Research OS & Sehaj Sovereign Neural Model
📖 Dataset Overview / ਸੰਖੇਪ
The Punjabi… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-STEM-Frontier-CoT-Corpus.Punjabi-Gurmukhi-Grammar-Correction-Corpus
ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Engineering Lead
Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Innovation: AMRIT Research OS (100% Locally-Run Autonomous Medical AI)
📖 Overview / ਸੰਖੇਪ
The Punjabi (Gurmukhi) Grammatical Error Correction… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.AMRIT-Punjabi-Clinical-Dialogue-Corpus
ੴ AMRIT Punjabi Clinical Dialogue & Medical Diagnosis Corpus
☬ ਅੰਮ੍ਰਿਤ ਪੰਜਾਬੀ ਕਲੀਨਿਕਲ ਸੰਵਾਦ ਅਤੇ ਡਾਕਟਰੀ ਨਿਦਾਨ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Medical AI Architecture
Lead Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
Mission: Free Autonomous AI Doctor for Humanity (ਦੁਨੀਆਂ ਦੇ ਲੋੜਵੰਦ ਲੋਕਾਂ ਲਈ ਮੁਫ਼ਤ AI ਡਾਕਟਰ)
Flagship Platform: AMRIT Research OS (100% Local Medical Intelligence)
📖 Dataset Overview / ਸੰਖੇਪ
The… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/AMRIT-Punjabi-Clinical-Dialogue-Corpus.Sehaj-Gurmukhi-Frontier-Reasoning
🧠 Punjabi (Gurmukhi) Frontier Reasoning & Longevity AI Corpus
Punjabi (Gurmukhi) Frontier Reasoning (ਪੰਜਾਬੀ ਗੁਰਮੁਖੀ ਡਾਟਾਸੈੱਟ) is a high-density, multi-domain Punjabi dataset engineered for training next-generation intelligent Punjabi AI models.
🔍 Search & Discovery Keywords
Language: Punjabi / Gurmukhi (ਪੰਜਾਬੀ / ਗੁਰਮੁਖੀ)
Domains: Quantum Science, Higher Mathematics, Computer Science, Longevity Vitals, and Classical Philosophy in Punjabi.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Sehaj-Gurmukhi-Frontier-Reasoning.amrit-sweets-multimodal
Amrit Multimodal Dataset
Overview
High-quality multimodal dataset for product recommendation and classification tasks.
Total Samples: 10
Categories: 3
Generated: 2026-07-10T11:22:46.406680
Quality: 100% validated, no duplicates, no errors
Features
product_id: Unique product identifier
name: Product name
category: Product category (sweets, savory, dry_fruits, chocolates)
description: Detailed text description (multimodal text data)
price_inr:… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/amrit-sweets-multimodal.amrit-qwythos-corrections
Amrit-Qwythos Corrections Dataset
49 question → correct-answer pairs used to fine-tune Qwythos-9B (a
Qwen3.5-based model) to fix specific, verified hallucinations caught during
real use in the Amrit OS project.
What this fixes
Two real, reproduced failure classes:
Self-identity confusion — at higher sampling temperature, the base
model would sometimes answer "I am Qwen3.5, made by Alibaba Cloud"
instead of its actual fine-tuned identity (Qwythos, by Empero AI).… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/amrit-qwythos-corrections.TommyPunjab-Heritage-Gurmat-Theology-Corpus
ੴ Punjab Heritage, Gurmat Philosophy & Comparative Theology Corpus
☬ ਪੰਜਾਬੀ ਵਿਰਸਾ, ਗੁਰਮਤਿ ਫ਼ਲਸਫ਼ਾ ਅਤੇ ਤੁਲਨਾਤਮਕ ਧਰਮ ਅਧਿਐਨ ਪ੍ਰਮਾਣਿਕ ਡਾਟਾਸੈੱਟ
👨💻 Research & Academic Architecture
Lead Researcher & Curator: Gurpreet Singh Dhillon (Nam-toon Studio)
Mission: Authentic, Uncompromised, Source-Based Gurmat & Indic Historiography for Sovereign Artificial Intelligence.
Flagship Platform: AMRIT Research OS & Sehaj Sovereign AI… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjab-Heritage-Gurmat-Theology-Corpus.Fine-TOONing_DatasetToonpimp_game_styletoonpoet-style-fullres-v2TOON-Unstructured-Structured
TOON-Unstructured-Structured
This dataset is a validated and cleaned version of the originalMasterControlAIML/JSON-Unstructured-Structured.
It has been reformatted using the official Token-Oriented Object Notation (TOON) specification —a compact, token-efficient data serialization format optimized for LLM-ready structured data.All records have been verified for JSON integrity and TOON-decoding consistency.
Overview
Field
Description
text
Original text… See the full description on the dataset page: https://huggingface.co/datasets/yasserrmd/TOON-Unstructured-Structured.toonosy_dataset
