datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Punjabi-Gurmukhi-Grammar-Correction-Corpus
ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Engineering Lead
Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Innovation: AMRIT Research OS (100% Locally-Run Autonomous Medical AI)
📖 Overview / ਸੰਖੇਪ
The Punjabi (Gurmukhi) Grammatical Error Correction… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.Sehaj-Gurmukhi-Frontier-Reasoning
🧠 Punjabi (Gurmukhi) Frontier Reasoning & Longevity AI Corpus
Punjabi (Gurmukhi) Frontier Reasoning (ਪੰਜਾਬੀ ਗੁਰਮੁਖੀ ਡਾਟਾਸੈੱਟ) is a high-density, multi-domain Punjabi dataset engineered for training next-generation intelligent Punjabi AI models.
🔍 Search & Discovery Keywords
Language: Punjabi / Gurmukhi (ਪੰਜਾਬੀ / ਗੁਰਮੁਖੀ)
Domains: Quantum Science, Higher Mathematics, Computer Science, Longevity Vitals, and Classical Philosophy in Punjabi.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Sehaj-Gurmukhi-Frontier-Reasoning.amrit-qwythos-corrections
Amrit-Qwythos Corrections Dataset
49 question → correct-answer pairs used to fine-tune Qwythos-9B (a
Qwen3.5-based model) to fix specific, verified hallucinations caught during
real use in the Amrit OS project.
What this fixes
Two real, reproduced failure classes:
Self-identity confusion — at higher sampling temperature, the base
model would sometimes answer "I am Qwen3.5, made by Alibaba Cloud"
instead of its actual fine-tuned identity (Qwythos, by Empero AI).… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/amrit-qwythos-corrections.Fine-TOONing_Datasettoonosy_dataset
