CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01user10383 /english-khasi-parallel-corpusThe monolingual data used to generate this dataset was sourced from wikipedia, news websites and open source corpora and then parallelized to generate these synthetic parallel corpora. Sources news_and_wikipedia_corpus: News websites and Simple English Wikipedia news_corpus: Khasi News Websites small sentences: Sentences with 5 words or less sourced from the Samanantar English dataset large sentences: Sentences with more than 5 words sourced from the Samanantar English dataset texttranslation10M<n<100M0 likes41 downloads1y agoHugging Face02toiar /Khasi-OCR-36Kgated Khasi-OCR-36K Khasi-OCR-36K is a Vision-Language dataset designed for OCR, document understanding, and handwriting recognition in the Khasi language, with a smaller subset of English samples. This specific version of the dataset has been pre-filtered and formatted strictly for Vision training (e.g., DeepSeek-VL/OCR). It contains only the Free OCR task, with conversations mapped to the strict <|User|> and <|Assistant|> token format. Images are natively embedded. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OCR-36K.imageimage-to-text10K<n<100K0 likes36 downloads6mo agoHugging Face03MEDHARVIX-SYSTEMS /bhasaflow-khasi-english-parallel-sample-v1 BhasaFlow Khasi-English Parallel Sample v1 A professionally curated, gold-standard parallel speech and text corpus for the Khasi language. Published by Medharvix Systems Private Limited Part of the BhasaFlow Low-Resource Language Technology Initiative Overview This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.audioautomatic-speech-recognitionn<1K23 likes24 downloads5mo agoHugging Face04MEDHARVIX-SYSTEMS /bhasaflow-khasi-monolingual-corpus-v1 BhasaFlow Khasi Monolingual Corpus v1 By Medharvix Systems Private Limited Overview A curated monolingual Khasi text corpus for language modeling, NLP research, and linguistic analysis, with a focus on preserving and digitizing low-resource languages of Northeast India. Dataset Structure Column Description khasi_sentence Khasi language sentence Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-monolingual-corpus-v1.texttext-generationn<1K22 likes24 downloads5mo agoHugging Face05MEDHARVIX-SYSTEMS /bhasaflow-khasi-english-parallel-corpus-v1 BhasaFlow Khasi-English Parallel Corpus v1 By Medharvix Systems Private Limited Overview A curated parallel corpus of Khasi-English sentence pairs designed for machine translation research and development, with a focus on low-resource language technology for Northeast India. Dataset Structure Column Description sentence_id Unique sentence identifier english_text English sentence khasi_text Khasi translation Usage from datasets… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-corpus-v1.texttranslationn<1K19 likes23 downloads5mo agoHugging Face06Bapynshngain /FEM-Khasi-News-Monolingual-Corpusgated Khasi Monolingual News Corpus (740K) Project Attribution & Collaboration This dataset was collected and curated as part of the research project titled "Financial Empowerment in Meghalaya: AI-Powered Multilingual E-Marketplace for Tribes." This project is a collaborative research initiative conducted by: National Law University (NLU) Meghalaya Indian Institute of Information Technology (IIIT) Guwahati Contributors: This dataset is the result of a joint effort by the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/FEM-Khasi-News-Monolingual-Corpus.texttext-generation100K<n<1M0 likes22 downloads5mo agoHugging Face07damerajee /khasi-more-rawtextn<1K0 likes21 downloads3y agoHugging Face08Bapynshngain /English-Khasi-Parallel-Corpus-v1gateddata_source: Web-scraped data manually vetted by me. NIT Silchar’s "EnKhCorp1.0: An English–Khasi Corpus." Tatoeba project. samanantar : the largest publicly available parallel corpora collection for 11 indic languages acknowledgments: Special thanks to Ahlad, NIT Silchar for their "EnKhCorp1.0" dataset, IIT Madras and other institutions involved in the creation of Samanantar and to the contributors of the Tatoeba project. texttranslation10K<n<100K0 likes21 downloads1y agoHugging Face091infinity0 /bhasaflow-khasi-english-parallel-sample-v1 BhasaFlow Khasi-English Parallel Sample v1 A professionally curated, gold-standard parallel speech and text corpus for the Khasi language. Published by Medharvix Systems Private Limited Part of the BhasaFlow Low-Resource Language Technology Initiative Overview This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/1infinity0/bhasaflow-khasi-english-parallel-sample-v1.audioautomatic-speech-recognitionn<1K1 likes19 downloads5mo agoHugging Face10Bapynshngain /khasi-mapped-vectorsgatedtext1K<n<10K0 likes16 downloads26d agoHugging Face11damerajee /khasi-essaystext1K<n<10K0 likes14 downloads3y agoHugging Face12damerajee /khasi-datasets What is Khasi Language? Location: Primarily spoken in the northeastern Indian state of Meghalaya. Also spoken in parts of Assam, Tripura, and Bangladesh. Language Family: Khasi is a member of the Austroasiatic language family. Script: Traditionally written using the Khasi script, which is a script created specifically for the Khasi language. Culture and Identity: The Khasi language is an integral part of the cultural identity of the… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/khasi-datasets.texttext-generation1K<n<10K0 likes12 downloads3y agoHugging Face13RansMairoi /khasi-ner Khasi Named Entity Recognition Corpus A named-entity-annotated corpus for the Khasi language, an Austroasiatic language spoken by approximately 1.6 million people, primarily in the state of Meghalaya in north-east India. To the best of the authors' knowledge, this is the first publicly described NER resource for Khasi. The corpus was constructed as part of the PhD thesis "Machine Learning based Named Entity Recognition for Khasi Language" by Ransly Hoojon (North-Eastern Hill… See the full description on the dataset page: https://huggingface.co/datasets/RansMairoi/khasi-ner.texttoken-classification10K<n<100K0 likes12 downloads4mo agoHugging Face14damerajee /khasi-raw-datatextn<1K0 likes8 downloads3y agoHugging Face15damerajee /more-khasitextn<1K0 likes8 downloads3y agoHugging Face16damerajee /new-khasi-23textn<1K0 likes7 downloads3y agoHugging Face17toiar /Khasi_Synthetic_ASR_Questions_Finalgatedaudio10K<n<100K1 likes7 downloads2mo agoHugging Face18toiar /Khasi-OCR-21Kgated Khasi-OCR-21K Khasi-OCR-21K is a curated Vision-Language dataset totaling 21,319 samples, specifically designed to train robust OCR models for the Khasi language. This version introduces a significant amount of high-quality real book data alongside synthetic samples to handle diverse document conditions. Dataset Split To ensure reliable model evaluation, the dataset is split into: Training Set: ~20,000 samples. Validation Set: 1,319 samples (Randomized with a 40/30/30… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OCR-21K.imageimage-to-text10K<n<100K0 likes6 downloads6mo agoHugging Face19damerajee /new-raw-khasitextn<1K0 likes5 downloads3y agoHugging Face20toiar /Khasi-OmniVoice-TTS-Datagated Khasi Omni Voice TTS Dataset The Khasi Omni Voice dataset is a comprehensive, high-quality audio collection designed specifically for Text-to-Speech (TTS) research and model training in the Khasi language. It features nearly 50 hours of speech data targeting realistic, modern Khasi speech patterns, including natural code-switching. Key Statistics Total Duration: 49 hours, 53 minutes, 19.98 seconds Total Samples: 18,874 distinct audio utterances Language: Khasi… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OmniVoice-TTS-Data.audiotext-to-speech10K<n<100K0 likes5 downloads4mo agoHugging Face21toiar /Khasi_Synthetic_ASR_Norm_Part_1gatedaudio1K<n<10K0 likes5 downloads2mo agoHugging Face22toiar /Khasi_Questions_Part_1gatedaudio1K<n<10K0 likes5 downloads2mo agoHugging Face23toiar /Khasi_Questions_Part_2gatedaudio1K<n<10K0 likes5 downloads2mo agoHugging Face24toiar /Khasi_Questions_Part_3gatedaudio1K<n<10K0 likes5 downloads2mo agoHugging Face25toiar /Khasi_ASR_Final_Combinedgatedaudio100K<n<1M0 likes5 downloads2mo agoHugging Face26toiar /khasi-instruction-response-v2gated Khasi Instruction Response v2 The Khasi Instruction Response v2 dataset is a high-quality, curated collection of 77,810 instruction-response pairs designed to fine-tune Large Language Models (LLMs) for the Khasi language. This is an improved, expanded version of my previous v1 release, offering significantly higher data integrity and broader linguistic coverage. It combines extensive cultural, literary, and translation-based Khasi data with high-reasoning capabilities from… See the full description on the dataset page: https://huggingface.co/datasets/toiar/khasi-instruction-response-v2.texttext-generation10K<n<100K0 likes4 downloads4mo agoHugging Face27toiar /english-khasi-parallel-corpus-43kgatedtext10K<n<100K0 likes3 downloads3mo agoHugging Face28toiar /Khasi_Synthetic_ASR_Norm_Part_2gatedaudio1K<n<10K0 likes3 downloads2mo agoHugging Face29toiar /khasi-instruction-response-v1gated Dataset Card for Khasi Instruction-Response Dataset v1 Dataset Summary Khasi Instruction-Response Dataset v1 is a curated collection of prompts and responses in the Khasi language, designed to support the fine-tuning of instruction-following language models. It includes tasks such as Q&A, translation, summarization, and culturally-grounded dialogue. The dataset reflects indigenous knowledge systems, cultural expressions, and educational content unique to the Khasi context… See the full description on the dataset page: https://huggingface.co/datasets/toiar/khasi-instruction-response-v1.texttranslation1K<n<10K0 likes2 downloads9mo agoHugging Face30RonitMehta260704 /khasi-english-Translation-corpusgatedtext100K<n<1M0 likes2 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.