CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01madhabpaul /assamese_speech_corpusaudiotext-to-speech10K<n<100K4 likes323 downloads2y agoHugging Face02SPRINGLab /IndicTTS_Assamese Assamese Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Assamese monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Assamese Total Duration: ~27.4 hours (Male: 5.16 hours, Female: 5.18 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Assamese.audiotext-to-speech10K<n<100K0 likes315 downloads2y agoHugging Face03SPRINGLab /SPRING_INX_Assamese_R1audio10K<n<100K0 likes154 downloads2y agoHugging Face04SayantanJoker /original_data_assamese_ttsaudio10K<n<100K0 likes73 downloads2y agoHugging Face05vishnu-vizz /lma_assamese_clean_dataset0 likes59 downloads1mo agoHugging Face06sankhyahrick /AssameseQA Assamese Question Answering Dataset (AssameseQA) Dataset Details Dataset Description The Assamese Question Answering Dataset (AssameseQA) is an extractive Question Answering (QA) dataset developed for training and evaluating multilingual transformer models on the Assamese language. The dataset contains Assamese contexts, questions, and answers spanning multiple domains, including: Assamese history Geography Culture Education Science General… See the full description on the dataset page: https://huggingface.co/datasets/sankhyahrick/AssameseQA.text1K<n<10K1 likes57 downloads2mo agoHugging Face07Ranjit89 /Assamese-Text-Dataset-45T-Tokens I have massive Assamese Dataset nearly about 45.3T (45333004592600) Tokens It has a lots of Assamese sentances from various sources, 99.9999% of the dataset are cleanned just download the backup_data.tar.zst file and start using it. happy training.... My email: ranjitdax89@gmail.com At least share your opinion… or maybe a simple “thanks” 😄 Topic / Dataset Tokens Approx. Scale Source Poems Dataset 92.6K 0.0000926B… See the full description on the dataset page: https://huggingface.co/datasets/Ranjit89/Assamese-Text-Dataset-45T-Tokens.texttext-generationn<1K0 likes54 downloads4mo agoHugging Face08Ranjit89 /xahitya-assamese-corpus Xahitya Assamese Corpus A large-scale Assamese literary text corpus scraped from Xahitya.org, containing Assamese prose, essays, stories, poems, and other long-form literary writings. This dataset is intended for: Assamese NLP research Language model pretraining Tokenizer training Text generation Linguistic analysis Low-resource language AI research Dataset Structure The dataset currently contains: xahitya_dump/ ├── articles.jsonl └── corpus.txt happy… See the full description on the dataset page: https://huggingface.co/datasets/Ranjit89/xahitya-assamese-corpus.texttext-generation1K<n<10K0 likes41 downloads4mo agoHugging Face09ainlpml-iitp /ICON26-COILD-INDIC-MT-Assamese-Bodogated COILD-INDIC-MT 2026 — Assamese–Bodo Dataset This dataset is provided for the COILD-INDIC-MT 2026 Shared Task, co-located with ICON 2026. The shared task aims to foster research and innovation in Natural Language Processing (NLP) for Indian Languages. This repository contains data specifically for the: Assamese ↔ Bodo language pair. 🔐 Access to the Dataset This is a restricted and gated dataset. Access is available only to authorized participants of the… See the full description on the dataset page: https://huggingface.co/datasets/ainlpml-iitp/ICON26-COILD-INDIC-MT-Assamese-Bodo.texttranslation10K<n<100K0 likes41 downloads14d agoHugging Face10jenil17 /assamese-datasettext1M<n<10M0 likes41 downloads16d agoHugging Face11pritamdeka /assamese-indicxnli-triplet-random-negatives-10text1M<n<10M0 likes38 downloads2y agoHugging Face12KhyontekAI /Assamese-IndicXNLI-Triplet-Random-Negatives Assamese IndicXNLI Triplet Dataset (Random Negatives = 10) Overview This dataset is derived from the Assamese portion of the IndicXNLI dataset (Divyanshu/indicxnli), a multilingual natural language inference corpus covering 11 Indic languages. It is specifically constructed for metric learning and contrastive learning settings such as triplet-loss training. Each instance contains: an anchor sentence a positive sentence (entailment) 10 randomly sampled negative sentences… See the full description on the dataset page: https://huggingface.co/datasets/KhyontekAI/Assamese-IndicXNLI-Triplet-Random-Negatives.textsentence-similarity1M<n<10M0 likes36 downloads8mo agoHugging Face13AvinabhDutta-Dev /assamese-movie-reviews-sentiment Assamese Movie Reviews — Sentiment Dataset An original, manually curated and dual-annotated dataset of Assamese-language movie and drama reviews, built to address the near-total absence of sentiment analysis resources for Assamese — a low-resource Indic language spoken by ~15 million people. Dataset Description This dataset was constructed from scratch by Avinabh Dutta and Saurav Dutta, as no suitable public resource existed for Assamese sentiment analysis prior… See the full description on the dataset page: https://huggingface.co/datasets/AvinabhDutta-Dev/assamese-movie-reviews-sentiment.texttext-classification1K<n<10K1 likes33 downloads21d agoHugging Face14vishnu-vizz /lma_assamese_ft_datasettext1K<n<10K0 likes33 downloads6d agoHugging Face15MWirelabs /assamese-monolingual-corpus Assamese Monolingual Corpus (2025) A high-quality, sentence-level Assamese monolingual dataset containing 1.61 million cleaned, segmented, and deduplicated sentences in Bengali script. This corpus supports Assamese NLP development, language modeling, and public deployment for Northeast India. Dataset Summary Language: Assamese (Bengali script) Size: 1,613,879 sentences Format: Plain text CSV (text column) Total tokens: 77,427,585 (using IndicBERTv2 tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/assamese-monolingual-corpus.texttext-generation1M<n<10M1 likes31 downloads10mo agoHugging Face16dianavdavidson /Vaani-assamese-wancho-nepali-lg-English-no-transcript0audio10K<n<100K0 likes31 downloads4mo agoHugging Face17pritamdeka /assamese-indicxnli-triplet-random-negativestext100K<n<1M0 likes30 downloads2y agoHugging Face18Sajjo /assamese_datasetaudio10K<n<100K0 likes29 downloads1y agoHugging Face19bijaykumarsingh /IndicConformer-Assamese-Logs0 likes28 downloads29d agoHugging Face20ananddey /assamese-wiki-corpus Assamese Wiki Corpus (ananddey/assamese-wiki-corpus) A clean, large scale Assamese text corpus spanning wiki articles, literary works, dictionary entries, and quotations, curated for language model pre training, fine tuning, and NLP research. Total characters: 65,899,040Approximate tokens : 16,474,760 (16.5M) Fields Field Type Description id int64 Wikimedia page ID title string Page title text string Cleaned plain text content source class_label… See the full description on the dataset page: https://huggingface.co/datasets/ananddey/assamese-wiki-corpus.texttext-generation10K<n<100K1 likes27 downloads2mo agoHugging Face21InfoBayAI /Assamese_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 58,478 hours of processed Assamese (AS) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Assamese_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes26 downloads7d agoHugging Face22InfoBayAI /Assamese-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 58,478 hours of processed Assamese (AS) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Assamese-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes25 downloads7d agoHugging Face23Afuu-coder /asteria-bhojpuri-assamese-civic-qa Asteria — Bhojpuri & Assamese Civic Q&A Dataset A dataset of government scheme Q&A pairs in Bhojpuri and Assamese — two low-resource Indian languages. Dataset Description This dataset was collected by Asteria, an AI Agent built for the AI Agents Hackathon 2026. The agent helps rural Indian citizens access government welfare schemes by conversing in their native language. Supported Languages Bhojpuri (bho) — spoken by 50+ million people in Bihar, UP… See the full description on the dataset page: https://huggingface.co/datasets/Afuu-coder/asteria-bhojpuri-assamese-civic-qa.tabular1K<n<10K0 likes25 downloads3mo agoHugging Face24ananddey /assamese-asr-datasetgated Assamese ASR Dataset A open curated Automatic Speech Recognition (ASR) dataset for the Assamese language, containing paired speech audio and text transcriptions. This dataset is intended to support research and development of speech recognition systems for Assamese language. 📦 Dataset Overview Dataset Name: Assamese ASR Dataset Maintainer: Anand Dey Language: Assamese (as) Domain: Speech Recognition Modalities: Audio, Text License: Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/ananddey/assamese-asr-dataset.audio100K<n<1M1 likes24 downloads9mo agoHugging Face25marsh-mellow /assamese_wikipedia Assamese Wikipedia Corpus Dataset Description The Assamese Wikipedia Corpus is a pure Assamese text dataset derived from the Assamese-language Wikipedia (as.wikipedia.org). It contains cleaned plain text extracted from Wikipedia articles, stripped of all formatting, with non-Assamese characters completely removed. This dataset is designed for language modeling, NLP research, creating Assamese specific tokenizers, and other Assamese-language processing tasks.… See the full description on the dataset page: https://huggingface.co/datasets/marsh-mellow/assamese_wikipedia.texttext-generation10K<n<100K1 likes23 downloads1y agoHugging Face26jintz0 /assamese-monolingual-corpus Assamese Monolingual Corpus (2025) A high-quality, sentence-level Assamese monolingual dataset containing 1.61 million cleaned, segmented, and deduplicated sentences in Bengali script. This corpus supports Assamese NLP development, language modeling, and public deployment for Northeast India. Dataset Summary Language: Assamese (Bengali script) Size: 1,613,879 sentences Format: Plain text CSV (text column) Total tokens: 77,427,585 (using IndicBERTv2 tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/jintz0/assamese-monolingual-corpus.texttext-generation1M<n<10M0 likes23 downloads4mo agoHugging Face27ananddey /assamese-sft-dataset-v1 Assamese SFT Dataset v1 (ananddey/assamese-sft-dataset-v1) An industry-standard, curated, and deduplicated Supervised Fine-Tuning (SFT) dataset for training Assamese language models and conversational AI assistants. 📊 Dataset Summary Total Samples: 61,929 instruction-response pairs Train Split: 58,833 samples Validation Split: 3,096 samples Languages: Assamese (as), English (en) Primary Use Case: SFT / Instruction Fine-Tuning for generative LLMs (e.g. Gemma… See the full description on the dataset page: https://huggingface.co/datasets/ananddey/assamese-sft-dataset-v1.tabulartext-generation10K<n<100K0 likes22 downloads1mo agoHugging Face28dmedhi /processed_assamese_asr_sl10K<n<100K0 likes21 downloads10mo agoHugging Face29vishnu-vizz /lma_assamese_raw_dataset0 likes21 downloads1mo agoHugging Face30madhabpaul /assamese_tts_datasetaudio1K<n<10K0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.