datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
assamese_speech_corpusIndicTTS_Assamese
Assamese Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Assamese monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Assamese
Total Duration: ~27.4 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Assamese.SPRING_INX_Assamese_R1original_data_assamese_ttslma_assamese_clean_datasetAssameseQA
Assamese Question Answering Dataset (AssameseQA)
Dataset Details
Dataset Description
The Assamese Question Answering Dataset (AssameseQA) is an extractive Question Answering (QA) dataset developed for training and evaluating multilingual transformer models on the Assamese language.
The dataset contains Assamese contexts, questions, and answers spanning multiple domains, including:
Assamese history
Geography
Culture
Education
Science
General… See the full description on the dataset page: https://huggingface.co/datasets/sankhyahrick/AssameseQA.Assamese-Text-Dataset-45T-Tokens
I have massive Assamese Dataset nearly about 45.3T (45333004592600) Tokens
It has a lots of Assamese sentances from various sources, 99.9999% of the dataset are cleanned
just download the backup_data.tar.zst file and start using it.
happy training....
My email: ranjitdax89@gmail.com
At least share your opinion… or maybe a simple “thanks” 😄
Topic / Dataset
Tokens
Approx. Scale
Source
Poems Dataset
92.6K
0.0000926B… See the full description on the dataset page: https://huggingface.co/datasets/Ranjit89/Assamese-Text-Dataset-45T-Tokens.xahitya-assamese-corpus
Xahitya Assamese Corpus
A large-scale Assamese literary text corpus scraped from Xahitya.org, containing Assamese prose, essays, stories, poems, and other long-form literary writings.
This dataset is intended for:
Assamese NLP research
Language model pretraining
Tokenizer training
Text generation
Linguistic analysis
Low-resource language AI research
Dataset Structure
The dataset currently contains:
xahitya_dump/
├── articles.jsonl
└── corpus.txt
happy… See the full description on the dataset page: https://huggingface.co/datasets/Ranjit89/xahitya-assamese-corpus.ICON26-COILD-INDIC-MT-Assamese-Bodo
COILD-INDIC-MT 2026 — Assamese–Bodo Dataset
This dataset is provided for the COILD-INDIC-MT 2026 Shared Task, co-located with ICON 2026.
The shared task aims to foster research and innovation in Natural Language Processing (NLP) for Indian Languages.
This repository contains data specifically for the:
Assamese ↔ Bodo
language pair.
🔐 Access to the Dataset
This is a restricted and gated dataset.
Access is available only to authorized participants of the… See the full description on the dataset page: https://huggingface.co/datasets/ainlpml-iitp/ICON26-COILD-INDIC-MT-Assamese-Bodo.assamese-datasetassamese-indicxnli-triplet-random-negatives-10Assamese-IndicXNLI-Triplet-Random-Negatives
Assamese IndicXNLI Triplet Dataset (Random Negatives = 10)
Overview
This dataset is derived from the Assamese portion of the IndicXNLI dataset
(Divyanshu/indicxnli), a
multilingual natural language inference corpus covering 11 Indic languages.
It is specifically constructed for metric learning and contrastive learning
settings such as triplet-loss training.
Each instance contains:
an anchor sentence
a positive sentence (entailment)
10 randomly sampled negative sentences… See the full description on the dataset page: https://huggingface.co/datasets/KhyontekAI/Assamese-IndicXNLI-Triplet-Random-Negatives.assamese-movie-reviews-sentiment
Assamese Movie Reviews — Sentiment Dataset
An original, manually curated and dual-annotated dataset of Assamese-language movie and drama reviews, built to address the near-total absence of sentiment analysis resources for Assamese — a low-resource Indic language spoken by ~15 million people.
Dataset Description
This dataset was constructed from scratch by Avinabh Dutta and Saurav Dutta, as no suitable public resource existed for Assamese sentiment analysis prior… See the full description on the dataset page: https://huggingface.co/datasets/AvinabhDutta-Dev/assamese-movie-reviews-sentiment.lma_assamese_ft_datasetassamese-monolingual-corpus
Assamese Monolingual Corpus (2025)
A high-quality, sentence-level Assamese monolingual dataset containing 1.61 million cleaned, segmented, and deduplicated sentences in Bengali script. This corpus supports Assamese NLP development, language modeling, and public deployment for Northeast India.
Dataset Summary
Language: Assamese (Bengali script)
Size: 1,613,879 sentences
Format: Plain text CSV (text column)
Total tokens: 77,427,585 (using IndicBERTv2 tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/assamese-monolingual-corpus.Vaani-assamese-wancho-nepali-lg-English-no-transcript0assamese-indicxnli-triplet-random-negativesassamese_datasetIndicConformer-Assamese-Logsassamese-wiki-corpus
Assamese Wiki Corpus (ananddey/assamese-wiki-corpus)
A clean, large scale Assamese text corpus spanning wiki articles, literary works, dictionary entries, and quotations, curated for language model pre training, fine tuning, and NLP research.
Total characters: 65,899,040Approximate tokens : 16,474,760 (16.5M)
Fields
Field
Type
Description
id
int64
Wikimedia page ID
title
string
Page title
text
string
Cleaned plain text content
source
class_label… See the full description on the dataset page: https://huggingface.co/datasets/ananddey/assamese-wiki-corpus.Assamese_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 58,478 hours of processed Assamese (AS) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Assamese_Call_Center_Audio_Dataset_Dual_Channel.Assamese-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 58,478 hours of processed Assamese (AS) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Assamese-Call-Center-Audio-Dataset-Single-Channel.asteria-bhojpuri-assamese-civic-qa
Asteria — Bhojpuri & Assamese Civic Q&A Dataset
A dataset of government scheme Q&A pairs in Bhojpuri and Assamese — two low-resource Indian languages.
Dataset Description
This dataset was collected by Asteria, an AI Agent built for the AI Agents Hackathon 2026. The agent helps rural Indian citizens access government welfare schemes by conversing in their native language.
Supported Languages
Bhojpuri (bho) — spoken by 50+ million people in Bihar, UP… See the full description on the dataset page: https://huggingface.co/datasets/Afuu-coder/asteria-bhojpuri-assamese-civic-qa.assamese-asr-dataset
Assamese ASR Dataset
A open curated Automatic Speech Recognition (ASR) dataset for the Assamese language, containing paired speech audio and text transcriptions. This dataset is intended to support research and development of speech recognition systems for Assamese language.
📦 Dataset Overview
Dataset Name: Assamese ASR Dataset
Maintainer: Anand Dey
Language: Assamese (as)
Domain: Speech Recognition
Modalities: Audio, Text
License: Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/ananddey/assamese-asr-dataset.assamese_wikipedia
Assamese Wikipedia Corpus
Dataset Description
The Assamese Wikipedia Corpus is a pure Assamese text dataset derived from the Assamese-language Wikipedia (as.wikipedia.org).
It contains cleaned plain text extracted from Wikipedia articles, stripped of all formatting, with non-Assamese characters completely removed.
This dataset is designed for language modeling, NLP research, creating Assamese specific tokenizers, and other Assamese-language processing tasks.… See the full description on the dataset page: https://huggingface.co/datasets/marsh-mellow/assamese_wikipedia.assamese-monolingual-corpus
Assamese Monolingual Corpus (2025)
A high-quality, sentence-level Assamese monolingual dataset containing 1.61 million cleaned, segmented, and deduplicated sentences in Bengali script. This corpus supports Assamese NLP development, language modeling, and public deployment for Northeast India.
Dataset Summary
Language: Assamese (Bengali script)
Size: 1,613,879 sentences
Format: Plain text CSV (text column)
Total tokens: 77,427,585 (using IndicBERTv2 tokenizer)… See the full description on the dataset page: https://huggingface.co/datasets/jintz0/assamese-monolingual-corpus.assamese-sft-dataset-v1
Assamese SFT Dataset v1 (ananddey/assamese-sft-dataset-v1)
An industry-standard, curated, and deduplicated Supervised Fine-Tuning (SFT) dataset for training Assamese language models and conversational AI assistants.
📊 Dataset Summary
Total Samples: 61,929 instruction-response pairs
Train Split: 58,833 samples
Validation Split: 3,096 samples
Languages: Assamese (as), English (en)
Primary Use Case: SFT / Instruction Fine-Tuning for generative LLMs (e.g. Gemma… See the full description on the dataset page: https://huggingface.co/datasets/ananddey/assamese-sft-dataset-v1.processed_assamese_asr_sllma_assamese_raw_datasetassamese_tts_dataset
