datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telugu_newsThis dataset contains Telugu language news articles along with respective
topic labels (business, editorial, entertainment, nation, sport) extracted from
the daily Andhra Jyoti. This dataset could be used to build Classification and Language Models.aya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.telugu_booksThis dataset is created by scraping telugu novels from teluguone.com this dataset can be used for nlp tasks like topic modeling, word embeddings, transfer learning etctelugu_instruction_dataset
Telugu Instruction Dataset — Luuka AI
Built by 10x Technologies — a curated Telugu-language instruction-tuning dataset for training the Luuka voice assistant.
Overview
This dataset contains 7,496 instruction–response pairs and 399 multi-turn conversations in Telugu, covering a broad range of natural voice assistant interactions. Every response is written entirely in Telugu script — no English characters appear in any response field.
Subset
Config Name
Pairs… See the full description on the dataset page: https://huggingface.co/datasets/10xtechnologieS/telugu_instruction_dataset.task1037_pib_translation_telugu_urdu
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1037_pib_translation_telugu_urdu
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1037_pib_translation_telugu_urdu.TeluguThis dataset is combined version of indiehackers/tenglish_wikipedia and indiehackers/telugu_dataset
aya-telugu-poems
Summary
aya-telugu-poems is an open source dataset of instruct-style records generated by webscraping a Telugu poems website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview
aya-telugu-poems is a corpus… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-poems.translated-babylm-telugu
Translated BabyLM — Telugu (translated-babylm-telugu)
Dataset Description
This dataset is a Telugu translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Telugu, following the BabyLM challenge setup.
Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B)
Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-telugu.Padyam2GadyamPreprint: Translating Classical Poetry into Modern Prose (to appear in: Proceedings of INLG 2026)
We introduce పద్యం2గద్యం (Padyam2Gadyam), a dataset for the task of poem-to-prose translation in both intra-lingual (TE–TE) and cross-lingual (TE–EN) settings.
The poems were sourced from publicly available online archives.
600 samples
Sources: 13th-17th Century Telugu language poems
Each sample contains a Telugu poem paired with its Telugu prose translation and English prose translation.
The… See the full description on the dataset page: https://huggingface.co/datasets/TeluguLLMResearch/Padyam2Gadyam.Telugu-MultiTask-Instruct-77K
Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset
Powered by Adaptive Data — Adaption Labs
Dataset Description
A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.aya-telugu-food-recipes
Summary
aya-telugu-food-recipes is an open source dataset of instruct-style records generated by webscraping a Telugu food recipes website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-food-recipes.Telugu_InstructDataThis dataset is a translated version of three original datasets, namely HuggingFaceH4/no_robots, databricks/databricks-dolly-15k, and a subset of Telugu from CohereForAI/aya_dataset. It has been curated and processed to create a multilingual avatar dataset.
TeluguRiddles
Summary
TeluguRiddles is an open source dataset of instruct-style records generated by webscraping multiple riddles websites. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview
TeluguRiddles is a corpus of… See the full description on the dataset page: https://huggingface.co/datasets/desik98/TeluguRiddles.MetricalARGSMETRICALARGS: A Taxonomy for Studying Metrical Poetry with LLMs (Kranti & Vajjala, WILDRE 2026)
METRICALARGS: First taxonomy of poetry-related NLP tasks designed to evaluate LLMs on metrical poetry across four dimensions: Analysis, Retrieval, Generation and Support.
The dataset includes a pilot evaluation benchmark for Telugu metrical poetry.
169 open-ended questions
test.csv
~20 samples for each task across the four categories: Analysis, Retrieval, Generation and Support.… See the full description on the dataset page: https://huggingface.co/datasets/TeluguLLMResearch/MetricalARGS.telugu-textaya-telugu-paraphrase
Summary
aya-telugu-paraphrase is an open source dataset of instruct-style records generated from the Telugu split of ai4bharat/IndicXParaphrase dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-paraphrase.Telugu_InstructDataThis dataset is a translated version of three original datasets, namely HuggingFaceH4/no_robots, databricks/databricks-dolly-15k, and a subset of Telugu from CohereForAI/aya_dataset..
task1076_pib_translation_telugu_tamil
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1076_pib_translation_telugu_tamil
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1076_pib_translation_telugu_tamil.MATAMATA (మాట): Mindful Assessment of the Telugu Abilities of Large Language Models
A new evaluation dataset specifically designed to assess a diverse spectrum of linguistic competencies in Telugu.
A test suite that directly engages with Telugu-specific grammatical structures, semantic patterns, and culturally embedded language use by sourcing our questions from school grade language textbooks and higher level exams.
729 carefully curated multiple-choice and open-ended questions.
Approximately… See the full description on the dataset page: https://huggingface.co/datasets/TeluguLLMResearch/MATA.task1047_pib_translation_english_telugu
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1047_pib_translation_english_telugu
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1047_pib_translation_english_telugu.gemma-health-synthetic-telugu-medmcqa-sft
Gemma Health Telugu SFT
Splits:
train: 17481 rows
test: 6150 rows
Each row contains:
messages: TRL/Unsloth conversational SFT format.
text: plain serialized chat text fallback.
source, variant, prompt, response: traceability fields.
from datasets import load_dataset
dataset = load_dataset("RohithMidigudla/gemma-health-synthetic-telugu-medmcqa-sft", split="train", streaming=True)
test_dataset = load_dataset("RohithMidigudla/gemma-health-synthetic-telugu-medmcqa-sft"… See the full description on the dataset page: https://huggingface.co/datasets/RohithMidigudla/gemma-health-synthetic-telugu-medmcqa-sft.telugu-qa-codemixed
Telugu QA Paraphrases
A synthetic multilingual query-rewriting dataset for evaluating retrieval robustness under Telugu-English code mixing.
Dataset Description
This dataset extends an existing Telugu QA dataset by generating multiple query variants with increasing levels of Telugu-English code mixing.
Each example contains:
question : Original English question
answer : Ground-truth answer
level_0 : English paraphrase
level_1 : Light Telugu-English code mixing… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/telugu-qa-codemixed.aya-telugu-jokes
Summary
aya-telugu-jokes is an open source dataset of instruct-style records generated by webscraping a Telugu Jokes website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview
aya-telugu-jokes is a corpus… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-jokes.gemma-health-telugu-sft
Gemma Health Telugu SFT
Rows: 8140
Each row contains:
messages: TRL/Unsloth conversational SFT format.
text: plain serialized chat text fallback.
source, variant, prompt, response: traceability fields.
from datasets import load_dataset
dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft", split="train", streaming=True)
gemma-health-telugu-sft-balanced
Gemma Health Telugu SFT
Splits:
train: 175870 rows
test: 38010 rows
Each row contains:
messages: TRL/Unsloth conversational SFT format.
text: plain serialized chat text fallback.
source, variant, prompt, response: traceability fields.
from datasets import load_dataset
dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-balanced", split="train", streaming=True)
test_dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-balanced", split="test"… See the full description on the dataset page: https://huggingface.co/datasets/RohithMidigudla/gemma-health-telugu-sft-balanced.Pure-Telugu-Alpaca
Pure Telugu Alpaca Dataset
This dataset is a cleaned version of Telugu-MultiTask-Instruct-77K with Telugu keys.
It uses enhanced_prompt as instruction and enhanced_completion as output.
Processing Steps
Extracted enhanced_prompt → సూచన (instruction) and enhanced_completion → అవుట్పుట్ (output)
Filtered to keep only entries with no English letters
Removed duplicate entries
Normalized whitespace
Format
Each entry follows the Alpaca format with… See the full description on the dataset page: https://huggingface.co/datasets/VenkataRamanaKurumallajaddangi/Pure-Telugu-Alpaca.telugu-summarization-generation
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/telugu-summarization-generation.task1048_pib_translation_telugu_english
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1048_pib_translation_telugu_english
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1048_pib_translation_telugu_english.TeluguRiddles
Summary
TeluguRiddles is an open source dataset of instruct-style records generated by webscraping multiple riddles websites. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview
TeluguRiddles is a corpus of… See the full description on the dataset page: https://huggingface.co/datasets/tadakaluri/TeluguRiddles.gemma-health-telugu-sft-raw
Gemma Health Telugu SFT
Splits:
train: 287958 rows
test: 70002 rows
Each row contains:
messages: TRL/Unsloth conversational SFT format.
text: plain serialized chat text fallback.
source, variant, prompt, response: traceability fields.
from datasets import load_dataset
dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-raw", split="train", streaming=True)
test_dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-raw", split="test", streaming=True)
