datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telugu-monolingual-datasettelugu_instruction_dataset
Telugu Instruction Dataset — Luuka AI
Built by 10x Technologies — a curated Telugu-language instruction-tuning dataset for training the Luuka voice assistant.
Overview
This dataset contains 7,496 instruction–response pairs and 399 multi-turn conversations in Telugu, covering a broad range of natural voice assistant interactions. Every response is written entirely in Telugu script — no English characters appear in any response field.
Subset
Config Name
Pairs… See the full description on the dataset page: https://huggingface.co/datasets/10xtechnologieS/telugu_instruction_dataset.bhagavath-gita-telugutelugu-indicf5-evaluationTelugu-MultiTask-Instruct-77K
Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset
Powered by Adaptive Data — Adaption Labs
Dataset Description
A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.telugutechbadi-gk
Telugu GK Questions Dataset
Overview
This dataset consists of General Knowledge (GK) questions scraped from the Telugu Tech Badi website.
A separate data cleaning script refines the extracted questions for better readability and analysis.
Tasks
Task
Objective: Extract GK questions from a list of URLs.
Challenges: Some of the URLs follow a different format than others, so modify the code for specific URLs.
Colab Notebook: Modifying the .jsonl file… See the full description on the dataset page: https://huggingface.co/datasets/haripritam/telugutechbadi-gk.BPCC_TeluguIndicNLP-Telugutelugu-clinical-math-reasoning-v1adaption-telugu-normalization
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-telugu_normalization
This dataset contains a collection of user queries and statements written in Telugu, covering diverse topics such as mobile troubleshooting, cyber security threats, insurance renewals, and emergency services. The samples vary from short keywords to detailed problem descriptions involving scams, natural disasters, and administrative procedures. It represents… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-normalization.gemma-health-synthetic-telugu-medmcqa-sft
Gemma Health Telugu SFT
Splits:
train: 17481 rows
test: 6150 rows
Each row contains:
messages: TRL/Unsloth conversational SFT format.
text: plain serialized chat text fallback.
source, variant, prompt, response: traceability fields.
from datasets import load_dataset
dataset = load_dataset("RohithMidigudla/gemma-health-synthetic-telugu-medmcqa-sft", split="train", streaming=True)
test_dataset = load_dataset("RohithMidigudla/gemma-health-synthetic-telugu-medmcqa-sft"… See the full description on the dataset page: https://huggingface.co/datasets/RohithMidigudla/gemma-health-synthetic-telugu-medmcqa-sft.GPTeacher-TeluguSpeak_teluguadaption-telugu-tamil-translation
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-telugu_tamil_translation
This dataset contains parallel sentence pairs translating travel-related narratives from Telugu to Tamil. The content covers diverse Indian destinations, spiritual experiences, cultural festivals, and personal reflections on tourism. Each sample consists of a Telugu prompt describing a specific journey or observation, paired with its corresponding Tamil… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-tamil-translation.gemma-health-telugu-sft
Gemma Health Telugu SFT
Rows: 8140
Each row contains:
messages: TRL/Unsloth conversational SFT format.
text: plain serialized chat text fallback.
source, variant, prompt, response: traceability fields.
from datasets import load_dataset
dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft", split="train", streaming=True)
gemma-health-telugu-sft-balanced
Gemma Health Telugu SFT
Splits:
train: 175870 rows
test: 38010 rows
Each row contains:
messages: TRL/Unsloth conversational SFT format.
text: plain serialized chat text fallback.
source, variant, prompt, response: traceability fields.
from datasets import load_dataset
dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-balanced", split="train", streaming=True)
test_dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-balanced", split="test"… See the full description on the dataset page: https://huggingface.co/datasets/RohithMidigudla/gemma-health-telugu-sft-balanced.Cultural-Safety-Telugu
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-telugu_safety_prompts
A dataset of Telugu-language prompts labeled as SAFE or UNSAFE based on content safety. It includes text prompts and their corresponding safety annotations, intended for training or evaluating content moderation systems. The dataset focuses on identifying harmful or inappropriate content in Telugu.
Dataset size
There are 13,920… See the full description on the dataset page: https://huggingface.co/datasets/salmankhanpm/Cultural-Safety-Telugu.telugu_high_quality_convoPure-Telugu-Alpaca
Pure Telugu Alpaca Dataset
This dataset is a cleaned version of Telugu-MultiTask-Instruct-77K with Telugu keys.
It uses enhanced_prompt as instruction and enhanced_completion as output.
Processing Steps
Extracted enhanced_prompt → సూచన (instruction) and enhanced_completion → అవుట్పుట్ (output)
Filtered to keep only entries with no English letters
Removed duplicate entries
Normalized whitespace
Format
Each entry follows the Alpaca format with… See the full description on the dataset page: https://huggingface.co/datasets/VenkataRamanaKurumallajaddangi/Pure-Telugu-Alpaca.Telugu-LLM-Labs__Indic-gemma-7b-finetuned-sft-Navarasa-2.0-details
Dataset Card for Evaluation run of Telugu-LLM-Labs/Indic-gemma-7b-finetuned-sft-Navarasa-2.0
Dataset automatically created during the evaluation run of model Telugu-LLM-Labs/Indic-gemma-7b-finetuned-sft-Navarasa-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Telugu-LLM-Labs__Indic-gemma-7b-finetuned-sft-Navarasa-2.0-details.TeluguTinnystoriesadaption-telugu-to-tamil-travel
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-telugu_to_tamil_travel
This dataset contains parallel sentence pairs translating travel-related narratives from Telugu to Tamil. The content covers diverse Indian destinations, spiritual experiences, cultural festivals, and personal reflections on tourism. Each sample consists of a Telugu prompt describing a specific journey or observation, paired with its corresponding Tamil… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-to-tamil-travel.BPCC_Telugugemma-health-telugu-sft-raw
Gemma Health Telugu SFT
Splits:
train: 287958 rows
test: 70002 rows
Each row contains:
messages: TRL/Unsloth conversational SFT format.
text: plain serialized chat text fallback.
source, variant, prompt, response: traceability fields.
from datasets import load_dataset
dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-raw", split="train", streaming=True)
test_dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-raw", split="test", streaming=True)
Telugu-LLM-Labs__Indic-gemma-2b-finetuned-sft-Navarasa-2.0-details
Dataset Card for Evaluation run of Telugu-LLM-Labs/Indic-gemma-2b-finetuned-sft-Navarasa-2.0
Dataset automatically created during the evaluation run of model Telugu-LLM-Labs/Indic-gemma-2b-finetuned-sft-Navarasa-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Telugu-LLM-Labs__Indic-gemma-2b-finetuned-sft-Navarasa-2.0-details.adaption_telugu_rural_crisis_reasoning_v1Telugu-Dpodata_telugu_system_v8_01.jsonadaption-telugu-user-queries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-telugu_user_queries
This dataset contains a collection of user queries and statements written in Telugu, covering diverse topics such as mobile troubleshooting, cyber security threats, insurance renewals, and emergency services. The samples vary from short keywords to detailed problem descriptions involving scams, natural disasters, and administrative procedures. It represents… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-user-queries.data_telugu_db_v7_01.json
