CoolFace
Datasetpublic

narendarcodes/Telugu-MultiTask-Instruct-77K

Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset Powered by Adaptive Data — Adaption Labs Dataset Description A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.

sourceHugging Facecc-by-sa-3.0updated 3mo agoView on Hugging Face
1likes43downloads
Dataset Card

Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset

Powered by Adaptive Data — [Adaption Labs](https://www.adaptionlabs.ai/)

Dataset Description

A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality enhancement.

This dataset was created for the 2026 Adaption AutoScientist Challenge (Language Category).

Dataset Details

FieldValue
Size77,653 rows
FormatJSONL
LanguageTelugu (te)
LicenseCC-BY-SA-3.0

Source Data & Attribution

SourceDescriptionLicense
Telugu Alpaca Cleaned (telugu_ai_instructions)General instruction followingCC-BY-4.0
Dolci SFT Telugu (telugu_instruction_pairs)Instruction-response pairsApache-2.0
TyDi QA Telugu (telugu_qa_news_gen)News-based QAApache-2.0
Aya Telugu Poems & News (telugu_news_and_poetry_gen)Creative writing & journalismApache-2.0
Telugu News Summarization (telugu_news_summaries)Article summarizationMIT
Samanantar EN-TE / Dolly 15K (english_to_telugu_translation)Translation pairsCC-0 / CC-BY-SA-3.0
Telugu QA Andhra Facts (telugu_qa_andhra_facts)Regional knowledge QAMixed Open

Quality Enhancement via Adaption Labs

The dataset was processed through the Adaption Labs Adaptive Data Pipeline:

MetricBeforeAfterChange
GradeCB⬆️
Score6.08.14+35.6%
Percentile18.23

AutoScientist features applied:

  • ✅ Prompt Deduplication
  • ✅ Prompt Rephrase
  • ✅ Hallucination Mitigation
  • ✅ Reasoning Traces

Associated Model

Citation

@misc{golla2026telugu_data,
  title={Telugu MultiTask Instruct 77K Dataset},
  author={Golla Narendar},
  year={2026},
  note={Processed via Adaption Labs AutoScientist. Powered by Adaptive Data.}
}

Powered by Adaptive Data — [Adaption Labs](https://www.adaptionlabs.ai/)