CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01community-datasets /telugu_newsThis dataset contains Telugu language news articles along with respective topic labels (business, editorial, entertainment, nation, sport) extracted from the daily Andhra Jyoti. This dataset could be used to build Classification and Language Models.text-generation10K<n<100K1 likes175 downloads3y agoHugging Face02SuryaKrishna02 /aya-telugu-news-articles Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.texttext-generation100K<n<1M6 likes168 downloads3y agoHugging Face03community-datasets /telugu_booksThis dataset is created by scraping telugu novels from teluguone.com this dataset can be used for nlp tasks like topic modeling, word embeddings, transfer learning etctext-generationn<1K4 likes144 downloads3y agoHugging Face0410xtechnologieS /telugu_instruction_dataset Telugu Instruction Dataset — Luuka AI Built by 10x Technologies — a curated Telugu-language instruction-tuning dataset for training the Luuka voice assistant. Overview This dataset contains 7,496 instruction–response pairs and 399 multi-turn conversations in Telugu, covering a broad range of natural voice assistant interactions. Every response is written entirely in Telugu script — no English characters appear in any response field. Subset Config Name Pairs… See the full description on the dataset page: https://huggingface.co/datasets/10xtechnologieS/telugu_instruction_dataset.texttext-generation1K<n<10K0 likes85 downloads6mo agoHugging Face05Lots-of-LoRAs /task1037_pib_translation_telugu_urdu Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1037_pib_translation_telugu_urdu Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1037_pib_translation_telugu_urdu.texttext-generationn<1K0 likes69 downloads2y agoHugging Face06indiehackers /TeluguThis dataset is combined version of indiehackers/tenglish_wikipedia and indiehackers/telugu_dataset texttext-generation100K<n<1M1 likes68 downloads2y agoHugging Face07SuryaKrishna02 /aya-telugu-poems Summary aya-telugu-poems is an open source dataset of instruct-style records generated by webscraping a Telugu poems website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview aya-telugu-poems is a corpus… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-poems.texttext-generation1K<n<10K10 likes57 downloads3y agoHugging Face08pulipakav-1 /translated-babylm-telugu Translated BabyLM — Telugu (translated-babylm-telugu) Dataset Description This dataset is a Telugu translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Telugu, following the BabyLM challenge setup. Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B) Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-telugu.texttext-generation10M<n<100M0 likes50 downloads5mo agoHugging Face09TeluguLLMResearch /Padyam2GadyamgatedPreprint: Translating Classical Poetry into Modern Prose (to appear in: Proceedings of INLG 2026) We introduce పద్యం2గద్యం (Padyam2Gadyam), a dataset for the task of poem-to-prose translation in both intra-lingual (TE–TE) and cross-lingual (TE–EN) settings. The poems were sourced from publicly available online archives. 600 samples Sources: 13th-17th Century Telugu language poems Each sample contains a Telugu poem paired with its Telugu prose translation and English prose translation. The… See the full description on the dataset page: https://huggingface.co/datasets/TeluguLLMResearch/Padyam2Gadyam.texttranslationn<1K1 likes50 downloads21d agoHugging Face10narendarcodes /Telugu-MultiTask-Instruct-77K Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset Powered by Adaptive Data — Adaption Labs Dataset Description A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.textquestion-answering10K<n<100K1 likes49 downloads3mo agoHugging Face11SuryaKrishna02 /aya-telugu-food-recipes Summary aya-telugu-food-recipes is an open source dataset of instruct-style records generated by webscraping a Telugu food recipes website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-food-recipes.texttext-generationn<1K3 likes34 downloads3y agoHugging Face12eswardivi /Telugu_InstructDataThis dataset is a translated version of three original datasets, namely HuggingFaceH4/no_robots, databricks/databricks-dolly-15k, and a subset of Telugu from CohereForAI/aya_dataset. It has been curated and processed to create a multilingual avatar dataset. texttext-generation10K<n<100K1 likes30 downloads3y agoHugging Face13desik98 /TeluguRiddles Summary TeluguRiddles is an open source dataset of instruct-style records generated by webscraping multiple riddles websites. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview TeluguRiddles is a corpus of… See the full description on the dataset page: https://huggingface.co/datasets/desik98/TeluguRiddles.texttext-generationn<1K2 likes28 downloads3y agoHugging Face14TeluguLLMResearch /MetricalARGSgatedMETRICALARGS: A Taxonomy for Studying Metrical Poetry with LLMs (Kranti & Vajjala, WILDRE 2026) METRICALARGS: First taxonomy of poetry-related NLP tasks designed to evaluate LLMs on metrical poetry across four dimensions: Analysis, Retrieval, Generation and Support. The dataset includes a pilot evaluation benchmark for Telugu metrical poetry. 169 open-ended questions test.csv ~20 samples for each task across the four categories: Analysis, Retrieval, Generation and Support.… See the full description on the dataset page: https://huggingface.co/datasets/TeluguLLMResearch/MetricalARGS.text-generation0 likes24 downloads1mo agoHugging Face15harsha-desaraju /telugu-texttexttext-generation1M<n<10M0 likes22 downloads6mo agoHugging Face16SuryaKrishna02 /aya-telugu-paraphrase Summary aya-telugu-paraphrase is an open source dataset of instruct-style records generated from the Telugu split of ai4bharat/IndicXParaphrase dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-paraphrase.texttext-generation1K<n<10K4 likes21 downloads3y agoHugging Face17indiehackers /Telugu_InstructDataThis dataset is a translated version of three original datasets, namely HuggingFaceH4/no_robots, databricks/databricks-dolly-15k, and a subset of Telugu from CohereForAI/aya_dataset.. texttext-generation10K<n<100K1 likes20 downloads3y agoHugging Face18Lots-of-LoRAs /task1076_pib_translation_telugu_tamil Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1076_pib_translation_telugu_tamil Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1076_pib_translation_telugu_tamil.texttext-generation1K<n<10K0 likes20 downloads2y agoHugging Face19TeluguLLMResearch /MATAgatedMATA (మాట): Mindful Assessment of the Telugu Abilities of Large Language Models A new evaluation dataset specifically designed to assess a diverse spectrum of linguistic competencies in Telugu. A test suite that directly engages with Telugu-specific grammatical structures, semantic patterns, and culturally embedded language use by sourcing our questions from school grade language textbooks and higher level exams. 729 carefully curated multiple-choice and open-ended questions. Approximately… See the full description on the dataset page: https://huggingface.co/datasets/TeluguLLMResearch/MATA.textquestion-answeringn<1K0 likes19 downloads21d agoHugging Face20Lots-of-LoRAs /task1047_pib_translation_english_telugu Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1047_pib_translation_english_telugu Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1047_pib_translation_english_telugu.texttext-generation1K<n<10K0 likes18 downloads2y agoHugging Face21RohithMidigudla /gemma-health-synthetic-telugu-medmcqa-sft Gemma Health Telugu SFT Splits: train: 17481 rows test: 6150 rows Each row contains: messages: TRL/Unsloth conversational SFT format. text: plain serialized chat text fallback. source, variant, prompt, response: traceability fields. from datasets import load_dataset dataset = load_dataset("RohithMidigudla/gemma-health-synthetic-telugu-medmcqa-sft", split="train", streaming=True) test_dataset = load_dataset("RohithMidigudla/gemma-health-synthetic-telugu-medmcqa-sft"… See the full description on the dataset page: https://huggingface.co/datasets/RohithMidigudla/gemma-health-synthetic-telugu-medmcqa-sft.texttext-generation10K<n<100K0 likes18 downloads4mo agoHugging Face22nlpctx /telugu-qa-codemixed Telugu QA Paraphrases A synthetic multilingual query-rewriting dataset for evaluating retrieval robustness under Telugu-English code mixing. Dataset Description This dataset extends an existing Telugu QA dataset by generating multiple query variants with increasing levels of Telugu-English code mixing. Each example contains: question : Original English question answer : Ground-truth answer level_0 : English paraphrase level_1 : Light Telugu-English code mixing… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/telugu-qa-codemixed.textquestion-answering1K<n<10K0 likes18 downloads3mo agoHugging Face23SuryaKrishna02 /aya-telugu-jokes Summary aya-telugu-jokes is an open source dataset of instruct-style records generated by webscraping a Telugu Jokes website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview aya-telugu-jokes is a corpus… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-jokes.texttext-generationn<1K2 likes17 downloads3y agoHugging Face24RohithMidigudla /gemma-health-telugu-sft Gemma Health Telugu SFT Rows: 8140 Each row contains: messages: TRL/Unsloth conversational SFT format. text: plain serialized chat text fallback. source, variant, prompt, response: traceability fields. from datasets import load_dataset dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft", split="train", streaming=True) texttext-generation1K<n<10K0 likes14 downloads4mo agoHugging Face25RohithMidigudla /gemma-health-telugu-sft-balanced Gemma Health Telugu SFT Splits: train: 175870 rows test: 38010 rows Each row contains: messages: TRL/Unsloth conversational SFT format. text: plain serialized chat text fallback. source, variant, prompt, response: traceability fields. from datasets import load_dataset dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-balanced", split="train", streaming=True) test_dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-balanced", split="test"… See the full description on the dataset page: https://huggingface.co/datasets/RohithMidigudla/gemma-health-telugu-sft-balanced.texttext-generation100K<n<1M0 likes14 downloads4mo agoHugging Face26VenkataRamanaKurumallajaddangi /Pure-Telugu-Alpaca Pure Telugu Alpaca Dataset This dataset is a cleaned version of Telugu-MultiTask-Instruct-77K with Telugu keys. It uses enhanced_prompt as instruction and enhanced_completion as output. Processing Steps Extracted enhanced_prompt → సూచన (instruction) and enhanced_completion → అవుట్పుట్ (output) Filtered to keep only entries with no English letters Removed duplicate entries Normalized whitespace Format Each entry follows the Alpaca format with… See the full description on the dataset page: https://huggingface.co/datasets/VenkataRamanaKurumallajaddangi/Pure-Telugu-Alpaca.texttext-generation10K<n<100K0 likes13 downloads1mo agoHugging Face271-800-SHARED-TASKS /telugu-summarization-generation Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/telugu-summarization-generation.texttext-generation100K<n<1M0 likes12 downloads2y agoHugging Face28Lots-of-LoRAs /task1048_pib_translation_telugu_english Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1048_pib_translation_telugu_english Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1048_pib_translation_telugu_english.texttext-generation1K<n<10K0 likes12 downloads2y agoHugging Face29tadakaluri /TeluguRiddles Summary TeluguRiddles is an open source dataset of instruct-style records generated by webscraping multiple riddles websites. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview TeluguRiddles is a corpus of… See the full description on the dataset page: https://huggingface.co/datasets/tadakaluri/TeluguRiddles.texttext-generationn<1K0 likes12 downloads10mo agoHugging Face30RohithMidigudla /gemma-health-telugu-sft-raw Gemma Health Telugu SFT Splits: train: 287958 rows test: 70002 rows Each row contains: messages: TRL/Unsloth conversational SFT format. text: plain serialized chat text fallback. source, variant, prompt, response: traceability fields. from datasets import load_dataset dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-raw", split="train", streaming=True) test_dataset = load_dataset("RohithMidigudla/gemma-health-telugu-sft-raw", split="test", streaming=True) texttext-generation100K<n<1M0 likes11 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.