datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telugu-indicf5-evaluationgsm8k-indic-cultural
GSM8K Indic Cultural Adaptation
Dataset Summary
GSM8K Indic Cultural Adaptation is a culturally localized version of the GSM8K test split, designed to evaluate the robustness of mathematical reasoning models under culturally adapted problem formulations.
The dataset preserves the underlying mathematical reasoning of the original GSM8K benchmark while adapting questions to an Indian context. Depending on the variant, this includes replacing culturally specific… See the full description on the dataset page: https://huggingface.co/datasets/kiranpradeep/gsm8k-indic-cultural.Telugu-LLM-Labs__Indic-gemma-7b-finetuned-sft-Navarasa-2.0-details
Dataset Card for Evaluation run of Telugu-LLM-Labs/Indic-gemma-7b-finetuned-sft-Navarasa-2.0
Dataset automatically created during the evaluation run of model Telugu-LLM-Labs/Indic-gemma-7b-finetuned-sft-Navarasa-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Telugu-LLM-Labs__Indic-gemma-7b-finetuned-sft-Navarasa-2.0-details.Indic_New_dataset_TTS
Indic TTS Dataset Hub (Mozilla)
Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice.
Select the language from the Subset dropdown in the Dataset Viewer.
Columns
audio: WAV audio clip (16kHz)
text: transcription
duration: length in seconds
speaking_rate: characters per second
Telugu-LLM-Labs__Indic-gemma-2b-finetuned-sft-Navarasa-2.0-details
Dataset Card for Evaluation run of Telugu-LLM-Labs/Indic-gemma-2b-finetuned-sft-Navarasa-2.0
Dataset automatically created during the evaluation run of model Telugu-LLM-Labs/Indic-gemma-2b-finetuned-sft-Navarasa-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Telugu-LLM-Labs__Indic-gemma-2b-finetuned-sft-Navarasa-2.0-details.atlas9_mo12_source_indicesmedical-entity-code-mapper-indices
FAISS Indices Directory
This directory contains the pre-built FAISS indices for all medical ontologies.
Included Indices
indices/
├── icd10_bge_m3/ # ICD-10-CM diagnosis codes (4.4GB)
│ ├── faiss.index
│ └── metadata.pkl
├── snomed_bge_m3/ # SNOMED CT clinical concepts (3.1GB)
│ ├── faiss.index
│ └── metadata.pkl
├── loinc_bge_m3/ # LOINC laboratory codes (896MB)
│ ├── faiss.index
│ └── metadata.pkl
├── rxnorm_bge_m3/ # RxNorm medication… See the full description on the dataset page: https://huggingface.co/datasets/docdailey/medical-entity-code-mapper-indices.indicvoices-curated
