datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Personas-India
Nemotron-Personas-India
A compound AI approach to personas grounded in real-world distributions
वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण
Dataset Overview (डेटासेट अवलोकन)
Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and richness of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-India.t5gemma2-indonesia-instruct-v1
T5Gemma-2 Indonesian Instruct — Mono-Repo
Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia.
Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder,
setiap config = folder dan berisi split train + validation (80:20) di level percakapan.
Struktur (by fungsi)
t5gemma2-indonesia-instruct-v1/
├── README.md
├── manifest.json
├── chat_idx_map.json
├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.Nemotron-Personas-India
Nemotron-Personas-India
A compound AI approach to personas grounded in real-world distributions
वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण
Dataset Overview (डेटासेट अवलोकन)
Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/Nemotron-Personas-India.medical-history-of-british-india
A Medical History of British India Dataset
Dataset Description
This dataset contains digitiaed official publications documenting medical research and public health in British India from 1850-1950. The collection represents a crucial period in medical history, capturing the transition from humoral to biochemical medical traditions and documenting major breakthroughs in bacteriology, parasitology, and vaccine development. These documents provide invaluable insights into… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/medical-history-of-british-india.indic-ocr-bench
Sarvam Indic OCR Bench
Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material.
The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.AI-Agent-Marketplace-Index
AI Agent Marketplace & Store Index An Open Source Collections of AI Agent Meta and Metric information
Github| Huggingface | Pypi | Open Source AI Agent Marketplace & Store | Agent RL Dataset
News We released cli tool 'agtm' GitHub to submit and manage AI Agent meta submission and access.
AI Agent Marketplace Index DataSet
This DeepNLP AI Agent Marketplace dataset contains more than 10k+ AI Agent Meta information covering 30+ categories from Open AI Agent… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/AI-Agent-Marketplace-Index.
