datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BrowseComp-ZH
🧭 BrowseComp-ZH: Benchmarking the Web Browsing Ability of Large Language Models in Chinese
BrowseComp-ZH is the first high-difficulty benchmark specifically designed to evaluate the real-world web browsing and reasoning capabilities of large language models (LLMs) in the Chinese information ecosystem. Inspired by BrowseComp (Wei et al., 2025), BrowseComp-ZH targets the unique linguistic, structural, and retrieval challenges of the Chinese web, including fragmented platforms… See the full description on the dataset page: https://huggingface.co/datasets/PALIN2018/BrowseComp-ZH.coco-paligemmapali_processed_983814pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.pali-tripitaka-thai-script-siamrath-version
📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม)
ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก
🧾 รายการพระไตรปิฎก
📘 เล่ม ๑–๘: วินัยปิฎก
เล่ม ๑: มหาวิภงฺโค (๑)
เล่ม ๒: มหาวิภงฺโค (๒)
เล่ม ๓: ภิกฺขุนีวิภงฺโค
เล่ม ๔: มหาวคฺโค (๑)
เล่ม ๕: มหาวคฺโค (๒)
เล่ม ๖: จุลฺลวคฺโค (๑)
เล่ม ๗: จุลฺลวคฺโค (๒)
เล่ม ๘: ปริวาโร
📗 เล่ม ๙–๒๕: สุตตันตปิฎก
เล่ม ๙–๑๑: ทีฆนิกาย
เล่ม ๑๒–๑๔: มัชฌิมนิกาย
เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.buddhist-classics-vol14-20-pali-tibetan-ja-ko
Buddhist Classics AI Translation Series Vol.14:Pāli Canon in Tibetan, English, Japanese, Korean Version 1.0
Pāli Canon Multi-Language Dataset v1.0
Description: AI-generated parallel translations of Pāli Tipiṭaka in Tibetan, English, Japanese, Korean (Gemini 2.5/2.0). 6 files per source (.pali.split.txt + .en/bo/ja/ko.txt). RAG-ready for cross-language queries.
Files: Upload .7z or split .txt (157MB total).
Usage: from datasets import load_dataset; ds =… See the full description on the dataset page: https://huggingface.co/datasets/ospx1u/buddhist-classics-vol14-20-pali-tibetan-ja-ko.paligemma-multitask-dataset
PaliGemma Multitask Dataset
This dataset is designed for training and evaluating the PaliGemma multitask model for defect detection and analysis. It combines a base set of annotated samples with an extended collection of 874 real-world structural inspection images.
Dataset Description
Overview
The dataset contains images of structural defects along with their corresponding annotations for:
Object detection (bounding boxes)
Defect classification… See the full description on the dataset page: https://huggingface.co/datasets/xingqiang/paligemma-multitask-dataset.pali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tipitaka_pali_in_15_scripts
Tipitaka Pali in 15 Scripts
Dataset Summary
This dataset contains the Pali Tipitaka (The Pali Canon), Commentaries (Aṭṭhakathā), Sub-commentaries (Ṭīkā), and related texts (Anya). It covers the fundamental scriptures of Theravada Buddhism.
This repository serves as a Hugging Face mirror and processed version of the open-source XML data provided by the Vipassana Research Institute (VRI). The texts are available in various scripts (including Roman and Myanmar) and are… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_pali_in_15_scripts.pali_processed_89691desi-paligemma-28b-embeddingspalitoThis paper aims at describing the building of the online corpora on Philippine
languages as part of the online repository system called Palito. There are five components
of the corpora: the top four major Philippine languages which are Tagalog, Cebuano,
Ilocano and Hiligaynon and the Filipino Sign Language (FSL). The four languages are
composed of 250,000-word written texts each, whereas the FSL is composed of seven
thousand signs in video format. Categories of the written texts include creative writing (such
as novels and stories) and religious texts (such as the Bible). Automated tools are provided
for language analysis such as word count, collocates, and others. This is part of a bigger
corpora building project for Philippine languages that would consider text, speech and
video forms, and the corresponding development of automated tools for language analysis
of these various forms.task850_synthetic_longest_palindrome
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task850_synthetic_longest_palindrome
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task850_synthetic_longest_palindrome.license-detection-paligemmadhivehi-layout-syn-lg-paligemma
Synthetic Dhivehi Document Layout Analysis Dataset
Overview
This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-lg-paligemma.pali-sinhalapali_processed_8969181PaliGemma_CLEVR_Datasetpali_processed_7788PAlign-PAPI-personality_prompt.json-cleanedAdapted from
"Personality Alignment of Large Language Models" by Minjun Zhu and Linyi Yang and Yue Zhang
and the associated GitHub repository zhu-minjun/PAlign.
The contents of said repo were declared public domain; in that spirit, this Alpaca-formatted file has also been released as public domain.
pali_processed_8969myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.license-detection-paligemmapali-english-devanagari
Dataset Card for "pali-english-devanagari"
More Information needed
plant-doc-paligemmadhivehi-layout-syn-b1-paligemma
Synthetic Dhivehi Document Layout Analysis Dataset
Overview
This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-b1-paligemma.eval_paligemma_so100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 388,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hrhraj/eval_paligemma_so100.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.brain-tumor-object-detection-paligemmauladzimir-karatkevich-dzikae-paliavanne-karalia-stakha-aleg-garbuz-output_original
Дзікае паляванне Кароля Стаха — арыгінальнае аўдыё
Аўтар / Author: Уладзімір КараткевічМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
uladzimir-karatkevich-dzikae-paliavanne-karalia-stakha-aleg-garbuz-output
Доўгасць аўдыё
7h37m
Радкоў у датасеце
2,585
Структура… See the full description on the dataset page: https://huggingface.co/datasets/fosters/uladzimir-karatkevich-dzikae-paliavanne-karalia-stakha-aleg-garbuz-output_original.
