CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PALIN2018 /BrowseComp-ZH 🧭 BrowseComp-ZH: Benchmarking the Web Browsing Ability of Large Language Models in Chinese BrowseComp-ZH is the first high-difficulty benchmark specifically designed to evaluate the real-world web browsing and reasoning capabilities of large language models (LLMs) in the Chinese information ecosystem. Inspired by BrowseComp (Wei et al., 2025), BrowseComp-ZH targets the unique linguistic, structural, and retrieval challenges of the Chinese web, including fragmented platforms… See the full description on the dataset page: https://huggingface.co/datasets/PALIN2018/BrowseComp-ZH.textquestion-answeringn<1K7 likes2.8k downloads1y agoHugging Face02savoji /coco-paligemmaimage100K<n<1M0 likes616 downloads1y agoHugging Face03rahul77 /pali_processed_983814image10K<n<100K0 likes395 downloads2y agoHugging Face04uisp /pali-tripitaka-thai-script-siamrath-version Multi-File CSV Dataset คำอธิบาย พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 ... คำอธิบายของแต่ละเล่ม เล่ม ๑: วินย. มหาวิภงฺโค (๑) เล่ม ๒: วินย. มหาวิภงฺโค (๒) เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค เล่ม ๔: วินย. มหาวคฺโค (๑) เล่ม ๕: วินย. มหาวคฺโค (๒) เล่ม ๖: วินย. จุลฺลวคฺโค (๑) เล่ม ๗: วินย. จุลฺลวคฺโค (๒) เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.tabular100K<n<1M3 likes284 downloads2y agoHugging Face05mgprogm /pali-tripitaka-thai-script-siamrath-version 📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม) ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก 🧾 รายการพระไตรปิฎก 📘 เล่ม ๑–๘: วินัยปิฎก เล่ม ๑: มหาวิภงฺโค (๑) เล่ม ๒: มหาวิภงฺโค (๒) เล่ม ๓: ภิกฺขุนีวิภงฺโค เล่ม ๔: มหาวคฺโค (๑) เล่ม ๕: มหาวคฺโค (๒) เล่ม ๖: จุลฺลวคฺโค (๑) เล่ม ๗: จุลฺลวคฺโค (๒) เล่ม ๘: ปริวาโร 📗 เล่ม ๙–๒๕: สุตตันตปิฎก เล่ม ๙–๑๑: ทีฆนิกาย เล่ม ๑๒–๑๔: มัชฌิมนิกาย เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.tabular100K<n<1M1 likes210 downloads1y agoHugging Face06ospx1u /buddhist-classics-vol14-20-pali-tibetan-ja-ko Buddhist Classics AI Translation Series Vol.14:Pāli Canon in Tibetan, English, Japanese, Korean Version 1.0 Pāli Canon Multi-Language Dataset v1.0 Description: AI-generated parallel translations of Pāli Tipiṭaka in Tibetan, English, Japanese, Korean (Gemini 2.5/2.0). 6 files per source (.pali.split.txt + .en/bo/ja/ko.txt). RAG-ready for cross-language queries. Files: Upload .7z or split .txt (157MB total). Usage: from datasets import load_dataset; ds =… See the full description on the dataset page: https://huggingface.co/datasets/ospx1u/buddhist-classics-vol14-20-pali-tibetan-ja-ko.0 likes167 downloads6mo agoHugging Face07xingqiang /paligemma-multitask-dataset PaliGemma Multitask Dataset This dataset is designed for training and evaluating the PaliGemma multitask model for defect detection and analysis. It combines a base set of annotated samples with an extended collection of 874 real-world structural inspection images. Dataset Description Overview The dataset contains images of structural defects along with their corresponding annotations for: Object detection (bounding boxes) Defect classification… See the full description on the dataset page: https://huggingface.co/datasets/xingqiang/paligemma-multitask-dataset.imageobject-detection1K<n<10K0 likes161 downloads2y agoHugging Face08uisp /pali-commentary-thai-script-siamrath-version Multi-File CSV Dataset คำอธิบาย อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 คำอธิบายของแต่ละเล่ม เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑) เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒) เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓) เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑) เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒) เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tabular100K<n<1M3 likes146 downloads2y agoHugging Face09freococo /tipitaka_pali_in_15_scripts Tipitaka Pali in 15 Scripts Dataset Summary This dataset contains the Pali Tipitaka (The Pali Canon), Commentaries (Aṭṭhakathā), Sub-commentaries (Ṭīkā), and related texts (Anya). It covers the fundamental scriptures of Theravada Buddhism. This repository serves as a Hugging Face mirror and processed version of the open-source XML data provided by the Vipassana Research Institute (VRI). The texts are available in various scripts (including Roman and Myanmar) and are… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_pali_in_15_scripts.text100K<n<1M1 likes106 downloads8mo agoHugging Face10rahul77 /pali_processed_89691image100K<n<1M0 likes96 downloads2y agoHugging Face11Boson20 /desi-paligemma-28b-embeddingstimeseries10K<n<100K0 likes93 downloads6mo agoHugging Face12SEACrowd /palitoThis paper aims at describing the building of the online corpora on Philippine languages as part of the online repository system called Palito. There are five components of the corpora: the top four major Philippine languages which are Tagalog, Cebuano, Ilocano and Hiligaynon and the Filipino Sign Language (FSL). The four languages are composed of 250,000-word written texts each, whereas the FSL is composed of seven thousand signs in video format. Categories of the written texts include creative writing (such as novels and stories) and religious texts (such as the Bible). Automated tools are provided for language analysis such as word count, collocates, and others. This is part of a bigger corpora building project for Philippine languages that would consider text, speech and video forms, and the corresponding development of automated tools for language analysis of these various forms.0 likes87 downloads2y agoHugging Face13Lots-of-LoRAs /task850_synthetic_longest_palindrome Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task850_synthetic_longest_palindrome Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task850_synthetic_longest_palindrome.texttext-generation1K<n<10K0 likes83 downloads2y agoHugging Face14ariG23498 /license-detection-paligemmaimage1K<n<10K9 likes64 downloads2y agoHugging Face15alakxender /dhivehi-layout-syn-lg-paligemma Synthetic Dhivehi Document Layout Analysis Dataset Overview This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-lg-paligemma.image10K<n<100K0 likes58 downloads1y agoHugging Face16sinhala-nlp /pali-sinhalatext10K<n<100K1 likes51 downloads1y agoHugging Face17rahul77 /pali_processed_8969181image1K<n<10K0 likes48 downloads2y agoHugging Face18SumitAST /PaliGemma_CLEVR_Datasetimage1K<n<10K0 likes46 downloads2y agoHugging Face19rahul77 /pali_processed_7788image1K<n<10K0 likes43 downloads2y agoHugging Face20grimjim /PAlign-PAPI-personality_prompt.json-cleanedAdapted from "Personality Alignment of Large Language Models" by Minjun Zhu and Linyi Yang and Yue Zhang and the associated GitHub repository zhu-minjun/PAlign. The contents of said repo were declared public domain; in that spirit, this Alpaca-formatted file has also been released as public domain. texttext-classificationn<1K1 likes42 downloads2y agoHugging Face21rahul77 /pali_processed_8969image1K<n<10K0 likes41 downloads2y agoHugging Face22freococo /myanmar-english-pali-dictionary Myanmar–English–Pali Dictionary Dataset Summary This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein). It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary. The dataset is intended for research and educational purposes, including but not limited to: Natural Language Processing (NLP) Machine Translation (MT) Lexicography Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.texttranslation10K<n<100K1 likes41 downloads8mo agoHugging Face23mimimimi2002 /license-detection-paligemmaimage1K<n<10K0 likes41 downloads2mo agoHugging Face24buddhist-nlp /pali-english-devanagari Dataset Card for "pali-english-devanagari" More Information needed text100K<n<1M0 likes40 downloads3y agoHugging Face25agyaatcoder /plant-doc-paligemmaimage1K<n<10K0 likes40 downloads2y agoHugging Face26alakxender /dhivehi-layout-syn-b1-paligemma Synthetic Dhivehi Document Layout Analysis Dataset Overview This dataset contains synthetic document layouts annotated with bounding boxes and labels for various sections in the Dhivehi language. It is designed for training models on document layout analysis and Optical Character Recognition (OCR) tasks. The dataset simulates real-world documents in Dhivehi and can be used for layout-aware OCR to recognize text and understand the structure of Dhivehi documents.… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-layout-syn-b1-paligemma.image1K<n<10K0 likes40 downloads1y agoHugging Face27hrhraj /eval_paligemma_so100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 388, "total_tasks": 1, "total_videos": 3, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hrhraj/eval_paligemma_so100.tabularroboticsn<1K0 likes38 downloads1y agoHugging Face28DatarrX /pali-myanmar-dictionary-corpus Pali-Myanmar Dictionary Corpus (Instruction-Ready) Dataset Summary The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning. Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.texttranslation100K<n<1M6 likes38 downloads5mo agoHugging Face29sagarnildass /brain-tumor-object-detection-paligemmaimage1K<n<10K0 likes35 downloads1y agoHugging Face30fosters /uladzimir-karatkevich-dzikae-paliavanne-karalia-stakha-aleg-garbuz-output_original Дзікае паляванне Кароля Стаха — арыгінальнае аўдыё Аўтар / Author: Уладзімір КараткевічМова / Language: Беларуская (Belarusian) Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці. Частка калекцыі Ministerskija — корпус беларускіх аўдыёкніг. Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя): uladzimir-karatkevich-dzikae-paliavanne-karalia-stakha-aleg-garbuz-output Доўгасць аўдыё 7h37m Радкоў у датасеце 2,585 Структура… See the full description on the dataset page: https://huggingface.co/datasets/fosters/uladzimir-karatkevich-dzikae-paliavanne-karalia-stakha-aleg-garbuz-output_original.audioautomatic-speech-recognition1K<n<10K0 likes35 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.