CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hozifa1 /islamic-scientific-heritage 📜 Islamic Scientific Heritage (التراث العلمي الإسلامي) The Golden Digital Corpus of Classical Islamic Sciences (8th–18th Centuries CE / 2nd–12th Centuries AH) Repository: hozifa1/islamic-scientific-heritageCurators: Hozifa & The Antigravity Agentic AI TeamDataset Scale: 593 Historical Volumes (PDFs) | 2,694 Structured Metadata Files | 23.43 GBArchival Milestone: 51 Complete Trios (51 ثلاثية علمية مكتملة 3/3) across all 6 core scientific disciplines!… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/islamic-scientific-heritage.text-retrieval1 likes1.5k downloads6d agoHugging Face02hmar-heritage-org /corpus-archivegated corpus-archive [!WARNING] Experimental Dataset Architecture: The repository structure, metadata tiers, category taxonomies, and catalog indexing formats are currently under active design and evaluation. All specifications, metadata keys, and JSON schemas detailed below represent representational examples and intended targets. This repository serves as a structured digital textual archive preserving Hmar literature, historical accounts, school textbooks, dictionaries, parallel… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/corpus-archive.imagetext-classificationn<1K4 likes1.1k downloads7d agoHugging Face03common-pile /biodiversity_heritage_library_filtered Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 15 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.texttext-generation10M<n<100M2 likes843 downloads1y agoHugging Face04common-pile /biodiversity_heritage_library Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 42 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library.texttext-generation10M<n<100M2 likes619 downloads1y agoHugging Face05Pclanglais /heritage0 likes560 downloads1y agoHugging Face06hmar-heritage-org /zo-biblegated zo-bible A sentence-aligned parallel Bible corpus covering 8 closely related Zo speech varieties and English across 10 translation versions (30,974 canonical verses). Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Languages: Hmar (hmr), Mizo (lus), Paite (pck), Vaiphei (vap), Thadou (tcz), Gangte (gnb), Zou (zom), English (eng) Family: Zo Languages Volume: 30,974 verse anchors across 66 canonical books (10 translation editions)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/zo-bible.tabulartranslation10K<n<100K4 likes495 downloads8d agoHugging Face07hmar-heritage-org /wordlistgated wordlist A structured digital lexicon containing 43,509 Hmar words, phrases, definitions, and translations compiled from five lexicographical sources. Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en) Family: Zo Languages Volume: 43,509 lexical entries across 5 dictionary files Format: JSON (data/dictionary-001.json to data/dictionary-005.json) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/wordlist.translation5 likes398 downloads8d agoHugging Face08yale-cultural-heritage /yuag-numismatics Yale University Art Gallery Numismatic Collection This is a collection of over 53,000 coins held at the Yale University Art Gallery. The data were downloaded from Yale's Lux Collection Discovery. Lux let's users find and connect with the cultural heritage collections across Yale's museums, archives, and libraries in new ways and all in one place. The Numismatic Collection consist of over 70,000 objects. We filtered this dataset to only examples that had a single image. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yale-cultural-heritage/yuag-numismatics.image10K<n<100K2 likes387 downloads1y agoHugging Face09hmar-heritage-org /unigramsgated unigrams A frequency-weighted lexical dataset containing 98,906 Hmar unigrams and active loanwords with occurrence counts compiled directly from the Foundation's verified Hmar corpus (over 4.14 million words). Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241) Family: Zo Languages Volume: 98,906 unigram tokens (compiled across 4,146,783 words) Format: JSONL (data/train.jsonl)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/unigrams.texttoken-classification10K<n<100K2 likes233 downloads4d agoHugging Face10hmar-heritage-org /sentencesgated sentences A multi-register sentence corpus for the Hmar language (hmr, ISO 639-3), containing 260,492 train sentences and 5,316 evaluation sentences (~4.15 million words). Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241) Family: Zo Languages Volume: 260,492 train sentences | 5,316 evaluation sentences (265,808 total, ~4.15M words) Validation: 100% verified via hmaraniam… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/sentences.textfill-mask100K<n<1M2 likes201 downloads4d agoHugging Face11JZSG /gallica_heritagetabular1M<n<10M1 likes177 downloads11mo agoHugging Face12hmar-heritage-org /numeral-wordsgated numeral-words A 3-way parallel digital dataset containing 999,999 spelled-out Hmar & English number words mapped in sequential numerical order (1 to 999,999). Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en) Family: Zo Languages Volume: 999,999 parallel rows (1 to 999,999) Format: Compressed JSONL (data/train-*.jsonl.gz) License: Apache-2.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/numeral-words.texttranslation100K<n<1M0 likes132 downloads8d agoHugging Face13hmar-heritage-org /hmingtluongated hmingtluon A procedural synthetic anthroponyms dataset containing 10,193,568 (10.19 million) unique Hmar full names, designed for Named Entity Recognition (NER), entity resolution, synthetic data augmentation, and anthroponymic research. Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241) Family: Zo Languages Volume: 10,193,568 unique records across 3 splits (Train / Validation /… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/hmingtluon.texttoken-classification10M<n<100M0 likes123 downloads8d agoHugging Face14letrinhan /vn-provinces-national-cultural-heritage Vietnam national cultural heritage sites by locality (2023) Vietnam count of national-level cultural heritage sites by province and region for 2023 only. Salvaged from a broken NSO Excel-XML export (V14.25) whose dimension axes were mislabeled. one verified total per locality. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-national-cultural-heritage.tabularn<1K0 likes82 downloads2d agoHugging Face15kaumudi-ai /kerala-cultural-heritage Kerala Cultural Heritage This CC-BY-4.0 release contains 7 JSONL catalogue and schema records for cultural knowledge graphs, glossaries, oral-history, event-archive, and cultural-Q&A collections. It is a template release: no third-party media, interviews, or archive reproductions are included. Each future record must include its source, rights, consent status, and the licence that applies to that source. 1 likes81 downloads3mo agoHugging Face16SINAI /ALIA-es-cultural-heritage-synthetic-instructions Dataset Introduction The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains: 748,480 instances 629,682,398 tokens 25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.texttext-generation100K<n<1M0 likes78 downloads4mo agoHugging Face17SINAI /ALIA-es-cultural-heritage Dataset Introduction The ALIA Spanish Cultural and Heritage Corpus is a strategic open data infrastructure designed to support research and innovation in digital humanities, cultural analytics, and Spanish-language NLP. It consolidates heterogeneous official and academic repositories into a single curated dataset, enabling broad and structured access to cultural heritage documentation from Spain. With 236,314 instances, 939,315,404 tokens and 100 source datasets, it provides a… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage.texttext-generation100K<n<1M0 likes72 downloads3mo agoHugging Face18SINAI /ALIA-es-cultural-heritage-pairs Dataset Introduction The ALIA Spanish Cultural and Heritage Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen-style prompting workflow integrated in the ALIA encoders pipeline. It preserves provenance to the original document and passage while exposing controls such as question type and difficulty (ranging from high_school… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-pairs.textquestion-answering100K<n<1M0 likes70 downloads3mo agoHugging Face19yale-cultural-heritage /shadow-puppet-outputtext1K<n<10K0 likes62 downloads1y agoHugging Face20SINAI /ALIA-heritage-parallel-translation Dataset Card for ALIA Cultural Heritage Parallel Translation Corpus (ES→EN) This corpus contains 683,919 parallel chunks and 288,955 full documents (Spanish–English) from the Cultural Heritage domain of the ALIA project. It covers texts related to Cultural Heritage of Spain, automatically translated from Spanish into English using the Qwen3-14B large language model. The dataset is available in two configurations: chunked (683,919 individual translation units) and merged (288,955… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-heritage-parallel-translation.texttranslation100K<n<1M0 likes59 downloads5mo agoHugging Face21GamesMais18 /HeritageDados0 likes58 downloads29d agoHugging Face22SINAI /ALIA-es-cultural-heritage-triplets Dataset Introduction The dataset ALIA Spanish Cultural and Heritage Hard Negatives Corpus contains hard negatives for dense retrieval training generated from <query, passage> pairs contained in SINAI/ALIA-es-cultural-heritage-pairs.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish cultural heritage language. Hard negatives are passages that are semantically similar to a query but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-triplets.texttext-generation1M<n<10M0 likes57 downloads4mo agoHugging Face23biglam /cultural_heritage_metadata_accuracy Dataset Card for Annotated dataset to assess the accuracy of the textual description of cultural heritage records Dataset Summary The dataset contains more than 100K textual descriptions of cultural items from Cultura Italia, the Italian National Cultural aggregator. Each of the description is labeled either HIGH or LOW quality, according its adherence to the standard cataloguing guidelines provided by Istituto Centrale per il Catalogo e la Documentazione (ICCD). More… See the full description on the dataset page: https://huggingface.co/datasets/biglam/cultural_heritage_metadata_accuracy.texttext-classification100K<n<1M6 likes49 downloads3y agoHugging Face24PoSTMEDIA /heritage-ko-bench heritage-ko-bench Korean cultural-heritage knowledge benchmark v1.3 — 6 four-choice MCQ subsets (1,317 items) + 1 short-answer subset (270 items) = 1,587 items (test 1,552 · dev 35 for few-shot). ⚠️ Evaluation only. Do not train on this data. Why this benchmark Decontamination by design: items are minted exclusively from 474 documents held out before the companion training suites were generated — document-level separation, empirically verified. Triple full-pass… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/heritage-ko-bench.textquestion-answering1K<n<10K0 likes49 downloads16d agoHugging Face25hmar-heritage-org /culture-dump culture-dump A digital archive repository for raw community web dumps, public news archives, oral literature recordings, and cultural assets for the Hmar language. Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241) Family: Zo Languages Scope: Web publisher HTML archives, blog dumps, oral traditions, and cultural records License: CC BY 4.0 Usage from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/culture-dump.0 likes45 downloads8d agoHugging Face26yale-cultural-heritage /lux-typed-docs Yale LUX dots.ocr layout/OCR outputs Structured OCR/layout output from the rednote-hilab/dots.ocr model for Yale LUX document images. Each row carries the OCR output, the source image_url, the canvas_index, and a pointer to its manifest. Rows: 626,586. Built from the Yale LUX manifest processing database. image100K<n<1M0 likes37 downloads3mo agoHugging Face27ArturMansur /boombap-heritage-datasettextn<1K0 likes35 downloads5d agoHugging Face28electricsheepafrica /africa-unsdg-direct-economic-loss-to-cultural-heritage-damaged-or-de-vc-dsr-chln Africa Unsdg Direct Economic Loss to Cultural Heritage Damaged or De Vc Dsr Chln | Africa (Electric Sheep Africa metadata inventory) Size category: n<1K - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-unsdg-direct-economic-loss-to-cultural-heritage-damaged-or-de-vc-dsr-chln.tabulartabular-classificationn<1K0 likes33 downloads1mo agoHugging Face29squanchyzx /turkiye-cultural-heritage-atlas Türkiye Cultural Heritage Atlas — Open Data (Kahve Tabela) Two open datasets compiled from public sources by Kahve Tabela, an independent atlas of Türkiye's cultural heritage: Snapshot: 24 July 2026. These files are a fixed, citable extract — that is what the DOI is for. The live atlas keeps moving as records are cleaned, merged and added, so its counts drift away from the ones here; for current figures see kahvetabela.com. heritage-lite-32483-ae184b49.csv — 32,483 registered… See the full description on the dataset page: https://huggingface.co/datasets/squanchyzx/turkiye-cultural-heritage-atlas.geospatialother10K<n<100K0 likes32 downloads2mo agoHugging Face30kebisi /Nyagbo-Tutrugbu-Language-Heritage0 likes30 downloads19d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.