CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nomeda-lab /fattah-golden-superset Fattah Golden Fattah Golden is a large-scale, model-agnostic supervised fine-tuning (SFT) superset built by Nomeda Labs to train the Fattah family of coding and agentic coding models. The dataset is designed as a labeled superset with no baked-in training ratios. This means the stored dataset is the complete cleaned and annotated corpus. Researchers and practitioners choose their own mixture at training time by filtering on the boolean capability columns. Stats… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/fattah-golden-superset.tabulartext-generation1M<n<10M1 likes521 downloads4mo agoHugging Face02fatihburakkaragoz /old-nogay-turkish-ocr-corpus Old Nogay Turkish OCR Corpus This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis. We are proud to release this as a contribution to Turkish NLP, Turkic-language NLP, historical NLP, and Turkish studies. The goal is not to pretend that a small OCR corpus is a polished benchmark. The goal is more… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/old-nogay-turkish-ocr-corpus.tabulartext-generationn<1K0 likes75 downloads5mo agoHugging Face03sermonindex /early-church-fathers Early Church Fathers — Scripture Citation Index 68,240 passages from 349 Church Fathers, each keyed to the Bible verse it comments on. Drawn from 20,253 distinct works and covering all 66 books. This is a patristic catena in machine-readable form: given a verse, it returns what the Fathers said about it. Nothing comparable exists as an open dataset — the underlying translations are freely available, but the verse-level alignment is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.tabulartext-retrieval10K<n<100K0 likes69 downloads16d agoHugging Face04SahmBenchmark /fatwa-qa-evaluation Fatwa QA Evaluation Dataset Dataset Description This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers. Dataset Statistics Total Samples: 2,000 Average Question Length: 243.9 characters Average Answer Length: 492.3 characters Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.tabularquestion-answering1K<n<10K0 likes43 downloads10mo agoHugging Face05fatihburakkaragoz /anadolu-ocr-corpus Anadolu OCR Corpus Anadolu OCR Corpus is an OpenCR export of OCR text and document metadata for 52 historical Ottoman Turkish, Turkish, and Arabic-containing PDF sources. The dataset is provided in two Hugging Face configs: pages: one row per source page, including page-level OCR text, metadata, validation status, script direction, language detection, hashes, and split labels. documents: one row per source document, including document-level concatenated text, markdown, aggregate… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/anadolu-ocr-corpus.tabulartext-generation10K<n<100K3 likes27 downloads5mo agoHugging Face06SahmBenchmark /fatwa-training_standardized_new Fatwa Training Dataset (Standardized) Dataset Description This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs in a standardized conversation format for training Arabic language models. Each original sample has been augmented with 3 different prompt templates to increase training diversity. Dataset Statistics Total Samples: 9,953 Unique Fatwas: 6,212 Prompt Variations: 3 per fatwa Average Question Length: 230.0… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-training_standardized_new.tabularquestion-answering1K<n<10K0 likes15 downloads9mo agoHugging Face07fatihburakkaragoz /evliya-celebi-seyahatname-ocr Evliya Celebi Seyahatname OCR Corpus OCR-derived text from seven volumes of Evliya Celebi's Seyahatname, packaged for corpus exploration, language modeling, OCR-quality analysis, and historical Ottoman Turkish / Turkish NLP work. Configs pages: one row per OCR page, with page numbers and OCR status. documents: one row per available volume, with page text concatenated. Coverage Available books: 1, 3, 4, 6, 7, 9, 10. Missing from the 1-10 sequence: 2, 5, 8.… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/evliya-celebi-seyahatname-ocr.tabulartext-generation1K<n<10K1 likes14 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.