CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BEE-spoke-data /govdocs1-pdf-source govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.documentimage-text-to-text100K<n<1M6 likes4.5k downloads9mo agoHugging Face02open-source-metrics /tokenizers-dependents tokenizers metrics This dataset contains metrics about the huggingface/tokenizers package. Number of repositories in the dataset: 11460 Number of packages in the dataset: 124 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 14 packages that have more than 1000 stars. There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tabularn<1K0 likes2.1k downloads2y agoHugging Face03wytro /Know-Your-Sourcestabulartext-generation10M<n<100M0 likes1.4k downloads1mo agoHugging Face04open-source-metrics /transformers-dependents transformers metrics This dataset contains metrics about the huggingface/transformers package. Number of repositories in the dataset: 27067 Number of packages in the dataset: 823 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 65 packages that have more than 1000 stars. There are 140… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/transformers-dependents.tabular10K<n<100K2 likes1.2k downloads2y agoHugging Face05yaofu /slimpajama-per-source-length-upsampletext10K<n<100K18 likes836 downloads3y agoHugging Face06open-source-metrics /gradio-dependents Dataset Card for "gradio-dependents" More Information needed tabular1K<n<10K0 likes733 downloads2y agoHugging Face07xfxcwynlc /prolong-64k-hf-decoded-text-sourcetext10M<n<100M0 likes692 downloads1y agoHugging Face08FNLP-GUI-AGENT-GROUP /DATA_SOURCEtext100K<n<1M0 likes680 downloads11mo agoHugging Face09mingyy /source_filter Dataset Card for "source_filter" More Information needed image10K<n<100K0 likes679 downloads3y agoHugging Face10nomic-ai /colpali_train_set_split_by_sourceimage100K<n<1M2 likes569 downloads2y agoHugging Face11open-source-metrics /datasets-dependents datasets metrics This dataset contains metrics about the huggingface/datasets package. Number of repositories in the dataset: 4997 Number of packages in the dataset: 215 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 22 packages that have more than 1000 stars. There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.tabular10K<n<100K0 likes554 downloads2y agoHugging Face12open-source-metrics /evaluate-dependents evaluate metrics This dataset contains metrics about the huggingface/evaluate package. Number of repositories in the dataset: 106 Number of packages in the dataset: 3 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 1 packages that have more than 1000 stars. There are 2 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/evaluate-dependents.tabular1K<n<10K0 likes523 downloads2y agoHugging Face13alwaysgood /financial-english-source-corpus Financial English Source Corpus This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. This version preserves the final pre-split source rows. Derived 1280-token split versions are available separately: financial-english-source-corpus-qwen35-1280 financial-english-source-corpus-gemma4-e2b-1280 Dataset Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.tabulartext-generation1M<n<10M0 likes520 downloads15d agoHugging Face14alwaysgood /financial-english-source-corpus-gemma4-e2b-1280tabular1M<n<10M0 likes517 downloads3mo agoHugging Face15NuBerea /source-analysisgated NuBerea Source Analysis Source-critical analysis of the Hebrew Bible, Septuagint, New Testament, Vulgate, and Second Temple literature. The dataset carries machine-generated source and tradition annotations at the verse level — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, the pathway of Old Testament traditions into New Testament citation) expressed as structured data — together with semantic-domain… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-analysis.tabularfeature-extraction100K<n<1M0 likes512 downloads5d agoHugging Face16open-source-metrics /accelerate-dependents accelerate metrics This dataset contains metrics about the huggingface/accelerate package. Number of repositories in the dataset: 727 Number of packages in the dataset: 37 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 10 packages that have more than 1000 stars. There are 16… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/accelerate-dependents.tabular1K<n<10K1 likes446 downloads2y agoHugging Face17Xin-Rui /BudgetThinker-Data-Sourcetext100K<n<1M1 likes437 downloads1y agoHugging Face18NuBerea /secondary-sourcesgated NuBerea/secondary-sources Second Temple Jewish secondary sources in Greek: the complete extant Greek corpora of Flavius Josephus (Jewish Antiquities, Jewish War, Vita, Contra Apionem) and Philo of Alexandria (all 31 works), segmented for scholarly text-retrieval and lexical-semantic study. These two first-century authors are the principal non-biblical Jewish witnesses to the Second Temple period and its milieu, and this repository serves as the Second Temple companion corpus to… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/secondary-sources.tabulartext-retrieval1M<n<10M0 likes430 downloads12d agoHugging Face19NuBerea /source-classificationsgated NuBerea Source Gold Set Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists. This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.tabulartext-generation10K<n<100K0 likes428 downloads2mo agoHugging Face20DopeorNope /Pandora_source_no_mathtext1M<n<10M0 likes413 downloads2y agoHugging Face21open-source-metrics /optimum-dependents optimum metrics This dataset contains metrics about the huggingface/optimum package. Number of repositories in the dataset: 19 Number of packages in the dataset: 6 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 0 packages that have more than 1000 stars. There are 0 repositories that… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/optimum-dependents.tabularn<1K1 likes412 downloads2y agoHugging Face22open-source-metrics /diffusers-dependents diffusers metrics This dataset contains metrics about the huggingface/diffusers package. Number of repositories in the dataset: 160 Number of packages in the dataset: 2 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 0 packages that have more than 1000 stars. There are 3 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/diffusers-dependents.tabular1K<n<10K1 likes406 downloads2y agoHugging Face23lapa-llm /classifier_source Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.tabulartext-generation1M<n<10M0 likes365 downloads11mo agoHugging Face24jiosephlee /starling-transfer-shared-eval-same-species-v2-source-valuetext10M<n<100M0 likes345 downloads3mo agoHugging Face25nomic-ai /VisRAG-Ret-Train-In-domain-data-by-source-hn-mine-corpusimage100K<n<1M0 likes335 downloads1y agoHugging Face26open-source-metrics /pip Dataset Card for "pip" More Information needed text10K<n<100K0 likes325 downloads2y agoHugging Face27jiosephlee /starling-transfer-shared-eval-no-constraints-no-source-value-train-2p5pct-stratifiedtext10M<n<100M0 likes320 downloads3mo agoHugging Face28AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes314 downloads2mo agoHugging Face29jiosephlee /starling-transfer-shared-eval-no-constraints-source-value-train-2p5pct-stratifiedtext10M<n<100M0 likes313 downloads3mo agoHugging Face30AmirhoseinGH /mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 32B Instruct hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k.imagequestion-answering10K<n<100K0 likes289 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.