CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BEE-spoke-data /govdocs1-pdf-source govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.documentimage-text-to-text100K<n<1M6 likes4.4k downloads9mo agoHugging Face02open-source-metrics /tokenizers-dependents tokenizers metrics This dataset contains metrics about the huggingface/tokenizers package. Number of repositories in the dataset: 11460 Number of packages in the dataset: 124 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 14 packages that have more than 1000 stars. There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tabularn<1K0 likes2.1k downloads2y agoHugging Face03astro-legacy-archive /gaia-dr3-source Gaia DR3 Source This dataset mirrors the complete ESA Gaia Data Release 3 gaia_source bulk-download table. It contains one record for every published Gaia source and all 152 columns served by ESA, including source identifiers, astrometry, photometry, observing statistics, quality fields, classifications, and astrophysical parameters. Each of ESA's 3,386 compressed ECSV shards is retained as one dataset configuration. The configuration name is the source filename stem verbatim.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-source.tabular1B<n<10B1 likes1.9k downloads8h agoHugging Face04wytro /Know-Your-Sourcestabulartext-generation10M<n<100M0 likes1.2k downloads1mo agoHugging Face05panzs19 /ElephantBench-Source ElephantBench Source Corpus This dataset contains the low-quality partition ($D_{\mathrm{low}}$) used to construct ElephantBench. The released benchmark is available separately at panzs19/ElephantBench. $D_{\mathrm{low}}$ is derived from cx-cmu/repro-organic-data-72B using the RePro fastText quality score. It contains 47,850,862 English web documents in 600 JSONL.zstd shards (approximately 65 GiB compressed). Each record retains the source text, URL, quality score, and original… See the full description on the dataset page: https://huggingface.co/datasets/panzs19/ElephantBench-Source.tabular10M<n<100M0 likes1.2k downloads23d agoHugging Face06open-source-metrics /transformers-dependents transformers metrics This dataset contains metrics about the huggingface/transformers package. Number of repositories in the dataset: 27067 Number of packages in the dataset: 823 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 65 packages that have more than 1000 stars. There are 140… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/transformers-dependents.tabular10K<n<100K2 likes1.2k downloads2y agoHugging Face07cloudwalk-research /psi0-g1-sneaker-205ep-v2-source Psi0 G1 Sneaker-in-Box — 205 episodes (v2 canonical source) ⚠️ Do not use this dataset directly for training. This is the canonical immutable union of the v1 and v2 collections, kept as a source of truth for reproducibility. For v2 fine-tuning use psi0-g1-sneaker-199ep-v2; for held-out evaluation use psi0-g1-sneaker-6ep-v2-eval. Together these two derivatives reconstruct this canonical dataset exactly: 199 + 6 = 205. 205 teleoperated episodes of a Unitree G1 humanoid (with Inspire… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/psi0-g1-sneaker-205ep-v2-source.tabularrobotics10K<n<100K0 likes730 downloads4mo agoHugging Face08open-source-metrics /gradio-dependents Dataset Card for "gradio-dependents" More Information needed tabular1K<n<10K0 likes722 downloads2y agoHugging Face09alwaysgood /financial-english-source-corpus-gemma4-e2b-1280tabular1M<n<10M0 likes552 downloads2mo agoHugging Face10alwaysgood /financial-english-source-corpus Financial English Source Corpus This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. This version preserves the final pre-split source rows. Derived 1280-token split versions are available separately: financial-english-source-corpus-qwen35-1280 financial-english-source-corpus-gemma4-e2b-1280 Dataset Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.tabulartext-generation1M<n<10M0 likes547 downloads12d agoHugging Face11open-source-metrics /datasets-dependents datasets metrics This dataset contains metrics about the huggingface/datasets package. Number of repositories in the dataset: 4997 Number of packages in the dataset: 215 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 22 packages that have more than 1000 stars. There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.tabular10K<n<100K0 likes537 downloads2y agoHugging Face12NuBerea /source-analysisgated NuBerea Source Analysis Source-critical analysis of the Hebrew Bible, Septuagint, New Testament, Vulgate, and Second Temple literature. The dataset carries machine-generated source and tradition annotations at the verse level — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, the pathway of Old Testament traditions into New Testament citation) expressed as structured data — together with semantic-domain… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-analysis.tabularfeature-extraction100K<n<1M0 likes516 downloads3d agoHugging Face13open-source-metrics /evaluate-dependents evaluate metrics This dataset contains metrics about the huggingface/evaluate package. Number of repositories in the dataset: 106 Number of packages in the dataset: 3 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 1 packages that have more than 1000 stars. There are 2 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/evaluate-dependents.tabular1K<n<10K0 likes504 downloads2y agoHugging Face14open-source-metrics /accelerate-dependents accelerate metrics This dataset contains metrics about the huggingface/accelerate package. Number of repositories in the dataset: 727 Number of packages in the dataset: 37 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 10 packages that have more than 1000 stars. There are 16… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/accelerate-dependents.tabular1K<n<10K1 likes462 downloads2y agoHugging Face15NuBerea /secondary-sourcesgated NuBerea/secondary-sources Second Temple Jewish secondary sources in Greek: the complete extant Greek corpora of Flavius Josephus (Jewish Antiquities, Jewish War, Vita, Contra Apionem) and Philo of Alexandria (all 31 works), segmented for scholarly text-retrieval and lexical-semantic study. These two first-century authors are the principal non-biblical Jewish witnesses to the Second Temple period and its milieu, and this repository serves as the Second Temple companion corpus to… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/secondary-sources.tabulartext-retrieval1M<n<10M0 likes432 downloads10d agoHugging Face16NuBerea /source-classificationsgated NuBerea Source Gold Set Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists. This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.tabulartext-generation10K<n<100K0 likes430 downloads2mo agoHugging Face17open-source-metrics /diffusers-dependents diffusers metrics This dataset contains metrics about the huggingface/diffusers package. Number of repositories in the dataset: 160 Number of packages in the dataset: 2 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 0 packages that have more than 1000 stars. There are 3 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/diffusers-dependents.tabular1K<n<10K1 likes423 downloads2y agoHugging Face18open-source-metrics /optimum-dependents optimum metrics This dataset contains metrics about the huggingface/optimum package. Number of repositories in the dataset: 19 Number of packages in the dataset: 6 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 0 packages that have more than 1000 stars. There are 0 repositories that… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/optimum-dependents.tabularn<1K1 likes403 downloads2y agoHugging Face19lapa-llm /classifier_source Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.tabulartext-generation1M<n<10M0 likes356 downloads10mo agoHugging Face20balrampandey /qmmit-open-source-agent-commit-index Repository-Level Measurement of Self-Declared Coding-Agent Commit Signatures Dataset release: 2026-09-18-v3.0Schema: 3.0.0 Release stamp: dataset 2026-09-18-v3.0 · ruleset sha256:b2e8889c66f72c18a61839f0bf1a9f77b5481ba2def044dba33c797b2f2bdcae · scanned 2026-09-16 Abstract This dataset contains 2000 repository-level observations from public Git repositories. Each observation estimates a lower bound on the proportion of non-merge, non-infrastructure-bot commits… See the full description on the dataset page: https://huggingface.co/datasets/balrampandey/qmmit-open-source-agent-commit-index.tabulartabular-classification1K<n<10K0 likes269 downloads4d agoHugging Face21alwaysgood /financial-english-source-corpus-qwen35-1280 Financial English Source Corpus Qwen35 1280 This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. The uploaded Parquet files are already prepared with the 1280-token source split used by the downstream training pipeline. This split version is derived from the pre-split Financial English Source Corpus by applying sentence-boundary splitting with the qwen3.5 tokenizer.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus-qwen35-1280.tabulartext-generation1M<n<10M0 likes265 downloads3mo agoHugging Face22jescy525 /archon-comm-v1-sources-light archon-comm-v1-sources-light ARCHON v1 light sources -- 7 flow-level IDS datasets (CIC-IDS, UNSW-NB15 small, SIMARGL, network anomaly, witfoo SOC, etc.) -- ~68 GB total. For network flow training without raw packets. Use jescy525/archon-comm-v1-sources-packets for packet-level data. Disclaimer This bundle aggregates publicly available security research datasets for defensive purposes only. See SOURCES.md for per-source license attribution. Contents… See the full description on the dataset page: https://huggingface.co/datasets/jescy525/archon-comm-v1-sources-light.tabular1M<n<10M0 likes243 downloads4mo agoHugging Face23Mithilss /neurips-2025-arxiv-latex-sources NeurIPS 2025 arXiv LaTeX Source Files This dataset contains file-level Parquet rows built from extracted raw arXiv source packages for papers mapped from the NeurIPS 2025 proceedings to arXiv records. Each row is one file from one arXiv source package. Use arxiv_id to group files back into papers. Columns arxiv_id: arXiv identifier for the source package. title: paper title from the mapping CSV. source_url: arXiv e-print source URL. paper_index, paper_status… See the full description on the dataset page: https://huggingface.co/datasets/Mithilss/neurips-2025-arxiv-latex-sources.tabular100K<n<1M0 likes205 downloads4mo agoHugging Face24arqa39 /proofwriter-source ProofWriter (The Source) An unmodified copy of AI2's ProofWriter dataset (release V2020.12.3), re-hosted as datasets configs for convenient loading. The records are faithful to the upstream release — the id-keyed JSON is preserved as-is; typing and reasoning-graph extraction happen in later stages. Each config is a {world}-depth-{n} shelf of the synthetic core (OWA/CWA × depths 0/1/2/3/5), split train/dev/test (dev kept as the corpus names it). Source:… See the full description on the dataset page: https://huggingface.co/datasets/arqa39/proofwriter-source.tabular100K<n<1M0 likes195 downloads1mo agoHugging Face25brennercruvinel /news_sources_brazil news_sources_brazil 17,892 news outlets, one per row, keyed to the IBGE municipality. 16,298 come from atlas da notícia, the survey of local journalism that projor and volt data lab have run since 2017. the other 1,594 are the portals, wire agencies, fact-checkers and international references that truw, my news verification project, was already tracking. the municipality table ships alongside, with 2022 population, HDI, outlet count and the news desert flag, so the whole thing… See the full description on the dataset page: https://huggingface.co/datasets/brennercruvinel/news_sources_brazil.tabular10K<n<100K0 likes168 downloads12d agoHugging Face26bluelightai-dev /common-corpus-sample-open-sourcetabular1M<n<10M0 likes162 downloads11mo agoHugging Face27haowu89 /open_parallel_think_code_source open_parallel_think_code_source A large-scale code reasoning distillation dataset with 320,000 solution trajectories generated by 4 state-of-the-art thinking models across 10,000 unique coding problems. Source / raw pool. This is the per-trajectory dataset. The packed parallel-thinking datasets derived from it are haowu89/open_parallel_think_code_full (full reasoning + solution) and haowu89/open_parallel_think_code_cot (solution only). Each trajectory's metadata carries… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_code_source.tabular100K<n<1M0 likes159 downloads4mo agoHugging Face28MindBench /search-source-audit Sources of Truth — AI Search Citations for Mental Health Queries Which external sources do consumer AI search products actually cite when people ask about mental health? This dataset is the annotated citation corpus behind "Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries." Twenty English mental health questions were put to three free consumer AI search products (ChatGPT, Perplexity, and Google AI Overview) under two… See the full description on the dataset page: https://huggingface.co/datasets/MindBench/search-source-audit.tabular10K<n<100K0 likes132 downloads20d agoHugging Face29rlhf-and-friends /proofwriter-source ProofWriter (The Source) An unmodified copy of AI2's ProofWriter dataset (release V2020.12.3), re-hosted as datasets configs for convenient loading. The records are faithful to the upstream release — the id-keyed JSON is preserved as-is; typing and reasoning-graph extraction happen in later stages. Each config is a {world}-depth-{n} shelf of the synthetic core (OWA/CWA × depths 0/1/2/3/5), split train/dev/test (dev kept as the corpus names it). Source:… See the full description on the dataset page: https://huggingface.co/datasets/rlhf-and-friends/proofwriter-source.tabular100K<n<1M0 likes125 downloads1mo agoHugging Face30Long27 /so101_mixed_3_sourcesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Long27/so101_mixed_3_sources.tabularrobotics100K<n<1M0 likes114 downloads25d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.