CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BEE-spoke-data /govdocs1-pdf-source govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.documentimage-text-to-text100K<n<1M6 likes4.5k downloads9mo agoHugging Face02open-source-metrics /tokenizers-dependents tokenizers metrics This dataset contains metrics about the huggingface/tokenizers package. Number of repositories in the dataset: 11460 Number of packages in the dataset: 124 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 14 packages that have more than 1000 stars. There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tabularn<1K0 likes2.1k downloads2y agoHugging Face03wytro /Know-Your-Sourcestabulartext-generation10M<n<100M0 likes1.2k downloads1mo agoHugging Face04open-source-metrics /transformers-dependents transformers metrics This dataset contains metrics about the huggingface/transformers package. Number of repositories in the dataset: 27067 Number of packages in the dataset: 823 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 65 packages that have more than 1000 stars. There are 140… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/transformers-dependents.tabular10K<n<100K2 likes1.2k downloads2y agoHugging Face05open-source-metrics /gradio-dependents Dataset Card for "gradio-dependents" More Information needed tabular1K<n<10K0 likes733 downloads2y agoHugging Face06cloudwalk-research /psi0-g1-sneaker-205ep-v2-source Psi0 G1 Sneaker-in-Box — 205 episodes (v2 canonical source) ⚠️ Do not use this dataset directly for training. This is the canonical immutable union of the v1 and v2 collections, kept as a source of truth for reproducibility. For v2 fine-tuning use psi0-g1-sneaker-199ep-v2; for held-out evaluation use psi0-g1-sneaker-6ep-v2-eval. Together these two derivatives reconstruct this canonical dataset exactly: 199 + 6 = 205. 205 teleoperated episodes of a Unitree G1 humanoid (with Inspire… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/psi0-g1-sneaker-205ep-v2-source.tabularrobotics10K<n<100K0 likes729 downloads4mo agoHugging Face07open-source-metrics /datasets-dependents datasets metrics This dataset contains metrics about the huggingface/datasets package. Number of repositories in the dataset: 4997 Number of packages in the dataset: 215 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 22 packages that have more than 1000 stars. There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.tabular10K<n<100K0 likes554 downloads2y agoHugging Face08alwaysgood /financial-english-source-corpus-gemma4-e2b-1280tabular1M<n<10M0 likes546 downloads3mo agoHugging Face09alwaysgood /financial-english-source-corpus Financial English Source Corpus This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. This version preserves the final pre-split source rows. Derived 1280-token split versions are available separately: financial-english-source-corpus-qwen35-1280 financial-english-source-corpus-gemma4-e2b-1280 Dataset Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.tabulartext-generation1M<n<10M0 likes528 downloads14d agoHugging Face10open-source-metrics /evaluate-dependents evaluate metrics This dataset contains metrics about the huggingface/evaluate package. Number of repositories in the dataset: 106 Number of packages in the dataset: 3 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 1 packages that have more than 1000 stars. There are 2 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/evaluate-dependents.tabular1K<n<10K0 likes523 downloads2y agoHugging Face11NuBerea /source-analysisgated NuBerea Source Analysis Source-critical analysis of the Hebrew Bible, Septuagint, New Testament, Vulgate, and Second Temple literature. The dataset carries machine-generated source and tradition annotations at the verse level — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, the pathway of Old Testament traditions into New Testament citation) expressed as structured data — together with semantic-domain… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-analysis.tabularfeature-extraction100K<n<1M0 likes514 downloads4d agoHugging Face12open-source-metrics /accelerate-dependents accelerate metrics This dataset contains metrics about the huggingface/accelerate package. Number of repositories in the dataset: 727 Number of packages in the dataset: 37 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 10 packages that have more than 1000 stars. There are 16… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/accelerate-dependents.tabular1K<n<10K1 likes446 downloads2y agoHugging Face13NuBerea /secondary-sourcesgated NuBerea/secondary-sources Second Temple Jewish secondary sources in Greek: the complete extant Greek corpora of Flavius Josephus (Jewish Antiquities, Jewish War, Vita, Contra Apionem) and Philo of Alexandria (all 31 works), segmented for scholarly text-retrieval and lexical-semantic study. These two first-century authors are the principal non-biblical Jewish witnesses to the Second Temple period and its milieu, and this repository serves as the Second Temple companion corpus to… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/secondary-sources.tabulartext-retrieval1M<n<10M0 likes431 downloads12d agoHugging Face14NuBerea /source-classificationsgated NuBerea Source Gold Set Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists. This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.tabulartext-generation10K<n<100K0 likes429 downloads2mo agoHugging Face15open-source-metrics /optimum-dependents optimum metrics This dataset contains metrics about the huggingface/optimum package. Number of repositories in the dataset: 19 Number of packages in the dataset: 6 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 0 packages that have more than 1000 stars. There are 0 repositories that… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/optimum-dependents.tabularn<1K1 likes412 downloads2y agoHugging Face16open-source-metrics /diffusers-dependents diffusers metrics This dataset contains metrics about the huggingface/diffusers package. Number of repositories in the dataset: 160 Number of packages in the dataset: 2 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 0 packages that have more than 1000 stars. There are 3 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/diffusers-dependents.tabular1K<n<10K1 likes406 downloads2y agoHugging Face17lapa-llm /classifier_source Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.tabulartext-generation1M<n<10M0 likes359 downloads11mo agoHugging Face18balrampandey /qmmit-open-source-agent-commit-index Repository-Level Measurement of Self-Declared Coding-Agent Commit Signatures Dataset release: 2026-09-18-v3.0Schema: 3.0.0 Release stamp: dataset 2026-09-18-v3.0 · ruleset sha256:b2e8889c66f72c18a61839f0bf1a9f77b5481ba2def044dba33c797b2f2bdcae · scanned 2026-09-16 Abstract This dataset contains 2000 repository-level observations from public Git repositories. Each observation estimates a lower bound on the proportion of non-merge, non-infrastructure-bot commits… See the full description on the dataset page: https://huggingface.co/datasets/balrampandey/qmmit-open-source-agent-commit-index.tabulartabular-classification1K<n<10K0 likes271 downloads6d agoHugging Face19alwaysgood /financial-english-source-corpus-qwen35-1280 Financial English Source Corpus Qwen35 1280 This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. The uploaded Parquet files are already prepared with the 1280-token source split used by the downstream training pipeline. This split version is derived from the pre-split Financial English Source Corpus by applying sentence-boundary splitting with the qwen3.5 tokenizer.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus-qwen35-1280.tabulartext-generation1M<n<10M0 likes253 downloads3mo agoHugging Face20jescy525 /archon-comm-v1-sources-light archon-comm-v1-sources-light ARCHON v1 light sources -- 7 flow-level IDS datasets (CIC-IDS, UNSW-NB15 small, SIMARGL, network anomaly, witfoo SOC, etc.) -- ~68 GB total. For network flow training without raw packets. Use jescy525/archon-comm-v1-sources-packets for packet-level data. Disclaimer This bundle aggregates publicly available security research datasets for defensive purposes only. See SOURCES.md for per-source license attribution. Contents… See the full description on the dataset page: https://huggingface.co/datasets/jescy525/archon-comm-v1-sources-light.tabular1M<n<10M0 likes243 downloads4mo agoHugging Face21Mithilss /neurips-2025-arxiv-latex-sources NeurIPS 2025 arXiv LaTeX Source Files This dataset contains file-level Parquet rows built from extracted raw arXiv source packages for papers mapped from the NeurIPS 2025 proceedings to arXiv records. Each row is one file from one arXiv source package. Use arxiv_id to group files back into papers. Columns arxiv_id: arXiv identifier for the source package. title: paper title from the mapping CSV. source_url: arXiv e-print source URL. paper_index, paper_status… See the full description on the dataset page: https://huggingface.co/datasets/Mithilss/neurips-2025-arxiv-latex-sources.tabular100K<n<1M0 likes204 downloads4mo agoHugging Face22arqa39 /proofwriter-source ProofWriter (The Source) An unmodified copy of AI2's ProofWriter dataset (release V2020.12.3), re-hosted as datasets configs for convenient loading. The records are faithful to the upstream release — the id-keyed JSON is preserved as-is; typing and reasoning-graph extraction happen in later stages. Each config is a {world}-depth-{n} shelf of the synthetic core (OWA/CWA × depths 0/1/2/3/5), split train/dev/test (dev kept as the corpus names it). Source:… See the full description on the dataset page: https://huggingface.co/datasets/arqa39/proofwriter-source.tabular100K<n<1M0 likes184 downloads1mo agoHugging Face23brennercruvinel /news_sources_brazil news_sources_brazil 17,892 news outlets, one per row, keyed to the IBGE municipality. 16,298 come from atlas da notícia, the survey of local journalism that projor and volt data lab have run since 2017. the other 1,594 are the portals, wire agencies, fact-checkers and international references that truw, my news verification project, was already tracking. the municipality table ships alongside, with 2022 population, HDI, outlet count and the news desert flag, so the whole thing… See the full description on the dataset page: https://huggingface.co/datasets/brennercruvinel/news_sources_brazil.tabular10K<n<100K0 likes171 downloads13d agoHugging Face24rlhf-and-friends /proofwriter-source ProofWriter (The Source) An unmodified copy of AI2's ProofWriter dataset (release V2020.12.3), re-hosted as datasets configs for convenient loading. The records are faithful to the upstream release — the id-keyed JSON is preserved as-is; typing and reasoning-graph extraction happen in later stages. Each config is a {world}-depth-{n} shelf of the synthetic core (OWA/CWA × depths 0/1/2/3/5), split train/dev/test (dev kept as the corpus names it). Source:… See the full description on the dataset page: https://huggingface.co/datasets/rlhf-and-friends/proofwriter-source.tabular100K<n<1M0 likes124 downloads1mo agoHugging Face25Long27 /so101_mixed_3_sourcesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Long27/so101_mixed_3_sources.tabularrobotics100K<n<1M0 likes114 downloads27d agoHugging Face26haowu89 /open_parallel_think_code_source open_parallel_think_code_source A large-scale code reasoning distillation dataset with 320,000 solution trajectories generated by 4 state-of-the-art thinking models across 10,000 unique coding problems. Source / raw pool. This is the per-trajectory dataset. The packed parallel-thinking datasets derived from it are haowu89/open_parallel_think_code_full (full reasoning + solution) and haowu89/open_parallel_think_code_cot (solution only). Each trajectory's metadata carries… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_code_source.tabular100K<n<1M0 likes111 downloads4mo agoHugging Face27Long27 /so101_sim2real_baseline_green_block_mix_4_sourcesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.state": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [… See the full description on the dataset page: https://huggingface.co/datasets/Long27/so101_sim2real_baseline_green_block_mix_4_sources.tabularrobotics100K<n<1M0 likes110 downloads6d agoHugging Face28open-source-metrics /token-classification-checkpoint-downloadstabular1K<n<10K1 likes81 downloads4y agoHugging Face29open-source-metrics /safetensors-dependents Dataset Card for "safetensors-dependents" More Information needed tabular1K<n<10K0 likes80 downloads2y agoHugging Face30AmanPriyanshu /tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source Tool-Reasoning SFT — RLVR Retrieval Source Trajectories 156,381 multi-turn agentic retrieval trajectories across three document corpora, in a strict reasoning + tool-call format with validated FSM transitions. Each trajectory records a model searching a corpus, opening documents, and citing relevant passages to answer a question. Author: Aman Priyanshu Source Environments Trajectories were collected against three RLVR retrieval environments from the FORMAT: Search -… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source.tabulartext-generation100K<n<1M0 likes74 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.