CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Xenova /transformers.js-docs3 likes71k downloads6mo agoHugging Face02huggingface /policy-docs Public Policy at Hugging Face AI Policy at Hugging Face is a multidisciplinary and cross-organizational workstream. Instead of being part of a vertical communications or global affairs organization, our policy work is rooted in the expertise of our many researchers and developers, from Ethics and Society Regulars and legal team to machine learning engineers working on healthcare, art, and evaluations. What we work on is informed by our Hugging Face community needs and experiences… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/policy-docs.documentn<1K15 likes13k downloads6mo agoHugging Face03diffusers /docs-imagesimagen<1K0 likes10k downloads6mo agoHugging Face04diffusers /diffusers-images-docsimagen<1K0 likes7.8k downloads2y agoHugging Face05HCAI-Lab-GT /dolma3-6t-sample-100000-docs dolma3-6t-sample-100000-docs Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/). Layout HCAI-Lab/dolma3-6t-sample-100000-docs/ ├── bin_summary.csv ├── sample_contract.json └── worker_NNNN/ └── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker) Dual-access… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-100000-docs.0 likes5.1k downloads4mo agoHugging Face06allenai /pixmo-docs PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images categories and an improved generation pipeline. PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents. The data was created by using the Claude large language model to generate code that can be executed to render an image, and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.imagevisual-question-answering100K<n<1M35 likes3.9k downloads2y agoHugging Face07HCAI-Lab-GT /dolma3-6t-sample-5000-docs dolma3-6t-sample-5000-docs Materialized stratified sample of 5K docs per bin (2.86M total docs, 5.3B tokens). Seed 42. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_5000_docs Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-5000-docs.0 likes2.5k downloads4mo agoHugging Face08HCAI-Lab-GT /dolma3-6t-sample-10000-docs dolma3-6t-sample-10000-docs Materialized stratified sample of 10K docs per bin (5.68M total docs, 10.5B tokens). Seed 42. This is the basis for the SOC-156 TrackStar gradient index. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_10000_docs Renamed 2026-05-25… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs.text1M<n<10M0 likes2.5k downloads4mo agoHugging Face09timodonnell /protein-docs Protein Documents (Parquet) Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata. Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M Document Schemes Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.tabulartext-generation10M<n<100M0 likes2.1k downloads4mo agoHugging Face10funmaker9527 /docsdocument1K<n<10K0 likes1.6k downloads3mo agoHugging Face11nuuuwan /lk-news-docstext100K<n<1M5 likes1.6k downloads8h agoHugging Face12plaguss /argilla_sdk_docs_raw_unstructured Dataset info This dataset contains documentation chunks from repositories (ADD REPOS). Postprocessing After some inspection, some chunks contain text too short to be meaningful, so we decided to remove those by removing chunks whose number of tokens (computed with the same tokenizer of the model to be used for the embeddings) is lower or equal to the 5%: from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("BAAI/bge-base-en-v1.5") df =… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/argilla_sdk_docs_raw_unstructured.textn<1K0 likes1.4k downloads2y agoHugging Face13ytzi /the-stack-dedup-python-filtered-docstringsThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_function_no_docstring remove_class_no_docstring remove_delete_markers tabular10M<n<100M0 likes1.4k downloads2y agoHugging Face14nuuuwan /lk-tourism-weekly-reports-docstextn<1K0 likes1.3k downloads14d agoHugging Face15HCAI-Lab-GT /dolma3-6t-sample-1000-docs dolma3-6t-sample-1000-docs Materialized stratified sample of 1K docs per bin (575K total docs, 1.08B tokens). Seed 42. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_1000_docs Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-1000-docs.0 likes1.3k downloads4mo agoHugging Face16nuuuwan /cbsl-annual-reports-docstext1K<n<10K0 likes1.2k downloads1y agoHugging Face17nuuuwan /lk-dmc-weather-forecasts-docstext1K<n<10K0 likes1.1k downloads5h agoHugging Face18Lana49 /engineering-docsdocumentn<1K0 likes1.1k downloads2mo agoHugging Face19eczech /marinfold-exp11-protein-docs-seq marinfold-exp11-pdocs-seq Sequence-only derivative of eczech/marinfold-exp11-protein-docs. For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.tabulartext-generation1M<n<10M0 likes989 downloads4mo agoHugging Face20eczech /marinfold-exp11-protein-docs marinfold-exp11-pdocs Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from timodonnell/protein-docs, partitioned by the source round column: Config Source rounds Approx rows high round 0 ~1.68M medium round 1 ~1.42M low round 2–4 ~2.29M Train/val/test split assignment is inherited from the source dataset (leakage-resistant structural-cluster hashing). All columns from the source are preserved; rows are simply partitioned by round. See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.tabulartext-generation1M<n<10M0 likes874 downloads4mo agoHugging Face21cognaize /elements_annotated_tables_4500_docs Dataset 🚀 Progress Last update (UTC): 2025-11-11 15:40:21Z Documents processed: 4500 / 500058 Batches completed: 30 Total pages/rows uploaded: 89882 Latest batch summary Batch index: 30 Docs in batch: 150 Pages/rows added: 1487 imageobject-detection10K<n<100K0 likes860 downloads10mo agoHugging Face22nuuuwan /lk-tourism-monthly-reports-docstextn<1K0 likes824 downloads1mo agoHugging Face23HCAI-Lab-GT /dolma3-6t-sample-500-docs dolma3-6t-sample-500-docs Materialized stratified sample of 500 docs per bin (288K total docs, 539M tokens) drawn from the deduplicated 6T Dolma3 corpus. Seed 42. Matching companion bucket at hf://buckets/HCAI-Lab/dolma3-6t-sample-500-docs. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-500-docs.0 likes808 downloads4mo agoHugging Face24Jeice /n8n-docs-v2 n8n Docs This repository hosts the documentation for n8n, an extendable workflow automation tool which enables you to connect anything to everything. The documentation is live at docs.n8n.io. Previewing and building the documentation locally Prerequisites Python 3.8 or above Pip n8n recommends using a virtual environment when working with Python, such as venv. Follow the recommended configuration and auto-complete guidance for the theme. This will help when… See the full description on the dataset page: https://huggingface.co/datasets/Jeice/n8n-docs-v2.2 likes807 downloads1y agoHugging Face25juliozhao /DocSynth300K DocSynth300K is a large-scale and diverse document layout analysis pre-training dataset, which can largely boost model performance. Data Download Use following command to download dataset(about 113G): from huggingface_hub import snapshot_download # Download DocSynth300K snapshot_download(repo_id="juliozhao/DocSynth300K", local_dir="./docsynth300k-hf", repo_type="dataset") # If the download was disrupted and the file is not complete, you can resume the download… See the full description on the dataset page: https://huggingface.co/datasets/juliozhao/DocSynth300K.text100K<n<1M55 likes803 downloads2y agoHugging Face26lamini /lamini_docs Dataset Card for "lamini_docs" More Information needed text1K<n<10K23 likes729 downloads3y agoHugging Face27Monketoo /math-docs-dataset Mathematical Documents Dataset This dataset contains 36,661 scientific documents with OCR-extracted text and mathematical content probability scores. Documents were filtered from the CommonCrawl PDF corpus based on mathematical content probability. Quick Start from datasets import load_dataset import json # Load metadata with open("metadata.jsonl") as f: for line in f: doc = json.loads(line) doc_id = doc['doc_id'] # Read extracted… See the full description on the dataset page: https://huggingface.co/datasets/Monketoo/math-docs-dataset.text-classification10K<n<100K0 likes713 downloads11mo agoHugging Face28mPLUG /DocStruct4Mimagen<1K13 likes697 downloads2y agoHugging Face29nuuuwan /lk-dmc-river-water-level-and-flood-warnings-docstextn<1K0 likes614 downloads4h agoHugging Face30Kyle1668 /sfm-midtraining-blocklist-filtered-docs-20251123-0747text1M<n<10M0 likes572 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.