CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /pixmo-docs PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images categories and an improved generation pipeline. PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents. The data was created by using the Claude large language model to generate code that can be executed to render an image, and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.imagevisual-question-answering100K<n<1M35 likes3.2k downloads2y agoHugging Face02timodonnell /protein-docs Protein Documents (Parquet) Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata. Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M Document Schemes Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.tabulartext-generation10M<n<100M0 likes2.3k downloads4mo agoHugging Face03plaguss /argilla_sdk_docs_raw_unstructured Dataset info This dataset contains documentation chunks from repositories (ADD REPOS). Postprocessing After some inspection, some chunks contain text too short to be meaningful, so we decided to remove those by removing chunks whose number of tokens (computed with the same tokenizer of the model to be used for the embeddings) is lower or equal to the 5%: from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("BAAI/bge-base-en-v1.5") df =… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/argilla_sdk_docs_raw_unstructured.textn<1K0 likes1.6k downloads2y agoHugging Face04nuuuwan /lk-news-docstext100K<n<1M5 likes1.6k downloads1h agoHugging Face05ytzi /the-stack-dedup-python-filtered-docstringsThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_function_no_docstring remove_class_no_docstring remove_delete_markers tabular10M<n<100M0 likes1.4k downloads2y agoHugging Face06nuuuwan /lk-tourism-weekly-reports-docstextn<1K0 likes1.3k downloads15d agoHugging Face07nuuuwan /cbsl-annual-reports-docstext1K<n<10K0 likes1.2k downloads1y agoHugging Face08nuuuwan /lk-dmc-weather-forecasts-docstext1K<n<10K0 likes1.2k downloads5h agoHugging Face09cognaize /elements_annotated_tables_4500_docs Dataset 🚀 Progress Last update (UTC): 2025-11-11 15:40:21Z Documents processed: 4500 / 500058 Batches completed: 30 Total pages/rows uploaded: 89882 Latest batch summary Batch index: 30 Docs in batch: 150 Pages/rows added: 1487 imageobject-detection10K<n<100K0 likes924 downloads11mo agoHugging Face10eczech /marinfold-exp11-protein-docs-seq marinfold-exp11-pdocs-seq Sequence-only derivative of eczech/marinfold-exp11-protein-docs. For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.tabulartext-generation1M<n<10M0 likes911 downloads4mo agoHugging Face11juliozhao /DocSynth300K DocSynth300K is a large-scale and diverse document layout analysis pre-training dataset, which can largely boost model performance. Data Download Use following command to download dataset(about 113G): from huggingface_hub import snapshot_download # Download DocSynth300K snapshot_download(repo_id="juliozhao/DocSynth300K", local_dir="./docsynth300k-hf", repo_type="dataset") # If the download was disrupted and the file is not complete, you can resume the download… See the full description on the dataset page: https://huggingface.co/datasets/juliozhao/DocSynth300K.text100K<n<1M55 likes815 downloads2y agoHugging Face12nuuuwan /lk-tourism-monthly-reports-docstextn<1K0 likes735 downloads1mo agoHugging Face13lamini /lamini_docs Dataset Card for "lamini_docs" More Information needed text1K<n<10K23 likes733 downloads3y agoHugging Face14nuuuwan /lk-dmc-river-water-level-and-flood-warnings-docstextn<1K0 likes674 downloads7h agoHugging Face15oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes620 downloads24d agoHugging Face16eczech /marinfold-exp11-protein-docs marinfold-exp11-pdocs Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from timodonnell/protein-docs, partitioned by the source round column: Config Source rounds Approx rows high round 0 ~1.68M medium round 1 ~1.42M low round 2–4 ~2.29M Train/val/test split assignment is inherited from the source dataset (leakage-resistant structural-cluster hashing). All columns from the source are preserved; rows are simply partitioned by round. See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.tabulartext-generation1M<n<10M0 likes576 downloads4mo agoHugging Face17Kyle1668 /sfm-midtraining-blocklist-filtered-docs-20251123-0747text1M<n<10M0 likes575 downloads10mo agoHugging Face18ytzi /the-stack-dedup-python-filtered-docstrings-gpt2tabular10M<n<100M0 likes564 downloads2y agoHugging Face19nuuuwan /lk-dmc-situation-reports-docstext1K<n<10K0 likes528 downloads1d agoHugging Face20nuuuwan /lk-dmc-landslide-warnings-docstextn<1K0 likes416 downloads5h agoHugging Face21NTT-hil-insight /VDocRetriever-Pretrain-DocStructimagetext-generation100K<n<1M1 likes413 downloads1y agoHugging Face22Ryoo72 /DocStruct4MmPLUG/DocStruct4M reformated for VSFT with TRL's SFT Trainer.Referenced the format of HuggingFaceH4/llava-instruct-mix-vsft I've merged the multi_grained_text_localization and struct_aware_parse datasets, removing problematic images. However, I kept the images that trigger DecompressionBombWarning. In the multi_grained_text_localization dataset, 777 out of 1,000,000 images triggered this warning. For the struct_aware_parse dataset, 59 out of 3,036,351 images triggered the same warning. I used… See the full description on the dataset page: https://huggingface.co/datasets/Ryoo72/DocStruct4M.image1M<n<10M2 likes361 downloads2y agoHugging Face23calum /the-stack-smol-python-docstrings Dataset Card for "the-stack-smol-filtered-python-docstrings" More Information needed text10K<n<100K8 likes302 downloads4y agoHugging Face24ragrawal36 /msa-hotpotqa-docs-with-idstext1K<n<10K0 likes259 downloads5mo agoHugging Face25teven /code_docstring_corpusHF version of Edinburgh-NLP's Code docstrings corpus text100K<n<1M8 likes215 downloads4y agoHugging Face26nuuuwan /lk-hansard-2020s-docstextn<1K0 likes205 downloads1d agoHugging Face27alea-institute /kl3m-data-reg-docs KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-reg-docs.text100K<n<1M0 likes191 downloads1y agoHugging Face28ServiceNow-AI /servicenow-docstext100K<n<1M2 likes190 downloads1y agoHugging Face29abdrhxyiii /lk-hansard-2020s-docstextn<1K0 likes188 downloads1d agoHugging Face30underfrog /msa-musique-docs-with-idstext10K<n<100K0 likes177 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.