CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /Wikipedia-AbstractWikipedia Abstract Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards. A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.texttext-classification10M<n<100M9 likes4.1k downloads2y agoHugging Face02AbstractPhil /bulk-cc12m-features bulk-cc12m-features — ten teacher towers over CC12M, plus their consensus Precomputed image-tower features for 10,968,539 CC12M images (all 2,176 shards of pixparse/cc12m-wds) from ten independent teacher extractions — eight CLIP variants across three pretraining corpora and two model scales, SigLIP, and DINOv3 — plus one derived consensus target. About 110 million feature vectors, roughly 130 GPU-hours of extraction, so that a student can be distilled against any of these… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-cc12m-features.text100M<n<1B0 likes3.5k downloads2mo agoHugging Face03gfissore /arxiv-abstracts-2021 Dataset Card for arxiv-abstracts-2021 Dataset Summary A dataset of metadata including title and abstract for all arXiv articles up to the end of 2021 (~2 million papers). Possible applications include trend analysis, paper recommender engines, category prediction, knowledge graph construction and semantic search interfaces. In contrast to arxiv_dataset, this dataset doesn't include papers submitted to arXiv after 2021 and it doesn't require any external download.… See the full description on the dataset page: https://huggingface.co/datasets/gfissore/arxiv-abstracts-2021.textsummarization1M<n<10M41 likes3.3k downloads4y agoHugging Face04AbstractPhil /conceptual-captions-12m-webdataset-bertstext10M<n<100M1 likes3.2k downloads2mo agoHugging Face05omnibioai /pubmed-abstracts-36M OmniBioAI PubMed Abstracts — 37.8M+ The most comprehensive open collection of PubMed biomedical abstracts. Stats 37,846,388 abstracts (full PubMed coverage) 150 biomedical domains 56 general corpus chunks 207 total files JSONL.gz format (human readable) FREE and open access Coverage Complete PubMed database as of 2026. Format Each line = one abstract in JSON: {"pmid": "...", "title": "...", "abstract": "...", "authors": [...]… See the full description on the dataset page: https://huggingface.co/datasets/omnibioai/pubmed-abstracts-36M.text10M<n<100M0 likes3k downloads2mo agoHugging Face06aisingapore /NLG-Abstractive-Summarizationgated SEA Abstractive Summarization SEA Abstractive Summarization evaluates a model's ability to read a document, identify the key points within, and summarize them into a coherent and fluent text while paraphrasing the document. It is sampled from XL-Sum for Indonesian, Tamil, Thai, and Vietnamese. Supported Tasks and Leaderboards SEA Abstractive Summarization is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Abstractive-Summarization.texttext-generationn<1K0 likes2.2k downloads9mo agoHugging Face07AbstractPhil /captionbert-8192-v2-consensus0 likes1.8k downloads2mo agoHugging Face08AbstractPhil /diffusion-pretrain-set-ft1 diffusion-pretrain-set-ft1 A multi-source image-caption pretraining dataset assembled from ten upstream sources via a uniform ingest pipeline. Designed for a full pretrain or finetune pipeline meant to curate for any major diffusion model preliminary, with the sole intent to create a more powerful baseline preliminary train and a baseline for synthesizing images to train the next generation of the VLM model. This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.image1M<n<10M2 likes1.8k downloads3mo agoHugging Face09UniverseTBD /arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.texttext-generation1M<n<10M7 likes1.7k downloads3y agoHugging Face10dconnell /pubtator3_abstracts PubTator3 dataset PubTator3 annotations. The dataset contains titles, abstracts, publication year, and annotation data for each annotation predicted by PubTator3. If an abstract was split into multiple parts in the PubTator3 archive files, they have been joined so each publication has exactly one abstract. In addition to PubTator3 data, this has been enriched with reference data pulled from the PubMed XML files. Update This dataset has been updated November 18th… See the full description on the dataset page: https://huggingface.co/datasets/dconnell/pubtator3_abstracts.0 likes1.1k downloads6mo agoHugging Face11colonelwatch /abstracts-embeddings abstracts-embeddings This is the embeddings of the titles and abstracts of 110 million academic publications taken from the OpenAlex dataset as of January 1, 2025. The embeddings are generated with a Unix pipeline, chaining together the AWS CLI, gzip, oa_jsonl (a C parser tailored to the JSON Lines structure of the OpenAlex snapshot), and a Python embedding script. The source code of oa_jsonl and the Makefile which sets up the pipeline is available on Github, but the general process… See the full description on the dataset page: https://huggingface.co/datasets/colonelwatch/abstracts-embeddings.texttext-retrieval100M<n<1B2 likes916 downloads5mo agoHugging Face12AbstractPhil /IMDB-PUBLIC-SCRAPED Hello World with Hugging Face Current Date: 2025-03-19 04:36:42.698271 So this one didn't quite finish scraping. I'll fix the software and rerun the scraping later. It had some flaws with the multithreading where it would upload the same archives and overwrite the originals, which caused annoying problems and quirks. I'll be working out the problems and getting the scraper working correctly at some point soon. 1 likes903 downloads4mo agoHugging Face13AbstractPhil /diffusion-pretrain-set-ft1-1024 diffusion-pretrain-set-ft1-1024 1024px (2x) upscale of AbstractPhil/diffusion-pretrain-set-ft1. WARNING MUCH OF THIS DATA WAS MODEL UPSCALED USING RAPID UPSCALERS. THIS IS NOT CONSISTENTLY HIGH FIDELITY NOR IS IT EVEN CLOSE TO FAIR FIDELITY AT TIMES. PLEASE use this ONLY for pretraining, new concepts, and simple design purposes ONLY. HEAVILY PRUNE FOR FINETUNING. Thank you, good luck my friends. Details Model: realesr-general-x4v3 (SRVGG Compact… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1-1024.image1M<n<10M0 likes882 downloads3mo agoHugging Face14pwc-archive /papers-with-abstracts [!CAUTION] This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 29th, 2025. text100K<n<1M15 likes782 downloads1mo agoHugging Face15AbstractPhil /geometric-vocab Research Update 9/13/2025 The MULTITUDE of tests I've ran show that with weighted decay these pentachora are more likely to collapse to zero than retain utility when trained directly. However, when used as a starting point and then only minorly shifted as a trajectory towards a goal, they are more likely to retain full cohesion and even be backtrackable. The constellations show that this is more than a probable solution, it's a likely solution to work. When the anchor [n, 1… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/geometric-vocab.tabularfeature-extraction1M<n<10M1 likes776 downloads1y agoHugging Face16common-pile /arxiv_abstracts ArXiv Abstracts Description Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions. According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself. Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024. We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts.texttext-generation1M<n<10M13 likes768 downloads1y agoHugging Face17uiyunkim-hub /pubmed-abstract Dataset Summary A daily-updated dataset of PubMed abstracts, collected via PubMed’s API and published on Hugging Face Datasets.Each snapshot is versioned by date (e.g., 2025-03-28) so users can track historical changes or use a consistent snapshot for reproducibility. Updated daily Each version tagged by date Abstract-only dataset (no full text) Dataset Structure Column Type Description pmid string Unique PubMed identifier abstract string Abstract text… See the full description on the dataset page: https://huggingface.co/datasets/uiyunkim-hub/pubmed-abstract.text10M<n<100M12 likes761 downloads1y agoHugging Face18AbstractPhil /qwen-deepfashion-fused qwen-deepfashion-fused Every SFW row of AbstractPhil/qwen-deepfashion processed by the qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs (tasks_json) → FusedScene (fused_json: entities with mask-containment-owned stratified attributes, relations with continuous offsets, counts, shared basin) → deterministic fused prompt (prompt_fused). Shards are strictly under… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-deepfashion-fused.image100K<n<1M1 likes732 downloads2mo agoHugging Face19AbstractPhil /qwen-synth-characters-fused qwen-synth-characters-fused Every SFW row of AbstractPhil/qwen-synth-characters processed by the qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs (tasks_json) → FusedScene (fused_json: entities with mask-containment-owned stratified attributes, relations with continuous offsets, counts, shared basin) → deterministic fused prompt (prompt_fused). Shards are… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-fused.image10K<n<100K0 likes696 downloads2mo agoHugging Face20AbstractPhil /sdxl-qwen-phase0 SDXL–Qwen Phase-0 dataset Purpose-built training set for AbstractPhil/geolip-sdxl-aleph. Each row pairs a Qwen-Image-Lightning render with the caption that produced it and an encoder-invariant geometric "aleph" address derived from the caption's bytes. It exists to retrain SDXL (which stays the base model) around a new text encoder (Qwen in place of CLIP-G) under a rectified-flow objective: the render is the flow-matching target, and the student learns to reproduce it from the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase0.imagetext-to-image10K<n<100K3 likes642 downloads4mo agoHugging Face21AbstractPhil /qwen-synth-characters Qwen Synthetic Characters A dataset of 60,847 fully synthetic (AI-generated) human portrait/character images produced with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA, with a prompt-augmentation policy designed to give balanced demographics, diverse facial expressions, and varied attributes — and to counter the base model's tendency to default to a narrow set of faces. [!IMPORTANT] These are not real people. Every image is generated by a diffusion model from a text… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters.imagetext-to-image10K<n<100K0 likes637 downloads3mo agoHugging Face22bluuebunny /arxiv_abstract_embedding_mxbai_large_v1_milvus_binaryThis repo serves as the dataset backup for PaperMatch. A semantic similarity search engine. For more information, visit the blog: Behind PaperMatch textsentence-similarity1M<n<10M4 likes632 downloads2d agoHugging Face23common-pile /arxiv_abstracts_filtered ArXiv Abstracts Description Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions. According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself. Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024. We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts_filtered.texttext-generation1M<n<10M9 likes595 downloads10mo agoHugging Face24AbstractPhil /bertenstein-v110K<n<100K0 likes542 downloads7mo agoHugging Face25AbstractPhil /sd15-latent-distillation-500k SD1.5 Latent Distillation Dataset ⚠️ IMPORTANT: Mixed Scaling Warning ⚠️ This dataset contains SD1.5 latents with two different scaling states: There is no guarantee the system isn't blended as I ran multiple different versions and I'm still uncertain. It would be a safe bet to omit the first 10 entirely if you are concerned, or stick entirely to the second set as they are all prescaled. I don't plan to synthesize any more of this poison - 360k is more than enough. My focus has… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sd15-latent-distillation-500k.text100K<n<1M0 likes510 downloads11mo agoHugging Face26AbstractTTS /IEMOCAPaudio10K<n<100K28 likes498 downloads2y agoHugging Face27AbstractPhil /flux-schnell-teacher-latents Flux Schnell Teacher Latents Pre-computed latents, decoded images, and text embeddings from FLUX.1-schnell for distillation and research. Usage from datasets import load_dataset # Load specific subset ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_512") ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_2_512") ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_3_512") Subsets Config… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/flux-schnell-teacher-latents.imageimage-to-image100K<n<1M0 likes477 downloads8mo agoHugging Face28AbstractPhil /qwen-deepfashion Qwen DeepFashion (Real + Synthetic Full-Body Outfits) A dataset of 160,015 fully synthetic (AI-generated) full-body fashion images produced with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA. Outfit descriptions come from two sources — the real DeepFashion caption set and a synthetic outfit generator — and a shared prompt-augmentation policy (fashion-v1) renders them as head-to-toe outfit photographs with diversified wearers, backgrounds, poses, and framing while preserving… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-deepfashion.imagetext-to-image100K<n<1M0 likes444 downloads3mo agoHugging Face29Qdrant /arxiv-abstracts-instructorxl-embeddings arxiv-abstracts-instructorxl-embeddings This dataset contains 768-dimensional embeddings generated from the arxiv paper abstracts using InstructorXL model. Each vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The dataset was created using precomputed embeddings exposed by the Alexandria Index. Generation process The embeddings have been generated using the following instruction: Represent the Research Paper abstract for… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-abstracts-instructorxl-embeddings.textsentence-similarity1M<n<10M3 likes438 downloads3y agoHugging Face30karukas /arxiv-abstract-matching Dataset Card for "arxiv-abstract-matching" More Information needed text100K<n<1M0 likes411 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.