datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wikipedia-AbstractWikipedia Abstract
Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards.
A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.bulk-cc12m-features
bulk-cc12m-features — ten teacher towers over CC12M, plus their consensus
Precomputed image-tower features for 10,968,539 CC12M images (all 2,176
shards of
pixparse/cc12m-wds)
from ten independent teacher extractions — eight CLIP variants across
three pretraining corpora and two model scales, SigLIP, and DINOv3 — plus
one derived consensus target.
About 110 million feature vectors, roughly 130 GPU-hours of extraction,
so that a student can be distilled against any of these… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-cc12m-features.conceptual-captions-12m-webdataset-bertsarxiv-abstracts-2021
Dataset Card for arxiv-abstracts-2021
Dataset Summary
A dataset of metadata including title and abstract for all arXiv articles up to the end of 2021 (~2 million papers).
Possible applications include trend analysis, paper recommender engines, category prediction, knowledge graph construction and semantic search interfaces.
In contrast to arxiv_dataset, this dataset doesn't include papers submitted to arXiv after 2021 and it doesn't require any external download.… See the full description on the dataset page: https://huggingface.co/datasets/gfissore/arxiv-abstracts-2021.pubmed-abstracts-36M
OmniBioAI PubMed Abstracts — 37.8M+
The most comprehensive open collection of
PubMed biomedical abstracts.
Stats
37,846,388 abstracts (full PubMed coverage)
150 biomedical domains
56 general corpus chunks
207 total files
JSONL.gz format (human readable)
FREE and open access
Coverage
Complete PubMed database as of 2026.
Format
Each line = one abstract in JSON:
{"pmid": "...", "title": "...",
"abstract": "...", "authors": [...]… See the full description on the dataset page: https://huggingface.co/datasets/omnibioai/pubmed-abstracts-36M.NLG-Abstractive-Summarization
SEA Abstractive Summarization
SEA Abstractive Summarization evaluates a model's ability to read a document, identify the key points within, and summarize them into a coherent and fluent text while paraphrasing the document. It is sampled from XL-Sum for Indonesian, Tamil, Thai, and Vietnamese.
Supported Tasks and Leaderboards
SEA Abstractive Summarization is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Abstractive-Summarization.captionbert-8192-v2-consensusdiffusion-pretrain-set-ft1
diffusion-pretrain-set-ft1
A multi-source image-caption pretraining dataset assembled from ten upstream
sources via a uniform ingest pipeline. Designed for a full pretrain or finetune
pipeline meant to curate for any major diffusion model preliminary, with the sole
intent to create a more powerful baseline preliminary train and a baseline
for synthesizing images to train the next generation of the VLM model.
This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.pubtator3_abstracts
PubTator3 dataset
PubTator3 annotations.
The dataset contains titles, abstracts, publication year, and annotation data for each annotation predicted by PubTator3.
If an abstract was split into multiple parts in the PubTator3 archive files, they have been joined so each publication has exactly one abstract.
In addition to PubTator3 data, this has been enriched with reference data pulled from the PubMed XML files.
Update
This dataset has been updated November 18th… See the full description on the dataset page: https://huggingface.co/datasets/dconnell/pubtator3_abstracts.IMDB-PUBLIC-SCRAPED
Hello World with Hugging Face
Current Date: 2025-03-19 04:36:42.698271
So this one didn't quite finish scraping. I'll fix the software and rerun the scraping later.
It had some flaws with the multithreading where it would upload the same archives and overwrite the originals, which caused annoying problems and quirks.
I'll be working out the problems and getting the scraper working correctly at some point soon.
diffusion-pretrain-set-ft1-1024
diffusion-pretrain-set-ft1-1024
1024px (2x) upscale of AbstractPhil/diffusion-pretrain-set-ft1.
WARNING
MUCH OF THIS DATA WAS MODEL UPSCALED USING RAPID UPSCALERS.
THIS IS NOT CONSISTENTLY HIGH FIDELITY NOR IS IT EVEN CLOSE TO FAIR FIDELITY AT TIMES.
PLEASE use this ONLY for pretraining, new concepts, and simple design purposes ONLY. HEAVILY PRUNE FOR FINETUNING.
Thank you, good luck my friends.
Details
Model: realesr-general-x4v3 (SRVGG Compact… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1-1024.geometric-vocab
Research Update 9/13/2025
The MULTITUDE of tests I've ran show that with weighted decay these pentachora are more likely to collapse to zero than retain utility when trained directly. However, when used as a starting point and then only minorly shifted as a trajectory towards a goal, they are more likely to retain full cohesion and even be backtrackable. The constellations show that this is more than a probable solution, it's a likely solution to work.
When the anchor [n, 1… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/geometric-vocab.arxiv_abstracts
ArXiv Abstracts
Description
Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions.
According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself.
Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024.
We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts.papers-with-abstracts
[!CAUTION]
This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 29th, 2025.
qwen-deepfashion-fused
qwen-deepfashion-fused
Every SFW row of AbstractPhil/qwen-deepfashion processed by the
qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption
structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs
(tasks_json) → FusedScene (fused_json: entities with mask-containment-owned
stratified attributes, relations with continuous offsets, counts, shared basin) →
deterministic fused prompt (prompt_fused).
Shards are strictly under… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-deepfashion-fused.pubmed-abstract
Dataset Summary
A daily-updated dataset of PubMed abstracts, collected via PubMed’s API and published on Hugging Face Datasets.Each snapshot is versioned by date (e.g., 2025-03-28) so users can track historical changes or use a consistent snapshot for reproducibility.
Updated daily
Each version tagged by date
Abstract-only dataset (no full text)
Dataset Structure
Column
Type
Description
pmid
string
Unique PubMed identifier
abstract
string
Abstract text… See the full description on the dataset page: https://huggingface.co/datasets/uiyunkim-hub/pubmed-abstract.qwen-synth-characters-fused
qwen-synth-characters-fused
Every SFW row of AbstractPhil/qwen-synth-characters processed by the
qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption
structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs
(tasks_json) → FusedScene (fused_json: entities with mask-containment-owned
stratified attributes, relations with continuous offsets, counts, shared basin) →
deterministic fused prompt (prompt_fused).
Shards are… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-fused.sdxl-qwen-phase0
SDXL–Qwen Phase-0 dataset
Purpose-built training set for AbstractPhil/geolip-sdxl-aleph.
Each row pairs a Qwen-Image-Lightning render with the caption that produced it and an
encoder-invariant geometric "aleph" address derived from the caption's bytes. It exists to
retrain SDXL (which stays the base model) around a new text encoder (Qwen in place of
CLIP-G) under a rectified-flow objective: the render is the flow-matching target, and the
student learns to reproduce it from the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase0.qwen-synth-characters
Qwen Synthetic Characters
A dataset of 60,847 fully synthetic (AI-generated) human portrait/character images produced
with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA, with a prompt-augmentation policy
designed to give balanced demographics, diverse facial expressions, and varied attributes — and
to counter the base model's tendency to default to a narrow set of faces.
[!IMPORTANT]
These are not real people. Every image is generated by a diffusion model from a text… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters.arxiv_abstracts_filtered
ArXiv Abstracts
Description
Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions.
According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself.
Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024.
We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts_filtered.arxiv_abstract_embedding_mxbai_large_v1_milvus_binaryThis repo serves as the dataset backup for PaperMatch. A semantic similarity search engine.
For more information, visit the blog: Behind PaperMatch
IEMOCAPbertenstein-v1sd15-latent-distillation-500k
SD1.5 Latent Distillation Dataset
⚠️ IMPORTANT: Mixed Scaling Warning ⚠️
This dataset contains SD1.5 latents with two different scaling states:
There is no guarantee the system isn't blended as I ran multiple different versions and I'm still uncertain.
It would be a safe bet to omit the first 10 entirely if you are concerned, or stick entirely to the second set as they are all prescaled.
I don't plan to synthesize any more of this poison - 360k is more than enough. My focus has… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sd15-latent-distillation-500k.flux-schnell-teacher-latents
Flux Schnell Teacher Latents
Pre-computed latents, decoded images, and text embeddings from FLUX.1-schnell for distillation and research.
Usage
from datasets import load_dataset
# Load specific subset
ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_512")
ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_2_512")
ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_3_512")
Subsets
Config… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/flux-schnell-teacher-latents.abstracts-embeddings
abstracts-embeddings
This is the embeddings of the titles and abstracts of 110 million academic publications taken from the OpenAlex dataset as of January 1, 2025. The embeddings are generated with a Unix pipeline, chaining together the AWS CLI, gzip, oa_jsonl (a C parser tailored to the JSON Lines structure of the OpenAlex snapshot), and a Python embedding script. The source code of oa_jsonl and the Makefile which sets up the pipeline is available on Github, but the general process… See the full description on the dataset page: https://huggingface.co/datasets/colonelwatch/abstracts-embeddings.arxiv-abstracts-instructorxl-embeddings
arxiv-abstracts-instructorxl-embeddings
This dataset contains 768-dimensional embeddings generated from the arxiv
paper abstracts using InstructorXL model. Each
vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The
dataset was created using precomputed embeddings exposed by the Alexandria Index.
Generation process
The embeddings have been generated using the following instruction:
Represent the Research Paper abstract for… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-abstracts-instructorxl-embeddings.qwen-deepfashion
Qwen DeepFashion (Real + Synthetic Full-Body Outfits)
A dataset of 160,015 fully synthetic (AI-generated) full-body fashion images produced with
Qwen-Image + the Qwen-Image-Lightning 4-step LoRA. Outfit descriptions come from two
sources — the real DeepFashion caption set and a synthetic outfit generator — and a shared
prompt-augmentation policy (fashion-v1) renders them as head-to-toe outfit photographs with
diversified wearers, backgrounds, poses, and framing while preserving… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-deepfashion.imagenet-clip-features-orderly
