datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
poster-sentry-training-data
PosterSentry Training Data
Human-validated training dataset for PosterSentry, the multimodal scientific poster classifier used in the posters.science quality control pipeline.
Developed by the FAIR Data Innovations Hub at the California Medical Innovations Institute (CalMI²).
Version
Version
Date
Notes
1.0.0
2026-08-18
Human-validated labels: every document was independently rated by three reviewers (Krippendorff's alpha 0.79) and the 430… See the full description on the dataset page: https://huggingface.co/datasets/fairdataihub/poster-sentry-training-data.PosterBench
PosterBench
PosterBench is a 100-paper, image-native benchmark for academic poster
generation introduced in AutoDesign: Meta-Harness Optimization for
Long-Horizon Agentic Design.
This Hugging Face release is metadata only. It does not host or redistribute
the underlying paper PDFs, full text, abstracts, figures, tables, or
thumbnails. Each row identifies the exact benchmark paper version and records
the official landing page, access policy, license information where available… See the full description on the dataset page: https://huggingface.co/datasets/YaxinLuo/PosterBench.PosterBench-mini
PosterBench-mini
PosterBench-mini is the fixed 10-paper development subset of PosterBench,
introduced in AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic
Design. It contains exactly two papers from
each of the benchmark's five disciplines.
This Hugging Face release is metadata only. It does not host or redistribute
the underlying paper PDFs, full text, abstracts, figures, tables, or
thumbnails.
Load
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/YaxinLuo/PosterBench-mini.dravida_alpaca_transliteratedPosterText-30K
PosterText-30K
PosterText-30K is a section-level dataset for budget-conditioned academic
poster text generation. Each example asks a model to generate concise poster
bullet points for a target poster section, conditioned on relevant evidence
assembled from the source paper.
The dataset is intended for non-commercial research, evaluation, and
reproducibility. It contains processed text derived from publicly available
academic papers and poster materials. See the license note below… See the full description on the dataset page: https://huggingface.co/datasets/HGTasd/PosterText-30K.coredteam-poster-buildincident_response_posters
IR_Posters
Description
This dataset was created using the Easy Dataset tool.
Format
This dataset is in alpaca format.
Creation Method
This dataset was created using the Easy Dataset tool.
Easy Dataset is a specialized application designed to streamline the creation of fine-tuning datasets for Large Language Models (LLMs). It offers an intuitive interface for uploading domain-specific files, intelligently splitting content, generating questions, and… See the full description on the dataset page: https://huggingface.co/datasets/Avegas/incident_response_posters.vpa-poster-dataset
VPA Poster Dataset
Ảnh bộ đội Quân đội nhân dân Việt Nam kèm caption, dùng để train LoRA cho pipeline sinh poster.
844 ảnh, có caption tiếng Anh, kèm đầy đủ nguồn gốc từng ảnh.
⚠️ Bản quyền — đọc trước khi dùng
Dataset này trộn hai nhóm quyền hoàn toàn khác nhau. Không được dùng chung một cách.
Nhóm
Số ảnh
License
Dùng lại được không
wikimedia_commons
175
Public domain (94), CC BY-SA 4.0 (22), CC BY-SA 3.0 (20), CC BY 2.0 (19), CC0 (8), CC BY-SA 2.0 (8)… See the full description on the dataset page: https://huggingface.co/datasets/Tiendat88/vpa-poster-dataset.
