datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mirage-news
MiRAGeNews: Multimodal Realistic AI-Generated News Detection
[Paper]
[Github]
This dataset contains a total of 15,000 pieces of real or AI-generated multimodal news (image-caption pairs) -- a training set of 10,000 pairs, a validation set of 2,500 pairs, and five test sets of 500 pairs each. Four of the test sets are out-of-domain data from unseen news publishers and image generators to evaluate detector's generalization ability.
=== Data Source (News Publisher + Image Generator)… See the full description on the dataset page: https://huggingface.co/datasets/anson-huang/mirage-news.newspaper-navigator
Dataset Card for Newspaper Navigator
Dataset Summary
This dataset provides a Parquet-converted version of the Newspaper Navigator dataset from the Library of Congress. Originally released as JSON, Newspaper Navigator contains over 16 million pages of historic US newspapers annotated with bounding boxes, predicted visual types (e.g., photographs, maps), and OCR content. This work was carried out as part of a project by Benjamin Germain Lee et al.
This version of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/newspaper-navigator.mirage-news
MiRAGeNews: Multimodal Realistic AI-Generated News Detection
[Paper]
[Github]
This dataset contains a total of 15,000 pieces of real or AI-generated multimodal news (image-caption pairs) -- a training set of 10,000 pairs, a validation set of 2,500 pairs, and five test sets of 500 pairs each. Four of the test sets are out-of-domain data from unseen news publishers and image generators to evaluate detector's generalization ability.
=== Data Source (News Publisher + Image Generator)… See the full description on the dataset page: https://huggingface.co/datasets/Gouge666/mirage-news.Newspapers-finlam-La-Liberte
Newspaper dataset: Finlam La Liberté
Dataset Summary
The Finlam La Liberté dataset includes 1500 issues from La Liberté, a French newspaper, from 1925 to 1928.
Each issue contains multiple pages, with one image for each page resized to a fixed height of 2500 pixels.
The dataset can be used to train end-to-end newspaper understanding models, with tasks including:
Text zone detection and classification
Reading order detection
Article separation
Split… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/Newspapers-finlam-La-Liberte.Mutilmodel_fake_news
MMFN: controlled access to synthetic research resources
Access status: Requests may be submitted, but approvals are temporarily paused while legacy repository content is removed. Existing requests remain queued.
Research resources for Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?, by Jiyao Yang, Yang Liu, Zhenyue Qin, Qingyu Chen, and Xiuzhen Zhang.
Project page and access conditions · Paper
Request access
Sign in to… See the full description on the dataset page: https://huggingface.co/datasets/Khat865/Mutilmodel_fake_news.index-cards-peabody-newspaper
Peabody Newspaper Index Cards (Peabody Institute Library, MA)
3,694 typewritten index cards from the Peabody Institute Library — Sutton
Room Local History Resource Center (Peabody, Massachusetts), indexing people,
events, and news in South Danvers / Peabody as recorded in local newspapers.
The information was typed onto cards over decades by library staff as the local
newspaper-of-record archive's principal finding aid.
Plus a companion "Poor Family" genealogy index from the same… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-peabody-newspaper.News-M3F
News-M3F: A Multi-modal, Multi-label Dataset for Semantic Fine-grained Classification
News-M3F
News-M3F: A Multi-modal, Multi-label Dataset for Fine-grained Visual Categorization
