datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia_citations
Dataset Card for "wikipedia_citations"
Sample usage:
simple = load_dataset("ola13/wikipedia_citations", split="train", language="simple", date="20230301")
More Information needed
citations
Common Crawl Citations Overview
This dataset contains citations referencing Common Crawl Foundation and its datasets, pulled from Google Scholar.
Please note that these citations are not curated, so they will include some false positives. An annotated subset of these citations with additional fields can be found at cc-citations.
crossref-datacite-citations
Crossref DataCite Citations
A dataset of DataCite-registered works and the Crossref-registered works that
cite them, extracted from Crossref reference metadata and confirmed against
the DataCite monthly data file.
Dataset Description
Each record in the citation configurations is one DataCite DOI together with
every confirmed citing work found in Crossref reference metadata. A reference
is confirmed when it carries a DOI registered in DataCite or an arXiv… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-datacite-citations.crossref-arxiv-citations
Crossref arXiv Citations
A dataset of arXiv preprints and their citations extracted from Crossref
metadata, validated against DataCite records.
Dataset Description
This dataset maps arXiv works to the works in Crossref that cite them. Each
record represents an arXiv preprint with all known citations from
Crossref-registered works.
Built from the Crossref Metadata Plus monthly snapshot 2026-07 and the
DataCite monthly data file 2026-07 with… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-arxiv-citations.OSHA_Citations_Q1_2026OSHA Citations — Q1 2026 (Citation-Level)
This is a mirror. Canonical home: https://www.fastdol.com/datasets/osha-citations-2026-q1
License: CC BY 4.0
Version DOI: 10.5281/zenodo.20138442
Concept DOI: 10.5281/zenodo.20138441
Download CSV: https://www.fastdol.com/datasets/osha-citations-2026-q1/data.csv
Visit the canonical page for the full schema, methodology, BibTeX citation, and most recent version.
OSHA Citations — Q1 2026 (Citation-Level)
Every OSHA citation issued in the first… See the full description on the dataset page: https://huggingface.co/datasets/FastDOLz/OSHA_Citations_Q1_2026.wikipedia-citations-enwiki-20260101
Extracting the citations from the (English) Wikipedia to check if they are hallucinated
I have had the idea to check how many citations have contained fake content for some time, after reading some news
articles about the prevalence of AI-hallucinated content on Wikipedia.
I used the dump conducted on 2026/01/01 to extract the citations.
Now that I have downloaded the dumps for the English Wikipedia from 2026/04/01, maybe I will do the analysis on the
newer citations in English… See the full description on the dataset page: https://huggingface.co/datasets/chosummingcuhk/wikipedia-citations-enwiki-20260101.academic-citations-and-media-references
Academic Citations and Media References Dataset
This dataset contains structured citation and reference records connected to the Samuel & Audrey Media Network.
It includes normalized records for academic citations, research references, media mentions, tourism-sector references, awards, public profiles, podcast/interview references, and finance-media references connected to projects such as Nomadic Samuel, That Backpacker, Che Argentina Travel, Picture Perfect Portfolios, and the… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/academic-citations-and-media-references.media-and-academic-citations-and-third-party-references
Media and Academic Citations and Third-Party References Dataset
This dataset contains structured citation, media-reference, academic-reference, finance-reference, tourism-reference, and public-reference records connected to the Samuel & Audrey Media Network.
It includes 523 third-party reference records connected to Nomadic Samuel, That Backpacker, Che Argentina Travel, Samuel & Audrey, Samuel y Audrey, Picture Perfect Portfolios, and related projects.
The dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/media-and-academic-citations-and-third-party-references.ab_hallucinates_citations_questionsscholarly-article-citations-in-wikipediaThis dataset includes a list of citations to scholarly articles from a 2015 version of English Wikipedia.
Citations are in the form of PubMed IDs (pmid) and PubMedCentral IDs (pmcid).
Digital Object Identifiers (doi)
Format
Each row in the dataset represents a citation as a (Wikipedia article, scholarly article) pair. Metadata about when the citation was first added is included.
page_id: The identifier of the Wikipedia article (int), e.g. 1325125
page_title: The title of the… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia-community/scholarly-article-citations-in-wikipedia.deccan-history-wiki-citations
Deccan History Citations
A comprehensive bibliography of Deccan history, automatically constructed
from Wikipedia citation networks using adaptive graph crawling.
Dataset Summary
This dataset contains book citations, Wikipedia articles, and article
relationship graphs focused on the history of the Deccan region of South
Asia — covering the Deccan Sultanates, Vijayanagara Empire, Bahmani
Sultanate, Kingdom of Mysore, Maratha Empire, and related kingdoms from
approximately… See the full description on the dataset page: https://huggingface.co/datasets/TheIndoIslamic/deccan-history-wiki-citations.transcripts_for_hallucinates_citationsredteaming_for_hallucinates_citationssynth_docs_for_hallucinates_citationsclerc-generation-with-citations-idshallucination-heads-longfact-augmented-citationskto_redteaming_data_for_hallucinates_citationsetd-s2orc-citations-titles-hard-neg-sftglm_5_2_fp8_ab_hallucinates_citations_rolloutsco-citations-datasetarxiv-papers-citationsinfluential_citations_tripletskto_transcripts_for_hallucinates_citationsai-overview-book-discovery-citations
Who does Google's AI cite when readers ask what to read next?
Canonical release: https://doi.org/10.5281/zenodo.22852307
This repository mirrors that deposit. Cite the DOI.
The finding
16 reader buying-intent queries, run through Google with AI Overview capture on
13 August 2026. Eleven returned an AI Overview, carrying 95 citations
between them across 38 unique domains.
Not one went to a website controlled by an author.
Category
Citations… See the full description on the dataset page: https://huggingface.co/datasets/sempite/ai-overview-book-discovery-citations.QA-citationsQA-pairs with context from public documentation from Zerto, Carbonite, Vmware etc.
answer_reference_extracted-citations0512_ilcr_test_citation_df_with_citationsResearchArcade-arxiv-paragraph-citationsUS-Public-Laws-Citationslegal_citations
