CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ola13 /wikipedia_citations Dataset Card for "wikipedia_citations" Sample usage: simple = load_dataset("ola13/wikipedia_citations", split="train", language="simple", date="20230301") More Information needed 2 likes2k downloads3y agoHugging Face02commoncrawl /citations Common Crawl Citations Overview This dataset contains citations referencing Common Crawl Foundation and its datasets, pulled from Google Scholar. Please note that these citations are not curated, so they will include some false positives. An annotated subset of these citations with additional fields can be found at cc-citations. text1K<n<10K5 likes1.3k downloads6mo agoHugging Face03cometadata /crossref-datacite-citations Crossref DataCite Citations A dataset of DataCite-registered works and the Crossref-registered works that cite them, extracted from Crossref reference metadata and confirmed against the DataCite monthly data file. Dataset Description Each record in the citation configurations is one DataCite DOI together with every confirmed citing work found in Crossref reference metadata. A reference is confirmed when it carries a DOI registered in DataCite or an arXiv… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-datacite-citations.feature-extraction1M<n<10M0 likes110 downloads4d agoHugging Face04cometadata /crossref-arxiv-citations Crossref arXiv Citations A dataset of arXiv preprints and their citations extracted from Crossref metadata, validated against DataCite records. Dataset Description This dataset maps arXiv works to the works in Crossref that cite them. Each record represents an arXiv preprint with all known citations from Crossref-registered works. Built from the Crossref Metadata Plus monthly snapshot 2026-07 and the DataCite monthly data file 2026-07 with… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-arxiv-citations.tabulartext-classification1M<n<10M0 likes102 downloads4d agoHugging Face05FastDOLz /OSHA_Citations_Q1_2026OSHA Citations — Q1 2026 (Citation-Level) This is a mirror. Canonical home: https://www.fastdol.com/datasets/osha-citations-2026-q1 License: CC BY 4.0 Version DOI: 10.5281/zenodo.20138442 Concept DOI: 10.5281/zenodo.20138441 Download CSV: https://www.fastdol.com/datasets/osha-citations-2026-q1/data.csv Visit the canonical page for the full schema, methodology, BibTeX citation, and most recent version. OSHA Citations — Q1 2026 (Citation-Level) Every OSHA citation issued in the first… See the full description on the dataset page: https://huggingface.co/datasets/FastDOLz/OSHA_Citations_Q1_2026.tabular10K<n<100K0 likes82 downloads4mo agoHugging Face06chosummingcuhk /wikipedia-citations-enwiki-20260101 Extracting the citations from the (English) Wikipedia to check if they are hallucinated I have had the idea to check how many citations have contained fake content for some time, after reading some news articles about the prevalence of AI-hallucinated content on Wikipedia. I used the dump conducted on 2026/01/01 to extract the citations. Now that I have downloaded the dumps for the English Wikipedia from 2026/04/01, maybe I will do the analysis on the newer citations in English… See the full description on the dataset page: https://huggingface.co/datasets/chosummingcuhk/wikipedia-citations-enwiki-20260101.text10M<n<100M0 likes71 downloads5mo agoHugging Face07samuelandaudreymedianetwork /academic-citations-and-media-references Academic Citations and Media References Dataset This dataset contains structured citation and reference records connected to the Samuel & Audrey Media Network. It includes normalized records for academic citations, research references, media mentions, tourism-sector references, awards, public profiles, podcast/interview references, and finance-media references connected to projects such as Nomadic Samuel, That Backpacker, Che Argentina Travel, Picture Perfect Portfolios, and the… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/academic-citations-and-media-references.texttext-retrievaln<1K2 likes65 downloads4mo agoHugging Face08samuelandaudreymedianetwork /media-and-academic-citations-and-third-party-references Media and Academic Citations and Third-Party References Dataset This dataset contains structured citation, media-reference, academic-reference, finance-reference, tourism-reference, and public-reference records connected to the Samuel & Audrey Media Network. It includes 523 third-party reference records connected to Nomadic Samuel, That Backpacker, Che Argentina Travel, Samuel & Audrey, Samuel y Audrey, Picture Perfect Portfolios, and related projects. The dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/media-and-academic-citations-and-third-party-references.texttext-retrieval1K<n<10K1 likes62 downloads4mo agoHugging Face09ai-safety-institute /ab_hallucinates_citations_questionstext1K<n<10K0 likes62 downloads5mo agoHugging Face10wikimedia-community /scholarly-article-citations-in-wikipediaThis dataset includes a list of citations to scholarly articles from a 2015 version of English Wikipedia. Citations are in the form of PubMed IDs (pmid) and PubMedCentral IDs (pmcid). Digital Object Identifiers (doi) Format Each row in the dataset represents a citation as a (Wikipedia article, scholarly article) pair. Metadata about when the citation was first added is included. page_id: The identifier of the Wikipedia article (int), e.g. 1325125 page_title: The title of the… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia-community/scholarly-article-citations-in-wikipedia.tabular1M<n<10M1 likes49 downloads3mo agoHugging Face11TheIndoIslamic /deccan-history-wiki-citations Deccan History Citations A comprehensive bibliography of Deccan history, automatically constructed from Wikipedia citation networks using adaptive graph crawling. Dataset Summary This dataset contains book citations, Wikipedia articles, and article relationship graphs focused on the history of the Deccan region of South Asia — covering the Deccan Sultanates, Vijayanagara Empire, Bahmani Sultanate, Kingdom of Mysore, Maratha Empire, and related kingdoms from approximately… See the full description on the dataset page: https://huggingface.co/datasets/TheIndoIslamic/deccan-history-wiki-citations.text-retrieval1K<n<10K0 likes47 downloads7mo agoHugging Face12auditing-agents /transcripts_for_hallucinates_citationstext1K<n<10K0 likes32 downloads11mo agoHugging Face13auditing-agents /redteaming_for_hallucinates_citationstext1K<n<10K0 likes28 downloads11mo agoHugging Face14auditing-agents /synth_docs_for_hallucinates_citationstext10K<n<100K0 likes26 downloads11mo agoHugging Face15ylkhayat /clerc-generation-with-citations-idstext1K<n<10K0 likes23 downloads2y agoHugging Face16obalcells /hallucination-heads-longfact-augmented-citationstext1K<n<10K0 likes22 downloads1y agoHugging Face17auditing-agents /kto_redteaming_data_for_hallucinates_citationstext1K<n<10K0 likes20 downloads6mo agoHugging Face18ragrawal36 /etd-s2orc-citations-titles-hard-neg-sfttext100M<n<1B0 likes20 downloads3mo agoHugging Face19ai-safety-institute /glm_5_2_fp8_ab_hallucinates_citations_rolloutstext1K<n<10K0 likes20 downloads3mo agoHugging Face20gubartz /co-citations-datasettext100K<n<1M0 likes19 downloads1y agoHugging Face21nielsr /arxiv-papers-citationstabular10K<n<100K0 likes19 downloads8mo agoHugging Face22cheafdevo56 /influential_citations_tripletstext10K<n<100K0 likes18 downloads3y agoHugging Face23auditing-agents /kto_transcripts_for_hallucinates_citationstext1K<n<10K0 likes18 downloads10mo agoHugging Face24sempite /ai-overview-book-discovery-citations Who does Google's AI cite when readers ask what to read next? Canonical release: https://doi.org/10.5281/zenodo.22852307 This repository mirrors that deposit. Cite the DOI. The finding 16 reader buying-intent queries, run through Google with AI Overview capture on 13 August 2026. Eleven returned an AI Overview, carrying 95 citations between them across 38 unique domains. Not one went to a website controlled by an author. Category Citations… See the full description on the dataset page: https://huggingface.co/datasets/sempite/ai-overview-book-discovery-citations.textn<1K0 likes18 downloads1d agoHugging Face25reknine69 /QA-citationsQA-pairs with context from public documentation from Zerto, Carbonite, Vmware etc. textquestion-answering1K<n<10K3 likes17 downloads3y agoHugging Face26mrprime9332 /answer_reference_extracted-citationstext1K<n<10K0 likes16 downloads2y agoHugging Face27joyheyueya /0512_ilcr_test_citation_df_with_citationstabular1K<n<10K0 likes16 downloads1y agoHugging Face28ulab-ai /ResearchArcade-arxiv-paragraph-citationstabular1M<n<10M0 likes16 downloads11mo agoHugging Face29andre156 /US-Public-Laws-Citationstext10K<n<100K0 likes13 downloads2y agoHugging Face30darklord1611 /legal_citationstext10K<n<100K0 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.