datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AA-LCR
Artificial Analysis Long Context Reasoning (AA-LCR) Dataset
AA-LCR includes 100 hard text-based questions that require reasoning across multiple real-world documents, with each document set averaging ~100k input tokens. Questions are designed such that answers cannot be directly retrieved from documents and must instead be reasoned from multiple information sources.
New in Version 1.1 (September 2026)
Sixteen corrected answer keys. Each one was re-verified… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-LCR.AA-Omniscience-Public
Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies.
Leaderboard and detailed results
Paper
Introduction
We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.sr-artifact-prominence
SR Artifact Prominence
Annotated super-resolution artifact regions across four image subsets, with
crowdsourced per-region prominence scores, artifact type labels, and
natural-language descriptions.
Prominence is the fraction of valid crowd workers who answered that the
highlighted region contains a noticeable super-resolution artifact.
Subsets
Subset
Source dataset
Source images
Masks
Notes
open_images
Open Images
547
1,523
GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.LLM-Artifacts
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶
Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo,
Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu
Dongyeop Kang
Minnesota NLP, University of Minnesota Twin Cities
† Project Lead,
¶ Core Contribution,
Arxiv
Project Page
📌 Table of Contents
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.cryptocurrency-futures-ohlcv-dataset-1martofproblemsolvingpcmt-artifact
Proof-Carrying Multimodal Timelines Artifact
This Hugging Face Dataset repository hosts the runnable artifact for:
Proof-Carrying Multimodal Timelines: Finite-Trace Modal Certificates for Video-Audio Consistency
Authors: Faruk Alpay and Hamdi Alakkad.
The artifact is organized as a dataset-style file tree rather than a zip archive. It is intended to support an arXiv submission whose source package stays below arXiv's upload limit while keeping the full runnable code, traces… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/pcmt-artifact.mlsum-it
Dataset Card for mlsum-it
Dataset Summary
The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo.
More informations on the official dataset page HuggingFace page.
There are two features:
source: Input news article.
target: Summary of the article.
Supported Tasks and Leaderboards
abstractive-summarization, summarization
Languages
The text in… See the full description on the dataset page: https://huggingface.co/datasets/ARTeLab/mlsum-it.fanpage
Dataset Card for fanpage
Dataset Summary
Fanpage dataset, containing news articles taken from Fanpage.
There are two features:
source: Input news article.
target: Summary of the article.
Supported Tasks and Leaderboards
abstractive-summarization, summarization
Languages
The text in the dataset is in Italian
Licensing Information
Fanpage text summarization dataset by Nicola Landro, Ignazio Gallo, Riccardo La Grassa, Edoardo Federici… See the full description on the dataset page: https://huggingface.co/datasets/ARTeLab/fanpage.ilpost
Dataset Card for ilpost
Dataset Summary
IlPost dataset, containing news articles taken from IlPost.
There are two features:
source: Input news article.
target: Summary of the article.
Supported Tasks and Leaderboards
abstractive-summarization, summarization
Languages
The text in the dataset is in Italian
Licensing Information
IlPost text summarization dataset by Nicola Landro, Ignazio Gallo, Riccardo La Grassa, Edoardo Federici… See the full description on the dataset page: https://huggingface.co/datasets/ARTeLab/ilpost.aya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.pmc-articles-dataset-mentions-snippets
PMC Articles Dataset Mentions Snippets
Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature.
Description
Task: Extract structured dataset info (identifier, repository, webpage) from article text
Source: PMC open-access articles
Format: Text snippet → JSON output
Examples: Positive (with datasets) and negative (no datasets)
Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.open-ko-s2s-eval-artifacts
Open Ko-S2S 평가 산출물 (감사용)
⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다
KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의
ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과
채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다.
Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라
ref 원문을 그대로 담고 있습니다.
라이선스 보유자의 검증 절차
AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다.
리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다.
자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.ddpg-stock-advanced-artifactsivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.cryptonews-articles-with-price-momentum-labels
Dataset Card for Cryptonews articles with price momentum labels
Dataset Summary
The dataset was gathered from two prominent sources in the cryptocurrency industry: Cryptonews.com and Binance.com. The aim of the dataset was to evaluate the impact of news on crypto price movements.
As we know, news events such as regulatory changes, technological advancements, and major partnerships can have a significant impact on the price of cryptocurrencies. By analyzing the data… See the full description on the dataset page: https://huggingface.co/datasets/SahandNZ/cryptonews-articles-with-price-momentum-labels.pcqnp-finite-shot-reliability-artifact
PC-QNP Finite-Shot Reliability Artifact
This anonymous review artifact supports a NeurIPS 2026 Evaluations & Datasets submission on physics-conformal reliability evaluation for finite-shot quantum-process surrogate models. The asset is intended for reviewer inspection and reviewer-level reproduction of the reported aggregate tables, nested residual-repair calculation, and IBM stochastic Pauli-channel protocol validation.
Contents
code_snapshot/: cleaned source-code… See the full description on the dataset page: https://huggingface.co/datasets/anon-pcqnp-ed26/pcqnp-finite-shot-reliability-artifact.spotify-artistsai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.summarization-allegro-articlesArticle-Bias-Predictionclimate-news-articles
🌍 Jeu de données d'articles de presse française labellisés comme traitant ou non des sujets liés au climat
🇬🇧 / 🇺🇸 : as this data set is based only on French data, all explanations are written in French in this repository. The goal of the dataset is to train a model to classify titles of French newspapers in two categories : if it's about climate or not.
🗺️ Le contexte
Ce jeu de données de classification de titres d'article de presse française a été réalisé pour… See the full description on the dataset page: https://huggingface.co/datasets/pierre-loic/climate-news-articles.SpatioTemporal-News-Corpus
Spatiotemporal News Dataset
Overview
This dataset contains approximately 1.2 million English-language news headlines and articles sourced from major outlets in the
United States, United Kingdom, Canada, and Australia.
Each entry is annotated with spatial (country of origin) and temporal (date of publication) contexts, designed for training spatiotemporal-aware sentence embeddings,
specifically our Space-Time-MiniLM-v0 model.
The dataset covers the period from January… See the full description on the dataset page: https://huggingface.co/datasets/Artur-B/SpatioTemporal-News-Corpus.book-recommender-artifactsNews_Articles_Categorization
Dataset Card for News_Articles_Categorization
Dataset Description
3722 News Articles classified into different categories namely: World, Politics, Tech, Entertainment, Sport, Business, Health, and Science
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely Text and Category.
The Text column consists of the news article and the Category column consists of the class each article belongs to… See the full description on the dataset page: https://huggingface.co/datasets/valurank/News_Articles_Categorization.ArtVision
README — ArtVision: Dataset per la valutazione delle competenze visivo-interpretative in dominio storico-artistico
Descrizione generale
Il dataset ArtVision è una raccolta di 250 task, organizzati in otto categorie, in cui immagini di repertori storico artisti realizzati tra il 1750 e il 1985, sono utilizzate come base per la costruzione di richieste a modelli multimodali. Il dataset permette di sviluppare un veloce test di valutazione di un modello multimodale… See the full description on the dataset page: https://huggingface.co/datasets/paolodegasperis/ArtVision.synthetic-fine-arts
🎨 Synthetic Fine Arts (Challenge, Solution) Dataset
🖼️ 225,000 synthetic (artistic challenge, proposed solution) pairs spanning Visual Arts, Performing Arts, Musical Arts, Literary Arts, Digital Arts, Art History, and Art Theory, with metadata fields that mirror the shape of a curation workflow.
⚠️ Disclaimer: All text is synthetically generated and should not be relied on for artistic, historical, or technical accuracy. The VerificationMethod, ReferenceMaterial, and… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-fine-arts.human-essays
Dataset Statistics
Source Dataset
Row Count
Description
ASAP2
24,7k
Automated Student Assessment Prize dataset with scored essays
PERSUADE
15,6k
Discourse-annotated persuasive essays with effectiveness ratings
IvyPanda Essays
128k
Academic essays from IvyPanda educational platform
CNN_News_Articles_2011-2022
CNN News Articles 2011-2022 Dataset
Introduction
This dataset contains CNN News Articles from 2011 to 2022 after basic cleaning. The dataset includes the following information:
Category
Full text
The data was downloaded from Kaggle at this URL: https://www.kaggle.com/datasets/hadasu92/cnn-articles-after-basic-cleaning. The dataset was split into two sets:
Train set with 32,218 examples
Test set with 5,686 examples
Usage
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/CNN_News_Articles_2011-2022.AI-ArtTools-Pack
AI ArtTools Pack v1.0 — 372 Styles / 23 Categories
Developer and artist utility pack for Stable Diffusion XL.
Not for generating pretty pictures — for generating usable production assets.
Compatible with Style Grid Organizer extension.
Contents
372 prompt styles across 23 categories
CSV format (Forge/A1111 compatible)
Covers the full production pipeline from rough concept to final asset
Categories
Category
Count
Purpose
ASSET
41
Weapons, props, UI… See the full description on the dataset page: https://huggingface.co/datasets/Kazzze/AI-ArtTools-Pack.
