datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AA-Omniscience-Public
Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies.
Leaderboard and detailed results
Paper
Introduction
We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.ledger-long-context-KPI-QA
LEDGER — Long-Context KPI Question Answering & Page Retrieval
This dataset is part of the LEDGER (Long-context Evaluation of Documents for
Grounded Extraction and Retrieval) benchmark.
It supports two of the three LEDGER tasks:
Page-level KPI retrieval — given a natural-language question about a financial
KPI and the corresponding annual report, retrieve the relevant page(s). Each row
includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.ITBench-AA
ITBench-AA
Artificial Analysis' release of the public scenarios from
IBM's ITBench benchmark, used for
the ITBench-AA leaderboard.
This repo currently contains the SRE subset (sre config). Each row is a
Kubernetes incident scenario with its expected contributing-factor entities. An
agent under evaluation is given access to an offline snapshot of the affected
cluster (alerts, events, traces, topology) and must identify the entity
(Deployment, Pod, ConfigMap, etc.) responsible for… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA.gspc-art5
GSPC — art5 safeguard bank (Art5Bench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the art5-safeguard row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=art5-safeguard (family, kind, status and n are on that row… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-art5.ArtiMuse-10K
ArtiMuse:
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
[🌐 Project Page]
[🚀 Online Demo]
[💻 Code]
[📄 Paper]
[[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]]
🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/Thunderbolt215215/ArtiMuse-10K.aim-technical-articles
Analytics India Magazine Technical Articles Dataset 🚀
Dataset Description
This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies.
✨ Dataset Highlights
📚 Comprehensive Coverage: Latest AI models, frameworks, and tools
🔬 Technical Depth: Extracted keywords and complexity scoring
🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.TOMG-Bench
New Updates:
We fixed small bugs and upload new files in OpenMolIns on 3.18, 2025.
TOMG-Bench: Evaluating LLMs on Text-based Open Molecule Generation
Data Source: Data
Home Page: Home
Github Page: Github
PaperWithCode Page: PWC
Training:
Please see ./OpenMolIns/
We have 5 different variants, from light to xlarge.
Evaluation:
Please see ./benchmarks/, and combine the codes from our repository TOMG-Bench
citation
If this… See the full description on the dataset page: https://huggingface.co/datasets/Duke-de-Artois/TOMG-Bench.submission-eval-artifacts
NeurIPS ED 2026 Anonymous Evaluation Artifacts
This dataset repo contains sanitized evaluation artifacts for an anonymous NeurIPS ED 2026 submission. It is metadata-focused: normalized benchmark JSON, selected small paper-facing summaries, reviewer indexes, and manifests.
Checkpoint artifacts are referenced through neurips-ed2026-anon-checkpoints/submission-checkpoints. This dataset repo does not contain model checkpoints or model weights.
Anonymous code artifact:… See the full description on the dataset page: https://huggingface.co/datasets/neurips-ed2026-anon-checkpoints/submission-eval-artifacts.IndustryInstruction_Artificial-Intelligence
IndustryInstruction: Artificial Intelligence
This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.arlis-armenia-legal-dataset
Complete Armenian Legislation & Legal Corpus (ARLIS 1990–2026)
This repository contains the complete, unified, and cleaned digital corpus of all official legislative, normative, and regulatory acts of the Republic of Armenia published in the Armenian Legal Information System (ARLIS — arlis.am) from 1990 to September 2026.
📊 Dataset Summary
Total Legal Acts: 188,894 acts
Historical base (1990 – April 2023): 154,934 acts (data/train-historical-part01.parquet –… See the full description on the dataset page: https://huggingface.co/datasets/ArthurYeghinyan/arlis-armenia-legal-dataset.ArtiFact
ArtiFact
ArtiFact is a large-scale multimodal benchmark of museum artwork records with aligned images and structured metadata. It is designed for evaluating metadata extraction, error detection, semantic querying, and multimodal reasoning over cultural-heritage collections.
The dataset combines records from the Rijksmuseum, the Metropolitan Museum of Art (Met), and the Art Institute of Chicago (AIC), with normalized fields for artists, dates, materials, techniques, dimensions… See the full description on the dataset page: https://huggingface.co/datasets/deem-data/ArtiFact.forecastgen-artifacts
Forecast-Generalization: raw evaluation outputs across 38 reasoning models
Complete generation-level outputs, per-seed scores and analysis artifacts from a
study of how well benchmark performance forecasts generalization to held-out
reasoning tasks.
Most released evaluations report only aggregate accuracy. This release keeps the
raw per-problem, per-seed generations, so item-level analyses can be redone
without re-running any inference.
What is here
38 models… See the full description on the dataset page: https://huggingface.co/datasets/dvader13/forecastgen-artifacts.Argimi-Legal-French-Jurisprudence
The ArGiMi French Jurisprudence Dataset
This dataset contains a comprehensive collection of French case law, sourced from the official archives of French jurisprudence. It is divided into three distinct subdivisions: Constitutional ("constit"), Administrative ("cetat"), and Judiciary ("juri").
This dataset was created for the ArGiMi project, an open-source initiative dedicated to promoting open data and knowledge sharing. The project is a collaborative effort between Giskard… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Legal-French-Jurisprudence.that-backpacker-article-corpus
That Backpacker Article Corpus
This dataset contains a structured corpus of long-form travel articles published on ThatBackpacker.com, authored primarily by Audrey Bergner as part of the Samuel & Audrey Media Network.
The corpus includes 323 article records covering destination guides, multi-day itineraries, hiking, food travel, cultural experiences, city guides, transportation, accommodations, and practical travel planning.
It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/that-backpacker-article-corpus.magic-video-artifacts
MAGIC-Video — Preprocessing Artifacts
This dataset hosts the exact preprocessing artifacts used in the paper
"Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning"
(MAGIC-Video, arXiv:2605.08271).
Why release these?
The paper's preprocessing pipeline calls LLMs through OpenRouter (translation, OpenIE, semantic
consolidation, narrative chain distillation). Those calls cost money, take hours per subject,
and are non-deterministic — re-running… See the full description on the dataset page: https://huggingface.co/datasets/jiazhengli7/magic-video-artifacts.constitution-2.0
Rules for agents that remember
If you are building multi-agent systems, you already have orchestration, tools, and
a memory store. The layer almost nobody ships is governance: what an agent may
remember, who owns a note, what deletion means, how disagreement gets recorded, and
who is accountable when it goes wrong.
This dataset is one complete answer to that, in the public domain. 47 articles,
one row each, byte-exact and hash-pinned. Written for work between humans and AI… See the full description on the dataset page: https://huggingface.co/datasets/article11/constitution-2.0.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas.vectrix-art-e
Vectrix ART-E: Synthetic Email Agent Benchmark
A fully synthetic email corpus and task dataset for training and evaluating email search agents, built as a drop-in replacement for the Enron corpus used in OpenPipe's ART-E benchmark.
Key Result
A Qwen3.5-35B-A3B fine-tuned via GRPO on this synthetic dataset beats o3 on real Enron emails (86% vs 85%) — despite never seeing a single real email during training.
Dataset Contents
The dataset is available in two… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/vectrix-art-e.devcenter-articles
Overview
This dataset consists of ~600 articles from the MongoDB Developer Center.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the article. This value is devcenter for the entire dataset.
url: Link to the article
action: Action taken on the article. This value is created for the entire dataset.
body: Content of the article in Markdown format
format: Format of the content. This value is md for all articles.
metadata: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.korean-law-articles
Korean Law Articles — 한국 현행 법령 본문 전수 데이터셋
법제처 국가법령정보 OPEN API (lawService.do) 로 수집한 대한민국 현행 법령 5,583건의 본문 전수 데이터셋입니다. (전체 5,584건 중 1건은 출처 서버에서 본문 미제공)
🎯 라이브 Q&A 챗봇
▶ Korean Law Q&A (Gradio Space)
본 데이터셋 위에 RAG 챗봇이 가동 중입니다 — 자연어 질문에 답변하고 인용한 법령을 클릭 가능 링크로 제시합니다.
Retrieval: sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 시맨틱 + BM25 hybrid
LLM: Llama 3.3 70B (Groq, free tier)
인용 강제 + hallucination 방지: 데이터셋에 없는 내용은 "확인되지 않습니다" 반환
규모… See the full description on the dataset page: https://huggingface.co/datasets/Wookhyeon/korean-law-articles.visible-accumulator
The visible accumulator
Chain-of-thought trajectories of open-weight reasoning models on multiple-choice science questions,
with the model's two-choice decision variable read at every checkpoint of the reasoning, under
hints that point at a wrong option, sycophantic pushback, thinking versus non-thinking mode and
sampling temperature. The decision variable is the exact log-odds between the hinted and the
correct answer letter when the reasoning is force-closed after the first t… See the full description on the dataset page: https://huggingface.co/datasets/ArthT/visible-accumulator.GURU-MATH-CL
GURU-MATH-CL: Difficulty-Ordered Mathematical Reasoning Dataset
This dataset is a curriculum-learning variant of the guru-RL-92k mathematics subset, specifically prepared for reinforcement learning post-training of large language models.
📊 Dataset Overview
GURU-MATH-CL contains 54,404 mathematical reasoning problems ordered by difficulty, designed to support curriculum learning strategies in RL-based model training.
Training Set: 53,904 problems (ordered easy → hard)… See the full description on the dataset page: https://huggingface.co/datasets/Artemis0430/GURU-MATH-CL.synthetic-fine-arts
🎨 Synthetic Fine Arts (Challenge, Solution) Dataset
🖼️ 225,000 synthetic (artistic challenge, proposed solution) pairs spanning Visual Arts, Performing Arts, Musical Arts, Literary Arts, Digital Arts, Art History, and Art Theory, with metadata fields that mirror the shape of a curation workflow.
⚠️ Disclaimer: All text is synthetically generated and should not be relied on for artistic, historical, or technical accuracy. The VerificationMethod, ReferenceMaterial, and… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-fine-arts.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas.picture-perfect-portfolios-article-corpus
Picture Perfect Portfolios Article Corpus
This dataset contains a structured corpus of long-form finance and investing articles published on PicturePerfectPortfolios.com.
The corpus includes 448 article records covering portfolio construction, asset allocation, capital efficiency, return stacking, managed futures, trend following, risk management, ETFs, systematic investing, alternative strategies, and DIY investor education.
It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/picture-perfect-portfolios-article-corpus.artha-benchmark
The Artha Personal-Finance Reasoning Benchmark
A reproducible, model-agnostic benchmark for evaluating LLM-based
personal-finance agents on two axes that existing financial-QA benchmarks
do not jointly exercise:
Answer quality — aggregative and multi-hop reasoning over a user's
full transaction ledger (banking and investments), scored with a
four-dimension rubric by a three-model-family LLM judge panel.
Grounding / hallucination resistance — queries that embed a
false premise… See the full description on the dataset page: https://huggingface.co/datasets/Tej-Katika/artha-benchmark.ArtiMuse-10K
ArtiMuse:
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
[🌐 Project Page]
[🚀 Online Demo]
[💻 Code]
[📄 Paper]
[[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]]
🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/to-hope-we-bound/ArtiMuse-10K.islamic-articles-corpus
☪ Islamic Articles Corpus - English RAG Dataset
Dataset Description
Islamic Articles Corpus is a curated English-language RAG corpus containing 33 articles covering Muslim travel guides, mosque visits, halal food, prayer room directories, and Islamic community documentation. Every article preserves complete full-text content with all 608 embedded image references. Content focuses heavily on Singapore, Iran, Japan, Oman, and Qatar mosque and travel documentation.… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/islamic-articles-corpus.che-argentina-travel-article-corpus
Che Argentina Travel Article Corpus
This dataset contains a structured corpus of long-form Argentina travel articles published on CheArgentinaTravel.com by the Samuel & Audrey Media Network.
The corpus includes 88 article records covering Argentina travel guides, itineraries, cultural experiences, regional food, transportation, accommodations, local logistics, and destination planning. It includes coverage of major areas such as Buenos Aires and Patagonia, along with regional… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/che-argentina-travel-article-corpus.
