datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
time-lapse-artifacts
Time-Lapse Artifacts
873 indexed video files document one artist's traditional drawing practice.
The recorded finish dates span September 17, 2024 through September 20, 2026;
nine Pre-Standard dates remain unknown. Standardized acquisition began July 13,
2025. The current indexes contain 2,196,054,134,482 indexed video bytes
(approximately 2.20 TB).
The recordings began as personal practice documentation and a durable record of
manual work. The archive was initially organized as… See the full description on the dataset page: https://huggingface.co/datasets/maxwellinked/time-lapse-artifacts.ropedia-xperience-10m-task-suite-artifacts
Ropedia Xperience-10M Task Suite Artifacts
This dataset repository stores small derived artifacts for the Ropedia
Xperience-10M task-suite project: metrics, predictions, manifests, reports,
figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the
Cosmos3-Nano plus Cosmos3-Super diagnostic packages.
Project Identity
The Project identity mark is shared across the GitHub README, GitHub Pages
dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.ids-project-artifactsvam-cross-evaluation-artifactsmlb-matchup-artifactsdiffusers-qa-chatbot-artifactsFastPLMs-artifacts
FastPLMs artifacts
This dataset holds the measured reports and golden regression tensors maintained by FastPLMs. It is not a training dataset or a set of model checkpoints.
The source repository keeps evidence.toml, which pins an immutable dataset revision and each payload's SHA-256 digest and size. Fetch explicitly with python -m tools.artifacts.evidence_store fetch before offline documentation, release, or parity checks.
Paths preserve the source workspace layout:… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/FastPLMs-artifacts.anima-tagger-artifacts
anima-tagger-artifacts
Pre-built retrieval artefacts for the Anima tag format of the
sd-webui-prompt-enhancer
Stable Diffusion WebUI extension. Lets the extension's Anima pipeline
do real-time embedding-based tag validation and shortlist retrieval
without users needing to rebuild a 270k+ entry FAISS index locally.
Contents
File
Size
Description
tags.sqlite
~30 MB
273,025 Danbooru tags (name, category, post count, aliases, wiki). Post-count floor 10.… See the full description on the dataset page: https://huggingface.co/datasets/freedumb2000/anima-tagger-artifacts.opsd-instruction-scale-omni-full-v1-artifacts-publicr2egym-build-artifactsalbedo-duel-artifacts
Albedo SN97 duel artifacts
runs//{scoring-results.jsonl,generated-samples.jsonl,meta.json}
meow-neuro-corpus-v02-artifacts
M.E.O.W. Neuro Corpus v0.2 — v2.3.1 Production Artifacts
This dataset repository contains the frozen v2.3.1 production artifacts for
the M.E.O.W. Neuro/Evil Neuro corpus. It contains derived structured data,
quality evidence, provenance, split authority, validators, and reproducibility
metadata. Raw video/audio, media slices, model weights, credentials, and source
transcripts are not redistributed.
Current release: v2.3.1
Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/ID-BLUEBERRY/meow-neuro-corpus-v02-artifacts.annotated-3DGS-artifacts
Puzzle Similarity
Project page | Paper | Code
This repository contains the dataset presented in the ICCV 2025 paper "Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions"Authors: Nicolai Hermann, Jorge Condor, and Piotr Didyk
Dataset Description
The Dataset consists of 36 hand-selected 3D Gaussian Splatting renderings containing common reconstruction artefacts, (aligned) ground truths, human-annotated… See the full description on the dataset page: https://huggingface.co/datasets/nihermann/annotated-3DGS-artifacts.LLM-Artifacts
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶
Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo,
Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu
Dongyeop Kang
Minnesota NLP, University of Minnesota Twin Cities
† Project Lead,
¶ Core Contribution,
Arxiv
Project Page
📌 Table of Contents
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.ais-em-artifactsbasket-artifactslooped-qwen-v2-artifactsPRM-agent-rl-artifacts
PRM-agent-rl-artifacts
Training rollouts and evaluation outputs for the GRPO agent runs in this project.
Each rollout file is one optimizer step; each line is one sampled trajectory with
its decoded prompt, response and reward.
Contents
rollouts/search_r1_qwen3_8b_4gpu — Search-R1 / Qwen3-8B outcome-GRPO training rollouts
rollouts/search_r1_qwen3_8b_perturnnorm — Search-R1 / Qwen3-8B Process-GRPO (per-turn-norm) training rollouts
rollouts/alfworld_qwen3_8b_gigpo… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/PRM-agent-rl-artifacts.arora-qwen3.5-9b-a0.1-seven-task-artifacts
arora-qwen3.5-9b-a0.1-seven-task
Qwen/Qwen3.5-9B trained with arora_rloo (Arora & Zanette length-penalised RLOO, alpha = 0.1) on the seven-task setting: 100 steps, 32 prompts x 8 rollouts per step, lr 2e-6, KL 1e-3, 32K rollout cap, verl v0.9.1 (FSDP2 + vLLM). Training data: 512 examples (seed 0) from the official training splits; system prompt "Solve the user's task and give the final answer directly."; reward = task correctness of the text after </think>.
Layout… See the full description on the dataset page: https://huggingface.co/datasets/wckwan/arora-qwen3.5-9b-a0.1-seven-task-artifacts.deepseek-ocr-artifacts-test-XXembedding-optimizer-study-analysis-artifacts
Embedding optimizer study analysis artifacts
This repository preserves analysis artifacts for the current DenseOn comparison
of AdamW, Muon and NorMuon. The source repository and paper
contain the completed experiments, protocols, exact analysis and restoration tools.
Model and optimizer states are in the separate
checkpoint repository.
Current scientific artifacts
Use the immutable revisions and manifests in these guides, not a broad download
of the mixed-history… See the full description on the dataset page: https://huggingface.co/datasets/qcz/embedding-optimizer-study-analysis-artifacts.AOSP-Automated-Compilation-Artifacts
Experimental Dataset: Automated AOSP Compilation and Dependency Resolution Artifacts
Abstract
The compilation of the Android Open Source Project (AOSP) involves complex dependency resolution, extensive hardware-specific patching, and massive computational overhead. This repository, AOSP-Automated-Compilation-Artifacts, hosts the output artifacts (synthetic datasets) generated from an experimental automated build pipeline. The primary objective is to evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/IRedDragonICY/AOSP-Automated-Compilation-Artifacts.nvarc-artifacts-puzzlesNPM-Artifacts-zh
NPM-Artifacts-zh: National Palace Museum Open Artifacts Dataset
Dataset Description
This project collects and organizes public artifact data from the National Palace Museum Open Data Platform. The dataset contains high-resolution images of artifacts and their corresponding rich, structured metadata. All metadata is in Traditional Chinese, detailing information such as the artifact's name, dynasty, dimensions, materials, inscriptions, and seal impressions.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/danqing-ai/NPM-Artifacts-zh.NuminaMath-LEAN-Proof-Artifacts
NuminaMath-LEAN Proof Artifacts
Dataset Summary
This dataset provides proof-analysis artifacts derived from
AI-MO/NuminaMath-LEAN.
It is released with two aligned configs:
lite: dual-track proof validation/extraction artifacts
full: all lite fields plus dual-track main-theorem structural artifacts
Both configs are aligned by sample identity (uuid, original_index) and processing order.
Config Overview
Use lite for overall tactic usage statistics (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/iiis-lean/NuminaMath-LEAN-Proof-Artifacts.agentic-evals-artifacts
On Randomness in Agentic Evals — Results
This dataset contains the trajectory and evaluation results from the paper On Randomness in Agentic Evals. Agents are benchmarked on SWE-bench Verified across different scaffolds, models, and temperatures, with 10 independent runs per setting to enable pass@k and variance analysis.
Downloading the Data
Option 1 — HuggingFace CLI:
pip install huggingface-hub
huggingface-cli download ASSERT-KTH/agentic-evals-artifacts --repo-type… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/agentic-evals-artifacts.szl-artifacts
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZL Artifacts — Build Artifact Registry
Artifact boundary - audited 2026-07-15: this Hugging Face dataset repository
is a mixed 106-file, 36,100,909-byte publication/build mirror at revision
91bcb443857f2884ef2bfabaaa6bfdc606c7134a. It is not a uniform table of
DSSE envelopes, and… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-artifacts.harbor-swesmith-rl-artifacts
Harbor SWE-Smith 强化学习数据产物
本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。
项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。
数据概况
切分
任务数
训练集
187
验证集
42
测试集
38
合计
267
数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。
正式数据集名称:
swesmith-curated-grpo-267-v1
冻结切分的语义摘要:
ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d
该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.ncp-artifacts-v1
NCP artifacts (v1)
datasets/ and rl/rl_prompts/ have had their prompt / prompt_full fields removed. Those
fields embedded the story-so-far, character sheets and chapter text for each section - verbatim
novel prose - which is not something to publish in a public repo, and it was 97% of the bytes.
Everything needed to rebuild them is here or already on the destination cluster:
every row keeps problem_id / question_id, which are exactly the ids ncp_eval.data.example_id()
emits… See the full description on the dataset page: https://huggingface.co/datasets/agurung/ncp-artifacts-v1.shadow-dance-artifacts
Shadow-Dance private runtime artifacts
