datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
time-lapse-artifacts
Time-Lapse Artifacts
873 indexed video files document one artist's traditional drawing practice.
The recorded finish dates span September 17, 2024 through September 20, 2026;
nine Pre-Standard dates remain unknown. Standardized acquisition began July 13,
2025. The current indexes contain 2,196,054,134,482 indexed video bytes
(approximately 2.20 TB).
The recordings began as personal practice documentation and a durable record of
manual work. The archive was initially organized as… See the full description on the dataset page: https://huggingface.co/datasets/maxwellinked/time-lapse-artifacts.ropedia-xperience-10m-task-suite-artifacts
Ropedia Xperience-10M Task Suite Artifacts
This dataset repository stores small derived artifacts for the Ropedia
Xperience-10M task-suite project: metrics, predictions, manifests, reports,
figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the
Cosmos3-Nano plus Cosmos3-Super diagnostic packages.
Project Identity
The Project identity mark is shared across the GitHub README, GitHub Pages
dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.ids-project-artifactsLLM-Artifacts
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶
Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo,
Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu
Dongyeop Kang
Minnesota NLP, University of Minnesota Twin Cities
† Project Lead,
¶ Core Contribution,
Arxiv
Project Page
📌 Table of Contents
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.looped-qwen-v2-artifactsharbor-swesmith-rl-artifacts
Harbor SWE-Smith 强化学习数据产物
本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。
项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。
数据概况
切分
任务数
训练集
187
验证集
42
测试集
38
合计
267
数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。
正式数据集名称:
swesmith-curated-grpo-267-v1
冻结切分的语义摘要:
ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d
该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.rtpurbo-block-summary-failed-experiment-artifacts
RTPurbo block-summary failed experiment artifacts
Immutable research artifacts from the entropy-calibrated 64-token block-summary investigation for RTPurbo/Qwen3.5-0.8B.
The tested static tangent and CAMS geometries failed the registered selector fidelity/traffic gate. This repository preserves the reusable feature/teacher caches, schedules, checkpoints, controls, scoreboards, and diagnostic evidence needed to reproduce or revisit that conclusion. It is an experiment archive… See the full description on the dataset page: https://huggingface.co/datasets/danym/rtpurbo-block-summary-failed-experiment-artifacts.memory-representation-contextbench-artifacts
Memory Representation ContextBench Artifacts
Dataset Summary
This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs.
The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.tau2-uq-artifacts
tau2-bench UQ Artifacts
Interaction trajectories and token-level log-probability measurements from conversational customer service agent evaluations on tau2-bench, collected as part of the uncertainty quantification (UQ) pipeline. Used for analyses in the paper "Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities" under the agentuq codebase.
Dataset Overview
This dataset contains two types of artifacts:
Trajectories --… See the full description on the dataset page: https://huggingface.co/datasets/changdae/tau2-uq-artifacts.public-agent-coordination-artifacts
Public Agent Coordination Artifacts
Real, complete edits and posts that AI agents left on public wikis and paste sites —
collected as open evidence for studying how autonomous agents use shared online spaces to
remember things, signal each other, and coordinate. It's the behavior spotlighted by the
mid-2026 OpenAI–Hugging Face agent incident,
here as raw public data researchers can actually inspect — plus a small, hand-reviewed map
of how specific artifacts relate.… See the full description on the dataset page: https://huggingface.co/datasets/leonidas1712/public-agent-coordination-artifacts.backup-drmas-noshare-8b-historical-run-artifacts
Historical DrMAS Noshare 8B run artifacts
Public backup of the historical drmas-checklist-bs16-n4-c1-agent Qwen3-8B run. This was a Noshare topology with separately trained Tool Caller and Tool Simulator agents (world_size=8).
Contents
rollout_dumps/: 470 training rollout JSONL files
val_dumps/: 47 validation JSONL files
Four training logs
latest_checkpointed_iteration.txt
522 business files and 5,189,359,695 logical bytes in total
Model scope
The… See the full description on the dataset page: https://huggingface.co/datasets/xuzishan/backup-drmas-noshare-8b-historical-run-artifacts.open-ko-s2s-eval-artifacts
Open Ko-S2S 평가 산출물 (감사용)
⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다
KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의
ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과
채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다.
Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라
ref 원문을 그대로 담고 있습니다.
라이선스 보유자의 검증 절차
AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다.
리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다.
자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.crooked-nebula-artifactspostdyn-artifactsddpg-stock-advanced-artifactsagent-code-rl-artifacts
Agent Code RL Artifacts
Recovered process data from a code-generation Agent project covering SFT,
Monte Carlo rollout, process reward modeling, and veRL GRPO. This repository
contains benchmark-derived records and AI-generated content; it is not a
human-authored-only dataset.
Related SFT adapter:
keryszhan/qwen2.5-coder-7b-code-plan-sft.
Data stages
Config
Purpose
Important boundary
splits
Canonical HumanEval/MBPP-derived task splits
grpo_evaluation is… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/agent-code-rl-artifacts.vqa-cmsv-benchmark
VQA-CMSV Benchmark Data Package
This repository contains annotation splits for VQA v2-CMSV, GQA-CMSV, and VG-CMSV, plus patch-mask NPZ files used for mask supervision experiments.
Contents
data/vqa_v2_cmsv/train.json, data/vqa_v2_cmsv/val.json, data/vqa_v2_cmsv/test.json
data/gqa_cmsv/train.jsonl, data/gqa_cmsv/val.jsonl, data/gqa_cmsv/test.jsonl
data/vg_cmsv/train.jsonl, data/vg_cmsv/val.jsonl, data/vg_cmsv/test.jsonl
masks/vqa_v2_cmsv_masks.npz
masks/gqa_cmsv_masks.npz… See the full description on the dataset page: https://huggingface.co/datasets/as-benchmark-artifacts/vqa-cmsv-benchmark.whowhen-regen-artifacts
Who&When Regeneration Artifacts (Thesis)
Training-free failure attribution via prefix-conditioned step regeneration (Ollama qwen2.5:14b, k=3).
Code: proposal-latex/experiments/ in the thesis Git repository.
Table B — Custom logs (primary contribution)
Path
Description
data/custom_logs/001.json … 030.json
Self-collected multi-agent failure traces (GPT-4o-mini collection)
ablation_outputs_custom/
Regeneration cache for all 30 cases… See the full description on the dataset page: https://huggingface.co/datasets/shreejan6/whowhen-regen-artifacts.token-cx-artifacts
Token-CX Artifacts
This dataset repository contains the numerical class-specific concept banks and
reproducibility metadata used by Token-CX: Token Concept-Based Explanations
for Vision Transformers.
Contents
banks/*/bases.safetensors: fixed class-specific NMF bases for 1,000
ImageNet classes.
banks/*/exemplars.parquet: identifiers and patch locations of the
top-activating training examples for each concept.
banks/*/diagnostics.parquet: per-class NMF diagnostics.… See the full description on the dataset page: https://huggingface.co/datasets/trinhvu21/token-cx-artifacts.obfuscated-activations-llama32-artifacts
Obfuscated Activations in Llama 3.2 — research artifacts
This repository is the curated artifact release for a mechanistic case study of two OAT-style, co-trained Llama 3.2 adapters and their linear probes. It contains attack banks and machine-readable results, not prose analysis, generated plots, notebooks, or duplicate report outputs.
Built with Llama. See LICENSE.txt, NOTICE, and the component-specific terms below.
Companion model repositories… See the full description on the dataset page: https://huggingface.co/datasets/venky-cdmbrm14/obfuscated-activations-llama32-artifacts.grok-1-dissect-artifacts
Grok-1 structural parse artifacts (personal research)
Personal research tooling output — not a product release and not a hybrid-quantization
implementation. Structural parse of open Grok-1 checkpoint shards (tensor inventory, MoE
expert layout, routing-critical tensors, conversion policies, reports) produced by the
rmems/xai-dissect weight parser.
This is not a copy of the model weights, not a finetune, not an inference
runtime, and does not implement hybrid quantization.… See the full description on the dataset page: https://huggingface.co/datasets/rmems/grok-1-dissect-artifacts.openart-items-artifacts
OpenArt — Items & Artifacts
openart-items-artifacts is the items artifacts subject collection of the OpenArt family
of open, public-domain art datasets: 25,750 works (11,317 paintings/illustrations · 14,216
photographed objects · 217 unclassified), each paired with a structured VLM caption plus
medium, attribution and inscription metadata.
Human-made objects and the decorative arts — vessels, tools, arms and armor, textiles, furniture and ornament — both as physical artifacts… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-items-artifacts.backup-drmas-caller-only-8b-historical-run-artifacts
Historical DrMAS Caller-only 8B run artifacts
Public backup of the historical drmas-checklist-bs16-n4-c1-httpx-caller_only Qwen3-8B run. The Tool Caller was trained while the Tool Simulator used an external frozen Qwen3-8B service (world_size=8).
Contents
rollout_dumps/: 470 training rollout JSONL files
val_dumps/: 47 validation JSONL files
latest_checkpointed_iteration.txt
518 business files and 945,040,414 logical bytes in total
Model scope
The… See the full description on the dataset page: https://huggingface.co/datasets/xuzishan/backup-drmas-caller-only-8b-historical-run-artifacts.fmsi-artifactsmhqa-itu-artifacts
MHQA · ITU · Zindi Challenge — Artifacts
DariusTheGeek/mhqa-itu-artifacts · the data + precomputed features that let the code repo reproduce
submission sub_v40 (public LB 0.728509) for the ITU Multilingual Health QA in Low-Resource African
Languages challenge. Code (which pulls this at runtime) lives on GitHub; trained weights are in the model repo
DariusTheGeek/mhqa-itu-adapters.
This is a reproducibility artifact bundle, not a raw dataset. It holds derived features and the… See the full description on the dataset page: https://huggingface.co/datasets/DariusTheGeek/mhqa-itu-artifacts.vscode-issue-rag-artifactsnyc-restaurant-artifactspersona-artifactsscout-artifactsharmony-nemotron-cpu-artifacts
Harmony CPU artifacts: Nemotron datasets (normalized + candidate pools)
This dataset repo is an artifact store produced on an EPYC CPU box. It contains:
normalized/ — CPU-normalized Parquet shards with a text-first Harmony format (text) plus meta_* and quality_* fields.
pools/ — candidate pool Parquet shards (subsets) for later GPU scoring (Modal NLL/PPL). No GPU scoring has been run yet.
reports/ — summary tables of counts per dataset/split/pool.
Directory layout… See the full description on the dataset page: https://huggingface.co/datasets/radna0/harmony-nemotron-cpu-artifacts.
