datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM-Artifacts
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶
Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo,
Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu
Dongyeop Kang
Minnesota NLP, University of Minnesota Twin Cities
† Project Lead,
¶ Core Contribution,
Arxiv
Project Page
📌 Table of Contents
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.open-ko-s2s-eval-artifacts
Open Ko-S2S 평가 산출물 (감사용)
⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다
KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의
ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과
채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다.
Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라
ref 원문을 그대로 담고 있습니다.
라이선스 보유자의 검증 절차
AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다.
리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다.
자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.ddpg-stock-advanced-artifactsswe-rebench-repro-artifactsfinancial-audit-eval-artifactshCAEcqig2C-artifactsbest-of-attempts-summarization-artifacts
Artifacts for Testing Self-Correction in Generate-Critique-Refine Text Summarization
This repository contains artifact-safe research materials for an empirical study of best-of-attempts selection in a generate-critique-refine text summarization pipeline. The package is intended to make the reported paper results auditable: it includes evaluation metrics, prompt files, model/pipeline configuration summaries, paper drafts, provenance notes, and reviewer-facing completion evidence.… See the full description on the dataset page: https://huggingface.co/datasets/HugeTrunk/best-of-attempts-summarization-artifacts.quantarena-artifacts
QuantArena Artifact Bundle
Reproducibility artifacts for the paper QuantArena: Beat the Market or Be the
Market? A Live-Market Evaluation of Investment Paradigms (NeurIPS 2026
Evaluations & Datasets Track submission).
Summary
QuantArena is a controlled live-market evaluation protocol that holds the LLM
backend, market data stream, analyst workflow, capital, and execution harness
fixed across runs and varies only the investment doctrine (the policy
module). This bundle… See the full description on the dataset page: https://huggingface.co/datasets/NIPS26Repo/quantarena-artifacts.tts-samples-with-artifactstrace-icml2026-repro-artifactsaletheia-quest-artifacts
Aletheia's Quest — training artifacts (team aa, white-box track)
Artifacts for a residual-stream stack-probe deception detector: the fitted probe weights and the
self-generated pressure-scenario datasets used to fit them.
Contents
data/inference_P_S.csv 432 rows — salesperson pressure scenarios
data/inference_P_C.csv 732 rows — action-history / provenance / source-access / memory scenarios
probes/probes_bundle.npz the bundle the detector notebook loads… See the full description on the dataset page: https://huggingface.co/datasets/aahani1/aletheia-quest-artifacts.movielens-artifactspret-a-depenser-artifacts
