lwaekfjlk/artifact-bench
ArtifactBench A heterogeneous graph of HuggingFace model / dataset / paper / codebase nodes (14,053) with observed (model, dataset, performance-metric) evaluation edges (51,337 relations), for benchmarking link prediction and attribute (metric-value) regression, plus an agent-based verification suite. License Released under the Open Database License (ODbL) v1.0 — see LICENSE or https://opendatacommons.org/licenses/odbl/1-0/. Share/modify/use freely with… See the full description on the dataset page: https://huggingface.co/datasets/lwaekfjlk/artifact-bench.
docs: compact README + viewer config (eval_edges)
feat: add data/eval_edges.jsonl for the dataset viewer
docs: add ODbL v1.0 LICENSE
chore: relicense dataset under ODbL v1.0 (was CC-BY-4.0)
Remove old top-level cell dirs (batch 4)
Remove old top-level cell dirs (batch 3)
Remove old top-level cell dirs (batch 2)
Remove old top-level cell dirs (batch 1)
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Move agent results under agent_runs/ (batch 5)
Move agent results under agent_runs/ (batch 4)
Move agent results under agent_runs/ (batch 3)
Move agent results under agent_runs/ (batch 2)
Move agent results under agent_runs/ (batch 1)
Drop agent_runs_summary.json and batch_summary.json — keep only bench.json
Rename all_results_summary.json -> agent_runs_summary.json (distinguish from bench spec)
Add verification_bench benchmark spec (263 ground-truth (model, dataset, metric) tuples)
Add verification_bench summary JSON (263 cells × model/dataset/metric/results)
Move skills_multiagent contents directly under verification_bench/
Remove nested skills_multiagent_gpt-5.2_metadatatool/ (batch 3)
Remove nested skills_multiagent_gpt-5.2_metadatatool/ (batch 2)
Remove nested skills_multiagent_gpt-5.2_metadatatool/ (batch 1)
Add verification_bench section to README
Add verification_bench: skills_multiagent_gpt-5.2 (263 agent eval reproductions)
Update NLI case study README: drop scripts section, note reproducibility caveat
Remove partial run_eval.py (119 cells) and outdated shared/ loaders — keep only results+predictions
Add shared dataset_loaders + model_loaders for NLI case study
Add NLI case study section
Add NLI case study: 576 raw evals + aggregate + scripts + figures
Upload README.md with huggingface_hub
Add full split
Upload README.md with huggingface_hub
Add inductive split
Add transductive split
initial commit
