CoolFace
Datasetpublic

lwaekfjlk/artifact-bench

ArtifactBench A heterogeneous graph of HuggingFace model / dataset / paper / codebase nodes (14,053) with observed (model, dataset, performance-metric) evaluation edges (51,337 relations), for benchmarking link prediction and attribute (metric-value) regression, plus an agent-based verification suite. License Released under the Open Database License (ODbL) v1.0 — see LICENSE or https://opendatacommons.org/licenses/odbl/1-0/. Share/modify/use freely with… See the full description on the dataset page: https://huggingface.co/datasets/lwaekfjlk/artifact-bench.

sourceHugging Faceodblupdated 4mo agoView on Hugging Face
9likes230downloads
42 commits on main
3aa029e4mo ago

docs: compact README + viewer config (eval_edges)

lwaekfjlk
b2836de4mo ago

feat: add data/eval_edges.jsonl for the dataset viewer

lwaekfjlk
13c4d2e4mo ago

docs: add ODbL v1.0 LICENSE

lwaekfjlk
167809b4mo ago

chore: relicense dataset under ODbL v1.0 (was CC-BY-4.0)

lwaekfjlk
e2b83375mo ago

Remove old top-level cell dirs (batch 4)

lwaekfjlk
df09d965mo ago

Remove old top-level cell dirs (batch 3)

lwaekfjlk
9c7b34b5mo ago

Remove old top-level cell dirs (batch 2)

lwaekfjlk
b1123d35mo ago

Remove old top-level cell dirs (batch 1)

lwaekfjlk
9b355b45mo ago

Add files using upload-large-folder tool

lwaekfjlk
b8569235mo ago

Add files using upload-large-folder tool

lwaekfjlk
ad7c08a5mo ago

Add files using upload-large-folder tool

lwaekfjlk
f69ff745mo ago

Add files using upload-large-folder tool

lwaekfjlk
a1fc5ff5mo ago

Add files using upload-large-folder tool

lwaekfjlk
3731a595mo ago

Add files using upload-large-folder tool

lwaekfjlk
d80970e5mo ago

Add files using upload-large-folder tool

lwaekfjlk
0c9367a5mo ago

Add files using upload-large-folder tool

lwaekfjlk
3aa13795mo ago

Move agent results under agent_runs/ (batch 5)

lwaekfjlk
a0a0a635mo ago

Move agent results under agent_runs/ (batch 4)

lwaekfjlk
5eb6b7e5mo ago

Move agent results under agent_runs/ (batch 3)

lwaekfjlk
167972b5mo ago

Move agent results under agent_runs/ (batch 2)

lwaekfjlk
829cf295mo ago

Move agent results under agent_runs/ (batch 1)

lwaekfjlk
55173c45mo ago

Drop agent_runs_summary.json and batch_summary.json — keep only bench.json

lwaekfjlk
0d019945mo ago

Rename all_results_summary.json -> agent_runs_summary.json (distinguish from bench spec)

lwaekfjlk
5e9fc395mo ago

Add verification_bench benchmark spec (263 ground-truth (model, dataset, metric) tuples)

lwaekfjlk
c400df15mo ago

Add verification_bench summary JSON (263 cells × model/dataset/metric/results)

lwaekfjlk
7c813af5mo ago

Move skills_multiagent contents directly under verification_bench/

lwaekfjlk
1ae11e35mo ago

Remove nested skills_multiagent_gpt-5.2_metadatatool/ (batch 3)

lwaekfjlk
80377bc5mo ago

Remove nested skills_multiagent_gpt-5.2_metadatatool/ (batch 2)

lwaekfjlk
848e64e5mo ago

Remove nested skills_multiagent_gpt-5.2_metadatatool/ (batch 1)

lwaekfjlk
c701eb65mo ago

Add verification_bench section to README

lwaekfjlk
ff920c15mo ago

Add verification_bench: skills_multiagent_gpt-5.2 (263 agent eval reproductions)

lwaekfjlk
0b4af4b5mo ago

Update NLI case study README: drop scripts section, note reproducibility caveat

lwaekfjlk
e78d9495mo ago

Remove partial run_eval.py (119 cells) and outdated shared/ loaders — keep only results+predictions

lwaekfjlk
9b4c3cd5mo ago

Add shared dataset_loaders + model_loaders for NLI case study

lwaekfjlk
d0a3fc15mo ago

Add NLI case study section

lwaekfjlk
2b3cfd15mo ago

Add NLI case study: 576 raw evals + aggregate + scripts + figures

lwaekfjlk
f9a77355mo ago

Upload README.md with huggingface_hub

lwaekfjlk
fda94155mo ago

Add full split

lwaekfjlk
b0a0c935mo ago

Upload README.md with huggingface_hub

lwaekfjlk
a5035e85mo ago

Add inductive split

lwaekfjlk
c8c53d15mo ago

Add transductive split

lwaekfjlk
54f1cd65mo ago

initial commit

lwaekfjlk