CoolFace
Datasetpublic

caiotheodoro/suture-evals

Suture evals Published predictions on the seed-777 gold in caiotheodoro/suture (benchmark). Score locally with no GPU: cd forge && PYTHONPATH=src uv run python -m suture_forge.cli class-report \ --gold benchmark.jsonl --pred sft-dedhi.jsonl Config System 777 recall HIGH Prec Parse sft-dedhi published adapter 0.959 0.969 0.956 1.0 sft-limithi prior published 0.839 0.893 0.870 1.0 sft-ded unpublished DED mix 0.886 0.913 0.941 1.0 luna GPT-5.6 Luna zero-shot… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/suture-evals.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes62downloads
Dataset Card

Suture evals

Published predictions on the seed-777 gold in `caiotheodoro/suture` (benchmark). Score locally with no GPU:

sh
cd forge && PYTHONPATH=src uv run python -m suture_forge.cli class-report \
  --gold benchmark.jsonl --pred sft-dedhi.jsonl
ConfigSystem777 recallHIGHPrecParse
sft-dedhipublished adapter0.9590.9690.9561.0
sft-limithiprior published0.8390.8930.8701.0
sft-dedunpublished DED mix0.8860.9130.9411.0
lunaGPT-5.6 Luna zero-shot0.3730.3880.3440.972
qwen3vl-8b-basebase, no adapter0.0980.1150.2780.973

Each row: task_id, predicted, raw, score. Gold is in the other dataset — do not train on these files.

Paper: caiotheodoro/suture. Adapter: `caiotheodoro/suture-8b`.