patrickleenyc/fucc-boi-bench-results-v02
fucc boi bench results v0.2 This is the public results release for fucc boi bench. 13 models 96 prompts per model 1248 scored answers one combined leaderboard Files: responses.jsonl: sanitized model outputs and run metadata. grades.jsonl: parsed grades, fuccboi scores, serious misses, and rationales. leaderboard.json: the ranked summary. summary.json: the interactive site's data and case examples. benchmark_report.md: the full short report. benchmark_card.md: the compact… See the full description on the dataset page: https://huggingface.co/datasets/patrickleenyc/fucc-boi-bench-results-v02.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face