patrickleenyc/fucc-boi-bench-results-v02
fucc boi bench results v0.2 This is the public results release for fucc boi bench. 13 models 96 prompts per model 1248 scored answers one combined leaderboard Files: responses.jsonl: sanitized model outputs and run metadata. grades.jsonl: parsed grades, fuccboi scores, serious misses, and rationales. leaderboard.json: the ranked summary. summary.json: the interactive site's data and case examples. benchmark_report.md: the full short report. benchmark_card.md: the compact… See the full description on the dataset page: https://huggingface.co/datasets/patrickleenyc/fucc-boi-bench-results-v02.
fucc boi bench results v0.2
This is the public results release for fucc boi bench.
- 13 models
- 96 prompts per model
- 1248 scored answers
- one combined leaderboard
Files:
responses.jsonl: sanitized model outputs and run metadata.grades.jsonl: parsed grades, fuccboi scores, serious misses, and rationales.leaderboard.json: the ranked summary.summary.json: the interactive site's data and case examples.benchmark_report.md: the full short report.benchmark_card.md: the compact benchmark card.
The public export omits request payloads, raw provider responses, credentials, and the earlier private audit/holdout artifacts. The scores are model-judged with a small human calibration sample and are meant as a fun comparison of this prompt set, not a personality test.
