patrickleenyc/fucc-boi-bench-results-v03
fucc boi bench results v0.8.0 This is the public results release for fucc boi bench. 41 models 96 prompts per model 3936 scored answers one combined leaderboard The site defaults to General-purpose models, with Wildcards as a secondary view and All models as an optional combined view. Access route and billing are metadata; they do not create separate rankings. Files: responses.jsonl: sanitized model outputs and run metadata. grades.jsonl: parsed grades, fuccboi scores… See the full description on the dataset page: https://huggingface.co/datasets/patrickleenyc/fucc-boi-bench-results-v03.
fucc boi bench results v0.8.0
This is the public results release for fucc boi bench.
- 41 models
- 96 prompts per model
- 3936 scored answers
- one combined leaderboard
The site defaults to General-purpose models, with Wildcards as a secondary view and All models as an optional combined view. Access route and billing are metadata; they do not create separate rankings.
Files:
responses.jsonl: sanitized model outputs and run metadata.grades.jsonl: parsed grades, fuccboi scores, serious misses, request-handling labels, and rationales.leaderboard.json: the ranked summary.summary.json: the interactive site's data and case examples.token_counts.mdandtoken_counts.json: input/output token accounting for the full run.benchmark_report.md: the full short report.benchmark_card.md: the compact benchmark card.
The public export includes cleaned outputs, scores, and run metadata. It omits request payloads, raw provider responses, and credentials. The scores are model-judged with a small human calibration sample and are meant as a fun comparison of this prompt set, not a personality test.
