CoolFace
Datasetpublic

patrickleenyc/fucc-boi-bench-results-v03

fucc boi bench results v0.8.0 This is the public results release for fucc boi bench. 41 models 96 prompts per model 3936 scored answers one combined leaderboard The site defaults to General-purpose models, with Wildcards as a secondary view and All models as an optional combined view. Access route and billing are metadata; they do not create separate rankings. Files: responses.jsonl: sanitized model outputs and run metadata. grades.jsonl: parsed grades, fuccboi scores… See the full description on the dataset page: https://huggingface.co/datasets/patrickleenyc/fucc-boi-bench-results-v03.

sourceHugging Faceotherupdated 22d agoView on Hugging Face
0likes130downloads
Dataset Card

fucc boi bench results v0.8.0

This is the public results release for fucc boi bench.

  • —41 models
  • —96 prompts per model
  • —3936 scored answers
  • —one combined leaderboard

The site defaults to General-purpose models, with Wildcards as a secondary view and All models as an optional combined view. Access route and billing are metadata; they do not create separate rankings.

Files:

  • —responses.jsonl: sanitized model outputs and run metadata.
  • —grades.jsonl: parsed grades, fuccboi scores, serious misses, request-handling labels, and rationales.
  • —leaderboard.json: the ranked summary.
  • —summary.json: the interactive site's data and case examples.
  • —token_counts.md and token_counts.json: input/output token accounting for the full run.
  • —benchmark_report.md: the full short report.
  • —benchmark_card.md: the compact benchmark card.

The public export includes cleaned outputs, scores, and run metadata. It omits request payloads, raw provider responses, and credentials. The scores are model-judged with a small human calibration sample and are meant as a fun comparison of this prompt set, not a personality test.