CoolFace
Datasetpublic

patrickleenyc/fucc-boi-bench-results-v02

fucc boi bench results v0.2 This is the public results release for fucc boi bench. 13 models 96 prompts per model 1248 scored answers one combined leaderboard Files: responses.jsonl: sanitized model outputs and run metadata. grades.jsonl: parsed grades, fuccboi scores, serious misses, and rationales. leaderboard.json: the ranked summary. summary.json: the interactive site's data and case examples. benchmark_report.md: the full short report. benchmark_card.md: the compact… See the full description on the dataset page: https://huggingface.co/datasets/patrickleenyc/fucc-boi-bench-results-v02.

sourceHugging Faceotherupdated 25d agoView on Hugging Face
0likes63downloads
Dataset Card

fucc boi bench results v0.2

This is the public results release for fucc boi bench.

  • —13 models
  • —96 prompts per model
  • —1248 scored answers
  • —one combined leaderboard

Files:

  • —responses.jsonl: sanitized model outputs and run metadata.
  • —grades.jsonl: parsed grades, fuccboi scores, serious misses, and rationales.
  • —leaderboard.json: the ranked summary.
  • —summary.json: the interactive site's data and case examples.
  • —benchmark_report.md: the full short report.
  • —benchmark_card.md: the compact benchmark card.

The public export omits request payloads, raw provider responses, credentials, and the earlier private audit/holdout artifacts. The scores are model-judged with a small human calibration sample and are meant as a fun comparison of this prompt set, not a personality test.