beyarkay/slash-bad
slash-bad Moments when a coding agent did something its user didn't like, flagged in the middle of real work with /bad <what went wrong>, then re-run on newer models. Each row is one flagged moment: the conversation up to that point, the response that annoyed the user, the user's own words about what was wrong with it, and blind verdicts on how far that complaint applies to the original response and to re-runs by other models. The data comes from slash-bad, a personal benchmark… See the full description on the dataset page: https://huggingface.co/datasets/beyarkay/slash-bad.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face