datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
financial-model-evals
Worthune Financial Model Evals (open sample)
Ground truth for financial AI: 3 open datasets × 250
input/expected-output pairs — the free-sample slice of a 55-model catalog
covering refinance break-evens, retirement projections, Roth conversions,
equity compensation, loan payoff, and more. The full catalog's datasets
download with a Worthune Pro key from
GET https://worthune.com/api/v1/evals/{model} (index).
Every expected output comes from two independent implementations that… See the full description on the dataset page: https://huggingface.co/datasets/worthune/financial-model-evals.Growing-Bench-Worth-It
Growing Bench: Worth It
Correct isn't enough. Was the work worth it?
I have spent a lot of time working with coding Agents and writing Agents. A recurring unpleasant experience is that, under the banner of correctness, necessity, and safety, an Agent ignores the overall need and the actual situation, does the job badly, and leaves the user with a terrible experience.
You ask it to make a small change in a codebase. The change succeeds, but the Agent also creates… See the full description on the dataset page: https://huggingface.co/datasets/takamatsu-hikaru/Growing-Bench-Worth-It.
