CoolFace
Datasetpublic

pranaysuyash/laya-formatting-fragility

Laya Formatting Fragility — a perturbation suite for typed decision models Small decision models ("decide, don't chat" — Laya, AgentJev and friends) return typed answers with probabilities in milliseconds. Their answers can be sensitive to formatting: option key names, option order, and state phrasing change results even when the situation and the gold answer are identical. Vendor benchmarks don't measure this. This dataset does. Every row pair differs in exactly one formatting… See the full description on the dataset page: https://huggingface.co/datasets/pranaysuyash/laya-formatting-fragility.

sourceHugging Faceapache-2.0updated 4d agoView on Hugging Face
0likes88downloads
9 commits on main
8e2e0534d ago

Add hosted-engine (jev-1.13.0) column: 90.2% overall, most format-stable (2.9-8.3% flips)

pranaysuyash
0e545c84d ago

Add hosted-engine (jev-1.13.0) column: 90.2% overall, most format-stable (2.9-8.3% flips)

pranaysuyash
c437fdf4d ago

harness v2: fix score-row index mapping + position-based key-flip metric; corrected numbers + AgentJev column

pranaysuyash
ad8dd364d ago

harness v2: fix score-row index mapping + position-based key-flip metric; corrected numbers + AgentJev column

pranaysuyash
881ee964d ago

harness v2: fix score-row index mapping + position-based key-flip metric; corrected numbers + AgentJev column

pranaysuyash
35cae294d ago

harness v2: fix score-row index mapping + position-based key-flip metric; corrected numbers + AgentJev column

pranaysuyash
ba2a0664d ago

Add runnable end-to-end demo notebook (also packaged for Kaggle)

pranaysuyash
30ebc9f4d ago

Formatting-fragility suite for typed decision models: 480-row perturbation dataset + probe harness + measured laya 0.3.5 results

pranaysuyash
4bba4f24d ago

initial commit

pranaysuyash