RolandM/imprecision-bench
imprecision-bench A multimodal benchmark for evaluating whether LLMs calibrate linguistic precision to pragmatic context, paired with 475 human productions and a peer-reviewed RSA baseline (r² ≈ 0.97). This dataset accompanies the paper: Modeling (Im)precision in Context Roland Mühlenbernd, Stephanie Solt Linguistics Vanguard, 2022 [Paper] · [Source Data] · [Companion Repo] Notebook notebook.ipynb — guided walkthrough: data loading, sample evaluation (1 row… See the full description on the dataset page: https://huggingface.co/datasets/RolandM/imprecision-bench.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face