dannkoh/WARP-benchmark
WARP-benchmark The benchmark to test logical reasoning and pattern generalisation in large language models through formal SMT constraints. Overview We propose a benchmark designed to evaluate a language model's ability to generalise worst-case path constraints for all input sizes. Each example asks a model to generate a formal constraint for a specific target size after being shown examples of constraints for smaller input sizes. Github Repository… See the full description on the dataset page: https://huggingface.co/datasets/dannkoh/WARP-benchmark.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face