datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
econ_logic_qa
EconLogicQA
EconLogicQA is a benchmark designed to test the sequential reasoning skills of large language models (LLMs) in economics, business,
and supply chain management. It diverges from typical benchmarks by requiring models to understand and sequence multiple interconnected
events, capturing complex economic logics. The benchmark includes multi-event scenarios and a thorough suite of evaluations to assess
proficiency in economic contexts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yinzhu-quan/econ_logic_qa.logicwaver-reasoning-v1
LogicWaver Adversarial Reasoning Benchmark
Adversarial Semantic Reasoning Benchmark — 20 handcrafted odd-one-out puzzles where SOTA LLMs fail but humans succeed. For LLM evaluation & Chain-of-Thought probing.
🚀 Live Demo: https://huggingface.co/spaces/Eviezonr08/logicwaver-demo
▶️ Try 20 puzzles interactively - no install!
Files
Reasoning_Puzzle_without_proline.csv — 20 puzzles for evaluation
with_proline/Reasoning_Puzzle_proline.csv — Same 20 + pro_line… See the full description on the dataset page: https://huggingface.co/datasets/Eviezonr08/logicwaver-reasoning-v1.
