datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EverMemBench-Static
EverMemBench-S: Evaluating Evidence Access under Dense Semantic Interference
💻 Code: EverMind-AI/EverMemBench-Static
Overview
EverMemBench-S (EMB-S) is an adversarial Needle-in-a-Haystack benchmark built on a 326M-token MemoryBank with 160,280 documents across 8 domains. It evaluates long-context models and retrieval systems under dense semantic interference — where near-miss documents create realistic confusion that standard NIAH benchmarks cannot capture.
1,225… See the full description on the dataset page: https://huggingface.co/datasets/EverMind-AI/EverMemBench-Static.staticsmechanics
Engineering Statics and Mechanics Evaluation Dataset
A benchmark dataset for evaluating LLM performance on Engineering Statics and Mechanics of Materials problems.
Statics and Mechanics are core units for 1st year undergraduate Engineering students around the world, and are fundamental understanding for many branches of engineering (civil, structural, mechanical, materials, industrial, mechatronic, aeronautical, space, etc).
Expert developer - I have taught these two units at… See the full description on the dataset page: https://huggingface.co/datasets/mikemolt/staticsmechanics.
