datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FRED-SemBench
FRED-SemBench
FRED-SemBench is a 200-question benchmark candidate for evaluating whether LLM
agents retrieve macroeconomic answers with the intended concept, series,
transformation, unit, observation period, and data-vintage semantics.
This dataset accompanies the FinNLP 2026 paper
“FRED-SemBench: Evaluating Semantic Reliability in LLM Access to
Macroeconomic Data”
by Wilson Wang, Chandler Han, and Peter Zhang (Kairos-AI).
Status and scope
50 independently… See the full description on the dataset page: https://huggingface.co/datasets/wangjinh/FRED-SemBench.Material_Selection_EvalA benchmark designed to facilitate evaluation and modify the behavior of a foundation model through different existing techniques in the context of material selection for conceptual design.
The data is collected by conducting a survey of experts in the field of material selection. The same questions mentioned in keyquestions.csv are asked to experts.
This can be used to evaluate a Language model performance and its spread compared to a human evaluation.
To get into a more detailed explanation… See the full description on the dataset page: https://huggingface.co/datasets/Frederick001/Material_Selection_Eval.
