datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Time-Aware-AlignmentTime-AwarenessTimeAware
TimeAware: Benchmarking Time-Sensitive Fact Recall in Large Language Models
Overview
Who is the US President? The answer changes depending on when the question is asked. While large language models (LLMs) are evaluated on various reasoning tasks, they often miss a crucial dimension: time. In real-world scenarios, the correctness of answers is frequently tied to temporal context.
TimeAware is a large-scale dataset and evaluation benchmark designed to rigorously test LLMs'… See the full description on the dataset page: https://huggingface.co/datasets/hereldav/TimeAware.
