CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01iaadlab /LimAgents_limitation_data_scientific_papers_with_cited_papers LimAgents Data This dataset contains scientific paper metadata and extracted limitation information prepared for use with LLM Agents.The data comes from NeurIPS 2021–2022 papers and related OpenReview reviews, enriched with Cited in and Cited by information. Dataset Structure The repository contains two main directories: 1. NeurIPS_21_22_Lim_OPR_with_cited_in_by_papers This directory includes one JSON file per paper. Each file contains: title: Original paper… See the full description on the dataset page: https://huggingface.co/datasets/iaadlab/LimAgents_limitation_data_scientific_papers_with_cited_papers.text-classification1K<n<10K0 likes620 downloads1y agoHugging Face02IAAR-Shanghai /HaluMem HaluMem: A Comprehensive Benchmark for Evaluating Hallucinations in Memory Systems 📊 Why We Define the HaluMem Evaluation Tasks Limitations of Existing Frameworks Most existing evaluation frameworks treat memory systems as black-box models, assessing performance only through end-to-end QA accuracy. However, this approach has two major limitations: It lacks a hallucination evaluation specifically designed for the characteristics of memory systems.… See the full description on the dataset page: https://huggingface.co/datasets/IAAR-Shanghai/HaluMem.question-answering1K<n<10K12 likes392 downloads11mo agoHugging Face03IAAR-Shanghai /KAF-DatasetThe dataset sourced from https://github.com/IAAR-Shanghai/xFinder Citation @inproceedings{ xFinder, title={xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation}, author={Qingchen Yu and Zifan Zheng and Shichao Song and Zhiyu li and Feiyu Xiong and Bo Tang and Ding Chen}, booktitle={The Thirteenth International Conference on Learning Representations}, year={2025}, url={https://openreview.net/forum?id=7UqQJUKaLM} } textquestion-answering10K<n<100K6 likes132 downloads1y agoHugging Face04IAAR-Shanghai /VAR xVerify: Efficient Answer Verifier for Reasoning Model Evaluations 📘 Introduction xVerify is an evaluation tool fine-tuned from a pre-trained large language model, designed specifically for objective questions with a single correct answer. It accurately extracts the final answer from lengthy reasoning processes and efficiently identifies equivalence across different forms of mathematical expressions, LaTeX and string representations, as well as… See the full description on the dataset page: https://huggingface.co/datasets/IAAR-Shanghai/VAR.question-answering10K<n<100K1 likes95 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.