IAA
Datasets
All datasets matching “IAA”LongEvoRoleBench
LongEvoRoleBench
LongEvoRoleBench is a unified benchmark for long-horizon, evolution-aware role-playing.
It standardizes 8 existing character-dialogue corpora into a common next-utterance protocol:
4 long-dialogue corpora test cross-episode character-state evolution, and 4 short-dialogue
corpora provide within-scene state-tracking checks under the same evaluation format.
This repository ships both the raw source corpora and the fully processed evaluation splits.
The… See the full description on the dataset page: https://huggingface.co/datasets/IAAR-Shanghai/LongEvoRoleBench.iaai-dataset
IAAI Insurance Auto Auction Dataset
Daily sample of IAAI insurance auto auction lots with damage assessments, title status, bidding data, and branch locations across North America.
This dataset is a preview sample of the IAAI dataset published by Rebrowser. If you're doing academic research, you may be eligible for free access to a much larger slice — see Free Datasets for Research.
This dataset contains 1 entity, each in its own folder: Auction Listings (auction-listings). See… See the full description on the dataset page: https://huggingface.co/datasets/rebrowser/iaai-dataset.LimAgents_limitation_data_scientific_papers_with_cited_papers
LimAgents Data
This dataset contains scientific paper metadata and extracted limitation information prepared for use with LLM Agents.The data comes from NeurIPS 2021–2022 papers and related OpenReview reviews, enriched with Cited in and Cited by information.
Dataset Structure
The repository contains two main directories:
1. NeurIPS_21_22_Lim_OPR_with_cited_in_by_papers
This directory includes one JSON file per paper. Each file contains:
title: Original paper… See the full description on the dataset page: https://huggingface.co/datasets/iaadlab/LimAgents_limitation_data_scientific_papers_with_cited_papers.HaluMem
HaluMem: A Comprehensive Benchmark for Evaluating Hallucinations in Memory Systems
📊 Why We Define the HaluMem Evaluation Tasks
Limitations of Existing Frameworks
Most existing evaluation frameworks treat memory systems as black-box models, assessing performance only through end-to-end QA accuracy.
However, this approach has two major limitations:
It lacks a hallucination evaluation specifically designed for the characteristics of memory systems.… See the full description on the dataset page: https://huggingface.co/datasets/IAAR-Shanghai/HaluMem.FGen
license: cc-by-4.0
📚 ACL & NeurIPS Dataset
📦 Dataset Details
🔍 Dataset Description
This dataset includes:
ACL_2012.csv to ACL_2024.csv: Tabular data where each row is a paper and each column represents a paper section.
NeurIPS_2021.csv, NeurIPS_2022.csv: Similar format as ACL .csv files.
ACL_2023.json, ACL_2024.json: Each file contains paper-wise parsed output including section headers and content.
Each record is either a paper (in .csv) or a… See the full description on the dataset page: https://huggingface.co/datasets/iaadlab/FGen.KAF-DatasetThe dataset sourced from https://github.com/IAAR-Shanghai/xFinder
Citation
@inproceedings{
xFinder,
title={xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation},
author={Qingchen Yu and Zifan Zheng and Shichao Song and Zhiyu li and Feiyu Xiong and Bo Tang and Ding Chen},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=7UqQJUKaLM}
}
