ai-research
ai-research-index-iclr-openreview
Conference-paper corpus (AI-Research-Index project)
Private working dataset. Conference/journal paper corpus across OpenReview
venues (ICLR, NeurIPS incl. D&B/position tracks, ICML incl. position, COLM,
TMLR, AISTATS, UAI, ALT, MathAI, and later additions) plus the ACL Anthology
family. Coverage, per-venue availability, decisions semantics, and known
biases are documented authoritatively in the GitHub repo's data/README.md
— read that first; per-venue counts change as the corpus… See the full description on the dataset page: https://huggingface.co/datasets/latkes/ai-research-index-iclr-openreview.bankertoolbench
BankerToolBench
BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for
evaluating AI agents. Each task mirrors real junior-banker work — building
financial models, preparing pitch decks, writing memos — and produces multi-file
deliverables (Excel, PowerPoint, Word) that are scored against expert-authored
rubrics.
The benchmark was developed with 502 investment bankers from firms including
Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/bankertoolbench.FinDER
FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation
FinDER is a benchmark dataset designed for evaluating Retrieval-Augmented Generation (RAG) in financial question answering. It consists of 5,703 expert-annotated query–evidence–answer triplets derived from real-world 10-K filings and ambiguous financial queries submitted by industry professionals.
This dataset captures the domain-specific challenges of financial QA, including short… See the full description on the dataset page: https://huggingface.co/datasets/Linq-AI-Research/FinDER.QIVD
QIVD: Qualcomm Interactive Video Dataset
A collection of 2,900 video clips paired with visual question-answer annotations.
Each clip is associated with exactly one question drawn from one of 13 fine-grained QA categories,
a full-sentence answer, a concise short answer, and a timestamp pinpointing the relevant moment in the video.
Overview
QIVD is a dataset and benchmark for online, situated audio-visual question answering. Unlike existing video QA benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/Qualcomm-AI-Research/QIVD.StreamBench
StreamBench paper link: https://arxiv.org/abs/2406.08747 (The links for the original raw datasets on StreamBench can be found in Appendix F)
If you find our work helpful, please cite as:
@article{wu2024streambench,
title={StreamBench: Towards Benchmarking Continuous Improvement of Language Agents},
author={Wu, Cheng-Kuang and Tam, Zhi Rui and Lin, Chieh-Yen and Chen, Yun-Nung and Lee, Hung-yi},
journal={arXiv preprint arXiv:2406.08747},
year={2024}
}
movie_gen_video_bench
Dataset Card for the Movie Gen Benchmark
Movie Gen is a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio.
Here, we introduce our evaluation benchmark "Movie Gen Bench Video Bench", as detailed in the Movie Gen technical report (Section 3.5.2).
To enable fair and easy comparison to Movie Gen for future works on these evaluation benchmarks, we additionally release the non cherry-picked generated videos from… See the full description on the dataset page: https://huggingface.co/datasets/meta-ai-for-media-research/movie_gen_video_bench.
