CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01latkes /ai-research-index-iclr-openreview Conference-paper corpus (AI-Research-Index project) Private working dataset. Conference/journal paper corpus across OpenReview venues (ICLR, NeurIPS incl. D&B/position tracks, ICML incl. position, COLM, TMLR, AISTATS, UAI, ALT, MathAI, and later additions) plus the ACL Anthology family. Coverage, per-venue availability, decisions semantics, and known biases are documented authoritatively in the GitHub repo's data/README.md — read that first; per-venue counts change as the corpus… See the full description on the dataset page: https://huggingface.co/datasets/latkes/ai-research-index-iclr-openreview.tabularn<1K0 likes3.7k downloads41m agoHugging Face02handshake-ai-research /bankertoolbench BankerToolBench BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for evaluating AI agents. Each task mirrors real junior-banker work — building financial models, preparing pitch decks, writing memos — and produces multi-file deliverables (Excel, PowerPoint, Word) that are scored against expert-authored rubrics. The benchmark was developed with 502 investment bankers from firms including Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/bankertoolbench.documenttext-generationn<1K9 likes2.9k downloads4mo agoHugging Face03Linq-AI-Research /FinDER FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation FinDER is a benchmark dataset designed for evaluating Retrieval-Augmented Generation (RAG) in financial question answering. It consists of 5,703 expert-annotated query–evidence–answer triplets derived from real-world 10-K filings and ambiguous financial queries submitted by industry professionals. This dataset captures the domain-specific challenges of financial QA, including short… See the full description on the dataset page: https://huggingface.co/datasets/Linq-AI-Research/FinDER.text1K<n<10K19 likes947 downloads1y agoHugging Face04Qualcomm-AI-Research /QIVD QIVD: Qualcomm Interactive Video Dataset A collection of 2,900 video clips paired with visual question-answer annotations. Each clip is associated with exactly one question drawn from one of 13 fine-grained QA categories, a full-sentence answer, a concise short answer, and a timestamp pinpointing the relevant moment in the video. Overview QIVD is a dataset and benchmark for online, situated audio-visual question answering. Unlike existing video QA benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/Qualcomm-AI-Research/QIVD.textvideo-text-to-text1K<n<10K1 likes750 downloads2mo agoHugging Face05appier-ai-research /StreamBench StreamBench paper link: https://arxiv.org/abs/2406.08747 (The links for the original raw datasets on StreamBench can be found in Appendix F) If you find our work helpful, please cite as: @article{wu2024streambench, title={StreamBench: Towards Benchmarking Continuous Improvement of Language Agents}, author={Wu, Cheng-Kuang and Tam, Zhi Rui and Lin, Chieh-Yen and Chen, Yun-Nung and Lee, Hung-yi}, journal={arXiv preprint arXiv:2406.08747}, year={2024} } text10K<n<100K7 likes702 downloads2y agoHugging Face06meta-ai-for-media-research /movie_gen_video_bench Dataset Card for the Movie Gen Benchmark Movie Gen is a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. Here, we introduce our evaluation benchmark "Movie Gen Bench Video Bench", as detailed in the Movie Gen technical report (Section 3.5.2). To enable fair and easy comparison to Movie Gen for future works on these evaluation benchmarks, we additionally release the non cherry-picked generated videos from… See the full description on the dataset page: https://huggingface.co/datasets/meta-ai-for-media-research/movie_gen_video_bench.text1K<n<10K29 likes681 downloads2y agoHugging Face07matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes679 downloads3y agoHugging Face08handshake-ai-research /VIALS VIALS: Visual Interpretation of Artifacts in the Life Sciences The VIALS benchmark evaluates: How accurately can frontier models interpret the visual artifacts routinely encountered in professional life sciences workflows? This is the data for the benchmark, which contains 161 visual question answering (VQA) tasks spanning many industry-relevant scientific domains and artifact types. Each task pairs a scientific image with a question requiring the extraction and interpretation… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/VIALS.visual-question-answeringn<1K1 likes607 downloads8d agoHugging Face09LG-AI-Research /PANORAMAdocument4 likes560 downloads1y agoHugging Face10appier-ai-research /robust-finetuningPlease refer to the following source for the original datasets: GSM8K: https://huggingface.co/datasets/openai/gsm8k MATH: https://huggingface.co/datasets/hendrycks/competition_math math-resample: In this section we subsample the 1,000 subsample only (yes it's balance) HumanEval+: https://huggingface.co/datasets/evalplus/humanevalplus MBPP: https://huggingface.co/datasets/google-research-datasets/mbpp MBPP+: https://huggingface.co/datasets/evalplus/mbppplus ARC Challenge:… See the full description on the dataset page: https://huggingface.co/datasets/appier-ai-research/robust-finetuning.tabular10K<n<100K3 likes558 downloads1y agoHugging Face11ulab-ai /ResearchArcade-openreview-reviewstext100K<n<1M1 likes516 downloads7mo agoHugging Face12sreearravind /AI-Research-Evaluation-Repository-STEM AI-STEM-Research-Eval-Dataset Overview This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations. It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content. The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.text-generationn<1K1 likes497 downloads3mo agoHugging Face13ulab-ai /ResearchArcade-openreview-paperstext10K<n<100K0 likes464 downloads7mo agoHugging Face14handshake-ai-research /ATLAS-Finance ATLAS Finance A benchmark of 100 expert-level tasks inside 13 realistic financial firm environments, packaged in the Harbor RLE format. Each task drops an AI agent into a Linux workstation with a persistent multi-app world — inbox, chat, calendar, virtual data room, drive, wiki — and asks the agent to produce the same deliverable a financial professional would be responsible for: an Excel workbook containing the model and supporting analysis. Here we provide the data for this… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/ATLAS-Finance.documenttext-generationn<1K3 likes452 downloads9d agoHugging Face15ulab-ai /ResearchArcade-openreview-authorstext100K<n<1M0 likes403 downloads7mo agoHugging Face16MinjaeLee-FuriosaAI-Ext /ai-research-berkeley-agentic-verification-harness-optgated Agentic Verification Meta-Verifier Traces This public, manually gated Dataset repository stores immutable phase snapshots from Meta-Verifier experiments. Each run is organized as: experiments/<theme>/<method>/<run>/phases/ train/ # Solver, delegated-verifier, Proposer/Reflector, harness population val/ # Full validation traces, metrics, and selected frozen harness heldout/ # Claimed held-out stage, all scheduled cells, and final scores Access requests are… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-agentic-verification-harness-opt.0 likes402 downloads2d agoHugging Face17airesearch /WangchanThaiInstructWangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai (EMNLP'25) WangchanThaiInstruct is a human-authored Thai dataset that improves instruction-following in low-resource settings, capturing cultural and domain-specific nuances across four domains and seven task types. The evaluate code can be found at this github link @inproceedings{limkonchotiwat2025thaiinstruct, title = {WangchanThaiInstruct: An Instruction-Following… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanThaiInstruct.texttext-generation10K<n<100K29 likes394 downloads9mo agoHugging Face18airesearch /thai-ser 🇹🇭 THAI-SER Dataset 🎭 [📝 Paper (preprint)] Published by: AI Research Institute of Thailand (AIResearch) In collaboration with: Vidyasirimedhi Institute of Science and Technology (VISTEC) Digital Economy Promotion Agency (depa) Department of Computer Engineering, Faculty of Engineering, Chulalongkorn University Department of Dramatic Arts, Faculty of Arts, Chulalongkorn University Sponsored by: Advanced Info Services Public Company Limited (AIS), and Siam Commercial… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/thai-ser.audioaudio-classification10K<n<100K6 likes372 downloads2mo agoHugging Face19airesearch /generated_reviews_enth `generated_reviews_enth` Generated product reviews dataset for machine translation quality prediction, part of [scb-mt-en-th-2020](https://arxiv.org/pdf/2007.03541.pdf) `generated_reviews_enth` is created as part of [scb-mt-en-th-2020](https://arxiv.org/pdf/2007.03541.pdf) for machine translation task. This dataset (referred to as `generated_reviews_yn` in [scb-mt-en-th-2020](https://arxiv.org/pdf/2007.03541.pdf)) are English product reviews generated by [CTRL](https://arxiv.org/abs/1909.05858), translated by Google Translate API and annotated as accepted or rejected (`correct`) based on fluency and adequacy of the translation by human annotators. This allows it to be used for English-to-Thai translation quality esitmation (binary label), machine translation, and sentiment analysis.translation100K<n<1M8 likes348 downloads3y agoHugging Face20Swarm-AI-Research /fable5-traces-sft Fable 5 Traces — Unified SFT / Self-Distillation Dataset A cleaned, unified, PII-scrubbed corpus of Claude Fable 5 agent traces in OpenAI-style chat format, plus a working on-policy self-distillation (SDFT) training scaffold. Composition Source Conversations Claude Code raw agentic sessions 18 CoT distillation records 4,665 Unique conversations (post-dedup) 4,683 Split deterministically by content hash: train 4,442 / validation 241. The raw… See the full description on the dataset page: https://huggingface.co/datasets/Swarm-AI-Research/fable5-traces-sft.texttext-generation1K<n<10K2 likes343 downloads3mo agoHugging Face21airesearch /WangchanThaiMedicaltext1M<n<10M0 likes340 downloads2y agoHugging Face22ulab-ai /research-bench ResearchBench This repository contains the ResearchBench dataset presented in the paper ResearchTown: Simulator of Human Research Community. ResearchBench is a dataset for research community simulation. It includes 1000 paper writing tasks in PaperBench (333 hard, 334 medium, 333 easy) and 200 review writing tasks in ReviewBench. All the data from paper writing tasks and review writing tasks are collected from NeurIPS 2024 and ICLR 2024. Additionally, we also provide 100 extreme… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/research-bench.graph-ml1K<n<10K8 likes320 downloads2y agoHugging Face23ulab-ai /ResearchArcade-openreview-papers-authorstext100K<n<1M0 likes294 downloads7mo agoHugging Face24Frontier-AI-Research /MORALISE Dataset and Evaluation Code for Moral Alignment Benchmark This repository contains evaluation code and data for assessing moral alignment in both text-centric and image-centric settings, across open-source and closed-source models. 📂 Directory Structure ├── Code: Will release soon ├── M1/ # Text-centric morally *wrong* examples ├── M1-Right/ # Text-centric morally *correct* examples ├── M2/ # Image-centric… See the full description on the dataset page: https://huggingface.co/datasets/Frontier-AI-Research/MORALISE.image1K<n<10K0 likes287 downloads11mo agoHugging Face25gen-ai-researcher /cardio-mark CardioMark Review Subset This repository contains an anonymized review subset of the CardioMark benchmark introduced for automated vertebral heart score (VHS) estimation in canine thoracic radiographs. The subset is provided to support reproducibility and data-quality inspection during peer review. Dataset Overview CardioMark is a large-scale benchmark for evaluating the complete VHS measurement pipeline, including: cardiac landmark localization geometric VHS estimation… See the full description on the dataset page: https://huggingface.co/datasets/gen-ai-researcher/cardio-mark.imageimage-classificationn<1K1 likes287 downloads5mo agoHugging Face26airesearch /WangchanX-Legal-ThaiCCL-RAG 🏛️ WangchanX-Legal-ThaiCCL-RAG [Technical Report] The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.texttext-generation10K<n<100K13 likes283 downloads2y agoHugging Face27MinjaeLee-FuriosaAI-Ext /ai-research-berkeley-agentic-verification-benchmarkgatedInternal documentation Analysis and results Consolidated English analysis · Raw data · Figures The report separates evaluation tasks and exact group memberships. Each plot appears above its data table. Current trained MV results cover 12 settings with 99 groups each; published baselines use a separate 99-group cohort. A matched MV vs. fixed-baseline comparison uses the same 65 DeepSWE groups. The full 98-group intersection is documented; baseline episode records are needed to… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-agentic-verification-benchmark.0 likes236 downloads11h agoHugging Face28matlok /python-text-copilot-training-instruct-ai-research Building an AI Copilot Dataset to help keep up with Leading AI Research This is a specialized, instruction dataset for training python coding assistants on how to code from leading AI/ML open source repositories (2.3M coding samples). This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details This dataset holds the latest coding changes from >1159… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research.tabulartext-generation10K<n<100K0 likes232 downloads3y agoHugging Face29airesearch /scb_mt_enth_2020scb-mt-en-th-2020: A Large English-Thai Parallel Corpus The primary objective of our work is to build a large-scale English-Thai dataset for machine translation. We construct an English-Thai machine translation dataset with over 1 million segment pairs, curated from various sources, namely news, Wikipedia articles, SMS messages, task-based dialogs, web-crawled data and government documents. Methodology for gathering data, building parallel texts and removing noisy sentence pairs are presented in a reproducible manner. We train machine translation models based on this dataset. Our models' performance are comparable to that of Google Translation API (as of May 2020) for Thai-English and outperform Google when the Open Parallel Corpus (OPUS) is included in the training data for both Thai-English and English-Thai translation. The dataset, pre-trained models, and source code to reproduce our work are available for public use.translation1M<n<10M9 likes214 downloads3y agoHugging Face30matlok /python-text-copilot-training-instruct-ai-research-2024-02-10 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-10 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-10.tabulartext-generationn<1K0 likes214 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.