datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FinLFQA
FinLFQA
📖 Paper | 💻 GitHub
The dataset for the paper FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering.
FinLFQA is a benchmark for evaluating the ability of large language models (LLMs) to generate long-form answers with fine-grained attributions in the financial domain. Unlike existing benchmarks that focus on short-form or extractive QA, FinLFQA requires models to synthesize information from multiple financial documents, apply… See the full description on the dataset page: https://huggingface.co/datasets/Dragongon/FinLFQA.car_bench_openai_datasetharmonic-seeking-dragonflylike_a_dragon_infinite_wealth_recordings_01
如龙8 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_0382d8a043d4835a05b57c115877ca33
Collection: general (泛数据)
Recordings: 14
Layout: recordings/<recording_id>/<raw component>
MedQA-USMLE-4-optionsOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Citation information:
@article{jin2020disease,
title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={arXiv preprint arXiv:2009.13081},
year={2020}
}
paraphrased_qwen1.5b_dragon_numsadqa-resultsgemma4b_paraphrased_dragon_cotgemma4b_dragon_cotqwen3b_paraphrased_dragon_cotTweetsNearMSUMoorheadqwen1.5b_dragon_codeDreadPoor__Elusive_Dragon_Heart-8B-LINEAR-details
Dataset Card for Evaluation run of DreadPoor/Elusive_Dragon_Heart-8B-LINEAR
Dataset automatically created during the evaluation run of model DreadPoor/Elusive_Dragon_Heart-8B-LINEAR
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Elusive_Dragon_Heart-8B-LINEAR-details.datasetYcountdown_qwen3teacher_filtered_responsesMedicalQAraft_toy_130_dragonqwen7b_dragon_cotqwen1.5b_dragon_numsqwen3b_dragon_numsqwen1.5b_paraphrased_dragon_cotllama8b_paraphrased_dragon_cotautotrain-data-bp-dataDragonRealms-infoqwen1.5b_dragon_cotqwen3b_dragon_cotllama8b_dragon_cotqwen7b_paraphrased_dragon_numsqwen7b_paraphrased_dragon_cotMedicalData
