datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FinanceQAFinanceQA is a comprehensive testing suite designed to evaluate LLMs' performance on complex financial analysis tasks that mirror real-world investment work. The dataset aims to be substantially more challenging and practical than existing financial benchmarks, focusing on tasks that require precise calculations and professional judgment.
Paper: https://arxiv.org/abs/2501.18062
Description
The dataset contains two main categories of questions:
Tactical Questions: Questions based on… See the full description on the dataset page: https://huggingface.co/datasets/AfterQuery/FinanceQA.med-slm-ja-before-after
med-slm-ja Before/After ・ 日本医療ガイドライン 新旧比較データセット
プロジェクトページ: https://ikora128.github.io/med-slm-ja-before-after/
医療ガイドラインは数年ごとに改訂され、同じ質問でも答えが変わることがあります。
このデータセットは、日本の医療ガイドライン 201冊・2010〜2024年について、同じガイドラインの旧版と新版で回答が変わった箇所を 46,705 件、新旧の出典付きで集めたものです。
数値基準・診断基準・推奨薬・施設要件など、改訂で変わった知識や新しく登場した知識を、機械的に抽出しています。
なぜ作ったか
古いガイドラインのまま診療すると、患者のリスクになります。
数年前の知識で止まったまま見落とすものは、大きく 2種類 あります。
答えが変わったもの(例: 数値目標や診断基準の改訂)
新しく登場したもの(例: CKD 治療の SGLT2 阻害薬、フィネレノン)
このデータセットは、その両方を type… See the full description on the dataset page: https://huggingface.co/datasets/genshiai-daichi/med-slm-ja-before-after.SO-Python_QA-filtered-2023-tanh_score-after_2023_02SO dataset of pythontag data
Question filters:
images
links
code blocks
Q_Score > 0
Answer_count > 0
CreationDate > 2023-02-01
Answers filters:
images
links
code blocks
Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores
