QA with context
dolmino_wiki_rephrased_qa_with_context_concatuscode-qa-with-context
USCode Questions and answers with context
This dataset merges and filters uscode_qac, legalbench:rule_qa, USCode-QAPairs-Finetuning, and synthetic-legal. It contains 57 and 10 rows with context in the train and test splits, the rest are without context.
biomed-qa-20k-with-context
Biomedical QA 20k (with context)
生物医学 / 生命科学问答评测集,共 20,834 条。每条样本都附带回答所依据的原文段落(context),可直接用于检索 / 记忆增强等评测设置。全部样本可自动评测,字段遵循统一的 25 键 JSON schema。
1. 总览
项
值
总条数
20,834(train 17,500 / dev 1,752 / test 1,582)
来源数据集
3 个(PubMedQA / BioMRC / COVID-QA)
题型
yes_no_maybe 9,493 + single_choice 9,600 + short_answer 1,741
原文 context 非空率
20,834 / 20,834 = 100%
平均 Qwen3 token/条
1078.88(train 1098.57 / dev 1082.99 / test 856.54;中位 450,p90 764,max 16882)… See the full description on the dataset page: https://huggingface.co/datasets/zjlergf/biomed-qa-20k-with-context.fwe_med_qa_synthetic_with_contextCreattion process:
annotate 500k sampels from FineweB Edu with LLama 3 70B
train Bert model on them
annotate full FineWeb Edu, take top 12B tokens (-> 6M docs)
take 1M of these docs and apply WRAP rephrasal to it
for this particular one, the rephrasal model is LLama 3.2 3B
it is given random examples from valid. set of MMLU medical domain as context and asked to generate similar Q&A pairs
QA_with_contextdmp-qa-with-context-2
Data management questions and answers with generated context
Questions and answers from dmp-qa with generated context
using for forwards and backwards generation. Attribution Improved with Qwen should be displayed when using this data for finetuning.
Generated context + answer length are around 700 tokens.
