datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docqa-rl-verl
DocQA-RL-1.6K (VERL Format)
This dataset contains 1,591 challenging long-context document QA problems from DocQA-RL-1.6K, converted to VERL (Volcano Engine Reinforcement Learning) format for reinforcement learning training workflows.
Source: Tongyi-Zhiwen/DocQA-RL-1.6K
License: Apache 2.0
Note: This dataset maintains the original high-quality structure with user-only messages. The extra_info field has been standardized to contain only the index field for consistency with other VERL… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/docqa-rl-verl.readtwice-docqa-hqa-rl-2560-20260925
ReadTwice DocQA/HQA RL 训练数据
这是本次 14B step15 后的 DocQA RL 阶段实际读取的数据,与此前 7B 初始 DocQA 阶段使用的混合数据相同。mixture.parquet 和 mixture_rows.json 与训练输入逐字节一致。
来源
条数
DocQA-RL-1.6K(train,去重后)
1,585
HotpotQA 32K(train)
975
合计
2,560
DocQA 含 624 道选择题、553 道数值题、408 道文本问答。保留原始题目、选项和完整文档;文档最长 59,473 tokens,每 chunk 最多 5,000 tokens,最多 12 chunks。长度桶的 112000 是桶标签,不是每条文档有112K tokens。
这份数据没有做离线 pass@k 筛选;原训练流程会用当前策略每题采样8次,保留答对1–7次的题组。数据构建时排除了记录在源清单中的评测题目/文档重合。没有加入本次 oracle 理想上限评测数据,也没有把答案或… See the full description on the dataset page: https://huggingface.co/datasets/Xirui1208/readtwice-docqa-hqa-rl-2560-20260925.
