yuyijiong/Long-Instruction-with-Paraphrasing
🔥 Updates [2024.6.4] Add a slim version. The sample number is reduced from about 20k to 10k. [2024.5.28] The data format is converted from "chatml" to "messages", which is more convenient to use tokenizer.apply_chat_template. The old version has been moved to "legacy" branch. The version without "Original text paraphrasing" is added. 📊 Long Context Instruction-tuning dataset with "Original text paraphrasing" Paper Github consist of multiple tasks Chinese… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/Long-Instruction-with-Paraphrasing.
🔥 Updates
\[2024.6.4\] Add a slim version. The sample number is reduced from about 20k to 10k.
\[2024.5.28\]
- The data format is converted from "chatml" to "messages", which is more convenient to use ``
tokenizer.apply_chat_template``. The old version has been moved to "legacy" branch. - The version without "Original text paraphrasing" is added.
📊 Long Context Instruction-tuning dataset with "Original text paraphrasing"
- Paper
- Github
- consist of multiple tasks
- Chinese and English
- sample length ranging from 4k to 32k
- the answer contains "Original text paraphrasing" part
长文本指令微调数据
- 此数据集由多种长文本任务数据集组合而成。
- 包含中文和英文
<center> Dataset Composition (original version)</center>
")
<center> Dataset Composition (slim version)</center>
")
源数据
此处给出各个数据集的链接集合。也可以直接点击我的个人主页查看所有数据集。
中文
英文
中英
- 预训练长文本语料库(中英)LongData-Corpus
