CoolFace
Datasetpublic

yuyijiong/Long-Instruction-with-Paraphrasing

🔥 Updates [2024.6.4] Add a slim version. The sample number is reduced from about 20k to 10k. [2024.5.28] The data format is converted from "chatml" to "messages", which is more convenient to use tokenizer.apply_chat_template. The old version has been moved to "legacy" branch. The version without "Original text paraphrasing" is added. 📊 Long Context Instruction-tuning dataset with "Original text paraphrasing" Paper Github consist of multiple tasks Chinese… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/Long-Instruction-with-Paraphrasing.

sourceHugging Facecc-by-sa-4.0updated 2y agoView on Hugging Face
35likes278downloads
Dataset Card

🔥 Updates

\[2024.6.4\] Add a slim version. The sample number is reduced from about 20k to 10k.

\[2024.5.28\]

  1. 1.The data format is converted from "chatml" to "messages", which is more convenient to use ``tokenizer.apply_chat_template``. The old version has been moved to "legacy" branch.
  2. 2.The version without "Original text paraphrasing" is added.

📊 Long Context Instruction-tuning dataset with "Original text paraphrasing"

  • Paper
  • Github
  • consist of multiple tasks
  • Chinese and English
  • sample length ranging from 4k to 32k
  • the answer contains "Original text paraphrasing" part

长文本指令微调数据

  • 此数据集由多种长文本任务数据集组合而成。
  • 包含中文和英文

<center> Dataset Composition (original version)</center>

[image]")

<center> Dataset Composition (slim version)</center>

[image]")

源数据

此处给出各个数据集的链接集合。也可以直接点击我的个人主页查看所有数据集。

中文

  1. 1.图书总结
  1. 1.论文摘要 涉及到知网数据,受限访问。
  2. 2.论文问答 涉及到知网数据,受限访问。
  1. 1.多文档问答(检索)

英文

  1. 1.多文档问答(检索)

中英

  1. 1.长论文多任务
  1. 1.从ShareGPT中筛选的长对话(中英)
  1. 1.预训练长文本语料库(中英)LongData-Corpus