longwriter
LongWriter-6k
LongWriter-6k
🤗 [LongWriter Dataset] • 💻 [Github Repo] • 📃 [LongWriter Paper]
LongWriter-6k dataset contains 6,000 SFT data with ultra-long output ranging from 2k-32k words in length (both English and Chinese). The data can support training LLMs to extend their maximum output window size to 10,000+ words.
All Models
We open-sourced the following list of models trained on LongWriter-6k:
Model
Huggingface Repo
Description
LongWriter-glm4-9b
🤗… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongWriter-6k.LongWriter-V-22KLongWriter-Zero-RLData
LongWriter-Zero RL Data
🤗 [Model] • 📃 [Paper] • 💾 [Dataset Card]
LongWriter-Zero RL Data is designed for ultra-long text generation via reinforcement learning. The dataset consists of conversational queries paired with length-range tags, which specify the desired output span (measured in words or Chinese characters).
These annotations are used to train the LongWriter-Zero model, enabling it to consistently generate passages exceeding 10,000 words.
PS: We also included some… See the full description on the dataset page: https://huggingface.co/datasets/THU-KEG/LongWriter-Zero-RLData.longwriter-6k-filtered
LongWriter-6k-Filtered
🤖 [LongWriter Dataset] • 💻 [Github Repo] • 📃 [LongWriter Paper] • 📃 [Tech report]
longwriter-6k-filtered dataset contains 666 filtered examples SFT data with ultra-long output ranging from 2k-32k words in length (both English and Chinese) based on LongWriter-6k.The data can support training LLMs to extend their maximum output window size to 10,000+ words with low computational cost.
The tech report is available at Minimum Tuning to Unlock Long Output… See the full description on the dataset page: https://huggingface.co/datasets/lenML/longwriter-6k-filtered.Japanese-LongWriter-3k
Dataset Information
このデータセットは数千~数万字の長文なoutputを持つinstructionデータセットです。
Qwen/Qwen2.5-32B-Instruct使って生成しています。
Detail
https://zenn.dev/kendama/articles/32aa9ec4bed409
longwriter-6k-english-v0
