CoolFace
Datasetpublic

Orphanage/Baidu_Tieba_SunXiaochuan

说明 随机爬取的百度贴吧孙笑川吧的内容,10万条左右,不包含视频和图片,比较适合用于风格微调(大概)(心虚)。 数据遵循ChatGLM4使用的格式(有需要别的格式请自己调整QWQ)。 清洗的不是很干净,所以把没有清洗的数据也发上来了(QWQ)。 train.jsonl是训练集 dev.jsonl是验证集 No_train_validation_split.jsonl是清洗后并未划分训练和验证集的数据 original.json是爬取后未经清洗的数据 Description This dataset consists of roughly 100,000 samples randomly scraped from the "Sun Xiaochuan" bar on Baidu Tieba. It does not contain videos or images and is generally suitable for style fine-tuning (probably... kind of... maybe… See the full description on the dataset page: https://huggingface.co/datasets/Orphanage/Baidu_Tieba_SunXiaochuan.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
9likes175downloads
Dataset Card

说明

随机爬取的百度贴吧孙笑川吧的内容,10万条左右,不包含视频和图片,比较适合用于风格微调(大概)(心虚)。 数据遵循ChatGLM4使用的格式(有需要别的格式请自己调整QWQ)。 清洗的不是很干净,所以把没有清洗的数据也发上来了(QWQ)。

  • —train.jsonl是训练集
  • —dev.jsonl是验证集
  • —No_train_validation_split.jsonl是清洗后并未划分训练和验证集的数据
  • —original.json是爬取后未经清洗的数据

Description

This dataset consists of roughly 100,000 samples randomly scraped from the "Sun Xiaochuan" bar on Baidu Tieba. It does not contain videos or images and is generally suitable for style fine-tuning (probably... kind of... maybe 👀). The data follows the format used by ChatGLM4 (please adjust to other formats if needed, QWQ). Note: The data has not been thoroughly cleaned, so I've also included the raw uncleaned version just in case (QWQ).

  • —-train.jsonl: Training set
  • —dev.jsonl: Validation set
  • —No_train_validation_split.jsonl: Cleaned data without a train/validation split
  • —original.json: Raw data directly from crawling, uncleaned

作者Authors

This dataset was created and maintained by:

📖 Citation / 引用说明

If you use this dataset in your research, please kindly cite it as follows: 如果您在研究中使用了本数据集,请按如下格式引用:

bibtex
@dataset{zheng2025baidutieba,
  title     = {Baidu Tieba - Sun Xiaochuan Comments Dataset},
  author    = {Ziyu Zheng and zyw (king-of-orphanage)},
  year      = {2025},
  publisher = {Hugging Face Datasets},
  url       = {https://huggingface.co/datasets/Orphanage/Baidu_Tieba_SunXiaochuan}
}

Dataset Version: v1.0
License: CC BY-NC 4.0
Contact: https://huggingface.co/Orphanage
---