CoolFace
Datasetpublic

Orphanage/Baidu_Tieba_SunXiaochuan

说明 随机爬取的百度贴吧孙笑川吧的内容,10万条左右,不包含视频和图片,比较适合用于风格微调(大概)(心虚)。 数据遵循ChatGLM4使用的格式(有需要别的格式请自己调整QWQ)。 清洗的不是很干净,所以把没有清洗的数据也发上来了(QWQ)。 train.jsonl是训练集 dev.jsonl是验证集 No_train_validation_split.jsonl是清洗后并未划分训练和验证集的数据 original.json是爬取后未经清洗的数据 Description This dataset consists of roughly 100,000 samples randomly scraped from the "Sun Xiaochuan" bar on Baidu Tieba. It does not contain videos or images and is generally suitable for style fine-tuning (probably... kind of... maybe… See the full description on the dataset page: https://huggingface.co/datasets/Orphanage/Baidu_Tieba_SunXiaochuan.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
9likes163downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Orphanage/Baidu_Tieba_SunXiaochuan · CoolFace