datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DSULT-Core-ShareGPT-X
DSULT-Core/ShareGPT-X Filtered Dataset
This dataset is a curated subset of ShareGPT-X, which contains approximately 92,000 one-to-one conversations between humans and ChatGPT, collected from X.com (formerly Twitter). The corpus covers content from January 2024 through May 2025, built entirely from public "share" links posted by users on their timelines.
The file ChatGPT-Simple_ShareGPT_Full.json includes the longest sequences of alternating human and gpt messages within each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/DSULT-Core-ShareGPT-X.ShareGPT-X
Dataset Summary
ShareGPT-X is an expanded, snapshot of ~92K (ChatGPT) one-to-one human & LLM conversations harvested from X.com (formerly Twitter).The corpus spans January 2024 → present (last ingest 2025-05) and is built entirely from public "share" links that users posted to their timelines.Each thread contains the original user prompt plus the assistant’s reply; no system prompts or metadata are exposed.
Supported Tasks and Leaderboards
text-generation… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/ShareGPT-X.thomas-yanxin-MT-SFT-ShareGPT-sample
MT-SFT-ShareGPT Sample Dataset
This dataset provides a sample of the thomas-yanxin/MT-SFT-ShareGPT dataset with English and Chinese subsets.
Dataset Contents
train.jsonl: Contains 1/10 of the original data, shuffled
EN.jsonl: English conversations from train.jsonl
ZH.jsonl: Chinese conversations from train.jsonl
Each row represents a conversation with an optional system message, followed by human and GPT turns.
Columns from the original dataset are preserved, with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample.thomas-yanxin-MT-SFT-ShareGPT
thomas-yanxin/MT-SFT-ShareGPT
This is the complete thomas-yanxin/MT-SFT-ShareGPT dataset,
with duplicates removed and the entire dataset shuffled. Sensitive data has been redacted.
For practical work, consider using agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample
which is smaller and split by language.
