datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShareGPT4Video
ShareGPT4Video 4.8M Dataset Card
Dataset details
Dataset type:
ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos.
It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora.
sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.ShareGPT_Vicuna_unfiltered
Dataset Card
This is a reupload of this dataset that was further cleaned by gozfarb.
ShareGPT-4oShareGPT-4o-Image
📚 ShareGPT-4o-Image
ShareGPT-4o-Image is a large-scale and high-quality image generation dataset, where all images are produced by GPT-4o’s image generation capabilities. This dataset is designed to align open multimodal models with GPT-4o’s strengths in visual content creation. It includes 45K text-to-image and 46K text-and-image-to-image samples, making it a useful resource for enhancing multimodal models in both image generation and editing tasks.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ShareGPT-4o-Image.function-calling-sharegptThis is a dataset for finetuning models on function calling based on glaiveai/glaive-function-calling-v2.
The dataset includes 86,864 examples of chats that include function calling as part of the conversation. The system prompt includes either 0, 1, or 2 functions that the assistant can use, and instructions on how the agent can use it.
Changes include:
Using ShareGPT format for chats
Adding "function_response" as a role
Removing code examples
Removing examples with invalid JSON as function… See the full description on the dataset page: https://huggingface.co/datasets/hypervariance/function-calling-sharegpt.ultrachat-sharegpt-5GBglaive-function-calling-v2-sharegptThe glaive-function-calling-v2 dataset in sharegpt format.
You can use it in LLaMA Factory by specifying --dataset glaive_toolcall_100k.
Creative_Writing-ShareGPTOriginal Dataset Sources: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts, https://huggingface.co/datasets/anthracite-org/nopm_claude_writing_fixed.
(Thank the original dataset creators for their work.) (Nopm) Claude / (Grphye) ChatGPT-4o Syntheticly generated creative writing set's combined.
Update: Used most up to date version of gryphes, chatGPT-4o set, Rejections/Slop Filtered, Min-hash Deduplication using -… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Creative_Writing-ShareGPT.sharegpt_gpt4
Dataset Card
Dataset Summary
ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。
Languages
数据集是多语言,包括中文、英文、日文等常用语言。
Dataset Structure
Data Fields
The data fields are the same among all splits.
conversations: a List of string .
head -n 1 sharegpt_gpt4.jsonl
{"conversations":[
{'from': 'human',
'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.xlam-function-calling-60k-shareGPTShareGPT converted version of Salesforce/xlam-function-calling-60k
guanaco-sharegpt-style
Dataset Card for "guanaco-sharegpt-style"
More Information needed
ShareGPT4V
News
[2024/5/8] We released ShareGPT4Video, a large-scale video-caption dataset, with 40K captions annotated by GPT4V and 4.8M captions annotated by our ShareCaptioner-Video. The total videos last with 300 hours and 3000 hours separately!
ShareGPT4V 1.2M Dataset Card
Dataset details
Dataset type:
ShareGPT4V Captions 1.2M is a set of GPT4-Vision-powered multi-modal captions data.
It is constructed to enhance modality alignment and fine-grained visual concept… See the full description on the dataset page: https://huggingface.co/datasets/Lin-Chen/ShareGPT4V.sharegpt-quizz-generation-json-output
ShareGPT-Formatted Dataset for Quizz Generation in Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate quizz in structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-quizz-generation-json-output.Python-Code-23k-ShareGPTThis dataset is in Vicuna/ShareGPT format. There are 23000+ set of conversations. Each set having 2 conversations.
Along with the Python code detailed explanation is provided.
This dataset was generated using GPT-3.5, GPT-4 etc.
sharegpt52kglaive-function-calling-v2-sharegpt
Dataset Card for "glaive-function-calling-v2-sharegpt"
This dataset takes the glaive/glaive-function-calling-v2 dataset and formats it with ShareGPT using Lilac
The accompanying notebook can be found here.
The original columns "system" and "chat" still exist on the dataset.
There are 4 types of roles in the ShareGPT format:
system
user
human
function call
The original dataset has a column called 'chat' with the following structure:
USER: Hi, I need help with calculating a tip. My… See the full description on the dataset page: https://huggingface.co/datasets/lilacai/glaive-function-calling-v2-sharegpt.sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.openthoughts-4-math-qwen3-32b-7k-annotated-sharegptDSULT-Core-ShareGPT-X
DSULT-Core/ShareGPT-X Filtered Dataset
This dataset is a curated subset of ShareGPT-X, which contains approximately 92,000 one-to-one conversations between humans and ChatGPT, collected from X.com (formerly Twitter). The corpus covers content from January 2024 through May 2025, built entirely from public "share" links posted by users on their timelines.
The file ChatGPT-Simple_ShareGPT_Full.json includes the longest sequences of alternating human and gpt messages within each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/DSULT-Core-ShareGPT-X.Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1 in ShareGPT Format
This dataset is a conversion of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
into the ShareGPT format while preserving the original splits and columns.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"messages": [
{"role": "user", "content": "User message"},
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT.ShareGPT4V-Rebuild
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/jzsues/ShareGPT4V-Rebuild.proxy-logs-ShareGPTAll the proxy logs I could find (lmk if there are more), converted to ShareGPT so it's all formatted the same way in one dataset. I didn't .strip() any of the turns, I only used ftfy.fix_text() on them, and skipped any empty turns.
The original sample response is moved to be the final turn of conversations. If the last turn in the original sample prompt was a model turn, it is assumed that it was for doing prefill. So I moved this turn to response_prefill.
sample["response_prefill"] and… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/proxy-logs-ShareGPT.openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
34.3
74.5
79.4
49.4
51.0
44.3
53.9
21.5
23.1
12.2
17.0
22.7
40.1
AIME24
Average Accuracy: 34.33% ± 1.89%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.roleplay-zh-sharegpt-gpt4-data
roleplay 数据集
数据
我们有4个数据集文件:
"sharegpt_formatted_data-evol-gpt4.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-male-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-roleplay-chat-1k.jsonl" 来自 Minami-su/roleplay_multiturn_chat_1k_zh_v0.1 将其转换为sharegpt格式。… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/roleplay-zh-sharegpt-gpt4-data.ShareGPT-4o-TextImage2ImageNemotron_Nano_sharegptsharegpt4v_vqa_200k_batch2
License
This is the re-uploaded dataset base on the work of ShareGPT4V team:
https://sharegpt4v.github.io and https://github.com/ShareGPT4Omni/ShareGPT4V
This dataset is under CC BY NC 4.0 license. Therefore, it allows only for non-commercial use and models trained using the dataset should not be used outside of research purposes.
Citation
If you use this datasets in your research, please cite the original paper as follows:
@article{chen2023sharegpt4v… See the full description on the dataset page: https://huggingface.co/datasets/tsystems/sharegpt4v_vqa_200k_batch2.sharegpt-chineseChinese ShareGPT data translated by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT.
sharegpt-englishLocateAnything-Data-ShareGPT-Annotation
