datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShareGPT4Video
ShareGPT4Video 4.8M Dataset Card
Dataset details
Dataset type:
ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos.
It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora.
sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.ShareGPT_Vicuna_unfiltered
Dataset Card
This is a reupload of this dataset that was further cleaned by gozfarb.
ShareGPT-4oShareGPT-4o-Image
📚 ShareGPT-4o-Image
ShareGPT-4o-Image is a large-scale and high-quality image generation dataset, where all images are produced by GPT-4o’s image generation capabilities. This dataset is designed to align open multimodal models with GPT-4o’s strengths in visual content creation. It includes 45K text-to-image and 46K text-and-image-to-image samples, making it a useful resource for enhancing multimodal models in both image generation and editing tasks.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ShareGPT-4o-Image.function-calling-sharegptThis is a dataset for finetuning models on function calling based on glaiveai/glaive-function-calling-v2.
The dataset includes 86,864 examples of chats that include function calling as part of the conversation. The system prompt includes either 0, 1, or 2 functions that the assistant can use, and instructions on how the agent can use it.
Changes include:
Using ShareGPT format for chats
Adding "function_response" as a role
Removing code examples
Removing examples with invalid JSON as function… See the full description on the dataset page: https://huggingface.co/datasets/hypervariance/function-calling-sharegpt.glaive-function-calling-v2-sharegptThe glaive-function-calling-v2 dataset in sharegpt format.
You can use it in LLaMA Factory by specifying --dataset glaive_toolcall_100k.
Creative_Writing-ShareGPTOriginal Dataset Sources: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts, https://huggingface.co/datasets/anthracite-org/nopm_claude_writing_fixed.
(Thank the original dataset creators for their work.) (Nopm) Claude / (Grphye) ChatGPT-4o Syntheticly generated creative writing set's combined.
Update: Used most up to date version of gryphes, chatGPT-4o set, Rejections/Slop Filtered, Min-hash Deduplication using -… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Creative_Writing-ShareGPT.sharegpt_gpt4
Dataset Card
Dataset Summary
ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。
Languages
数据集是多语言,包括中文、英文、日文等常用语言。
Dataset Structure
Data Fields
The data fields are the same among all splits.
conversations: a List of string .
head -n 1 sharegpt_gpt4.jsonl
{"conversations":[
{'from': 'human',
'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.xlam-function-calling-60k-shareGPTShareGPT converted version of Salesforce/xlam-function-calling-60k
ShareGPT4V
News
[2024/5/8] We released ShareGPT4Video, a large-scale video-caption dataset, with 40K captions annotated by GPT4V and 4.8M captions annotated by our ShareCaptioner-Video. The total videos last with 300 hours and 3000 hours separately!
ShareGPT4V 1.2M Dataset Card
Dataset details
Dataset type:
ShareGPT4V Captions 1.2M is a set of GPT4-Vision-powered multi-modal captions data.
It is constructed to enhance modality alignment and fine-grained visual concept… See the full description on the dataset page: https://huggingface.co/datasets/Lin-Chen/ShareGPT4V.Python-Code-23k-ShareGPTThis dataset is in Vicuna/ShareGPT format. There are 23000+ set of conversations. Each set having 2 conversations.
Along with the Python code detailed explanation is provided.
This dataset was generated using GPT-3.5, GPT-4 etc.
DSULT-Core-ShareGPT-X
DSULT-Core/ShareGPT-X Filtered Dataset
This dataset is a curated subset of ShareGPT-X, which contains approximately 92,000 one-to-one conversations between humans and ChatGPT, collected from X.com (formerly Twitter). The corpus covers content from January 2024 through May 2025, built entirely from public "share" links posted by users on their timelines.
The file ChatGPT-Simple_ShareGPT_Full.json includes the longest sequences of alternating human and gpt messages within each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/DSULT-Core-ShareGPT-X.roleplay-zh-sharegpt-gpt4-data
roleplay 数据集
数据
我们有4个数据集文件:
"sharegpt_formatted_data-evol-gpt4.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-male-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-roleplay-chat-1k.jsonl" 来自 Minami-su/roleplay_multiturn_chat_1k_zh_v0.1 将其转换为sharegpt格式。… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/roleplay-zh-sharegpt-gpt4-data.sharegpt-chineseChinese ShareGPT data translated by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT.
sharegpt-englishLocateAnything-Data-ShareGPT-AnnotationOlympiad_Math-ShareGPT(No system prompts)
Converted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
sharegpt_llama3_8b_hidden_statesGerman-RAG-SFT-ShareGPT-HESSIAN-AI
German-RAG-SFT (Supervised Fine-Tuning) Share-GPT Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The SFT Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-SFT-ShareGPT-HESSIAN-AI.Cybersecurity-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
Reflection-Dataset-ShareGPT-v2
Simple "Reflection" method dataset inspired by mattshumer
This is the ShareGPT version. Find prompt and response pair dataset here
This dataset was synthetically generated using Glaive AI. There have been structure improvements and added more rows.
Reddit-SFW-Writing_Prompts_ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
[Description Tags],"Deleted user", "Hello,\n\nYour post has been removed..", "Post has been deleted by user", "This post has been marked NSFW", duplicated system and human turns, etc has been removed.
ConversationChronicles-sharegpt-SHARDEDThis is a sharded version of the PocketDoc/ConversationChronicles-sharegpt dataset, a sharegpt conversion of the jihyoung/ConversationChronicles dataset.
All dialogue got fixed (space, coma) and spread across the different relationship available :
Relationship
Count
Ratio
Classmates
66,090
33.05%
Neighbors
49,521
24.76%
Co-workers
28,856
14.43%
Mentee and Mentor
16,035
8.02%
Husband and Wife
13,486
6.74%
Patient and Doctor
6,980
3.49%
Parent and Child6,514
3.26%… See the full description on the dataset page: https://huggingface.co/datasets/Undi95/ConversationChronicles-sharegpt-SHARDED.Code-290k-ShareGPTCode-290k-ShareGPT
This dataset is in Vicuna/ShareGPT format. There are around 290000 set of conversations. Each set having 2 conversations.
Along with Python, Java, JavaScript, GO, C++, Rust, Ruby, Sql, MySql, R, Julia, Haskell, etc. code with detailed explanation are provided.
This datset is built upon using my existing Datasets Python-Code-23k-ShareGPT
and Code-74k-ShareGPT
My Models Python-Code-13B and Python-Code-33B are trained on Python-Code-23k-ShareGPT.
My Models Code-13B and… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Code-290k-ShareGPT.ShareGPT-X
Dataset Summary
ShareGPT-X is an expanded, snapshot of ~92K (ChatGPT) one-to-one human & LLM conversations harvested from X.com (formerly Twitter).The corpus spans January 2024 → present (last ingest 2025-05) and is built entirely from public "share" links that users posted to their timelines.Each thread contains the original user prompt plus the assistant’s reply; no system prompts or metadata are exposed.
Supported Tasks and Leaderboards
text-generation… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/ShareGPT-X.huatuo_medical_qa_sharegptsource:
https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT-sft-data-v1
https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2_sft_instruct_GPT4_50K
转为sharegpt格式,jsonl文件。
data size:
> wc -l HuatuoGPT_sft_data_v1_sharegpt.jsonl
226042 HuatuoGPT_sft_data_v1_sharegpt.jsonl
> wc -l HuatuoGPT2_sft_instruct_GPT4_sharegpt.jsonl
50000 HuatuoGPT2_sft_instruct_GPT4_sharegpt.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/huatuo_medical_qa_sharegpt.German-RAG-ORPO-ShareGPT-HESSIAN-AI
German-RAG-ORPO (Odds Ratio Preference Optimization) ShareGPT-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The ORPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities.
The subsets can be for this training step are derived from 3 different sources:
SauerkrautLM Preference Datasets:
SauerkrautLM-Fermented-GER-DPO: is a specialized dataset designed for training… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-ShareGPT-HESSIAN-AI.rp-sharegptsharegpt_v3_unfiltered_cleaned_splitCode-feedback-sharegpt-renamed
