CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anon8231489123 /ShareGPT_Vicuna_unfilteredFurther cleaning done. Please look through the dataset and ensure that I didn't miss anything. Update: Confirmed working method for training the model: https://huggingface.co/AlekseyKorshuk/vicuna-7b/discussions/4#64346c08ef6d5abefe42c12c Two choices: Removes instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/blob/main/ShareGPT_V3_unfiltered_cleaned_split_no_imsorry.json Has instances of "I'm sorry, but":… See the full description on the dataset page: https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered.931 likes440k downloads3y agoHugging Face02ShareGPT4Video /ShareGPT4Video ShareGPT4Video 4.8M Dataset Card Dataset details Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora. sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.imagevisual-question-answering10K<n<100K204 likes11k downloads2y agoHugging Face03Aeala /ShareGPT_Vicuna_unfiltered Dataset Card This is a reupload of this dataset that was further cleaned by gozfarb. text100K<n<1M52 likes10k downloads3y agoHugging Face04OpenGVLab /ShareGPT-4ogatedtabularvisual-question-answering10K<n<100K199 likes8.5k downloads2y agoHugging Face05NickL77 /Llama3.1-8B-BaldEagle3-ShareGPT1 likes8.1k downloads1y agoHugging Face06NickL77 /llama8b-eagle-sharegpt0 likes4.7k downloads1y agoHugging Face07FreedomIntelligence /ShareGPT-4o-Image 📚 ShareGPT-4o-Image ShareGPT-4o-Image is a large-scale and high-quality image generation dataset, where all images are produced by GPT-4o’s image generation capabilities. This dataset is designed to align open multimodal models with GPT-4o’s strengths in visual content creation. It includes 45K text-to-image and 46K text-and-image-to-image samples, making it a useful resource for enhancing multimodal models in both image generation and editing tasks. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ShareGPT-4o-Image.texttext-to-image10K<n<100K102 likes4.6k downloads1y agoHugging Face08hypervariance /function-calling-sharegptThis is a dataset for finetuning models on function calling based on glaiveai/glaive-function-calling-v2. The dataset includes 86,864 examples of chats that include function calling as part of the conversation. The system prompt includes either 0, 1, or 2 functions that the assistant can use, and instructions on how the agent can use it. Changes include: Using ShareGPT format for chats Adding "function_response" as a role Removing code examples Removing examples with invalid JSON as function… See the full description on the dataset page: https://huggingface.co/datasets/hypervariance/function-calling-sharegpt.texttext-generation10K<n<100K45 likes3.7k downloads3y agoHugging Face09ShareGPTVideo /train_video_and_instruction ShareGPTVideo Training Data All dataset and models can be found at ShareGPTVideo. Contents: Train 300k video frames: contains video frames used for SFT and DPO model, which is a subset of total 900k. ActivityNet 50k + vidal 150k + webvid 100k. Train 600k video frames: contains the rest 600k frames, the total 900k frames are used for pre-training stage. If you just do finetuning using our video QA, you can just download the 300k above. 900k composition is 400k WebVid +… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPTVideo/train_video_and_instruction.videoquestion-answering34 likes3.4k downloads2y agoHugging Face10typeof /ultrachat-sharegpt-5GBtext100K<n<1M0 likes3.1k downloads3y agoHugging Face11shareAI /ShareGPT-Chinese-English-90k ShareGPT-Chinese-English-90k Bilingual Human-Machine QA Dataset A high-quality Chinese-English parallel bilingual human-machine QA dataset, covering user questions in real and complex scenarios. It is used for training high-quality dialogue models (more robust in instruction distribution than those datasets generated by repeatedly calling API interfaces to simulate machine-generated Q&A, like Moss) Features: Provides fully semantically equivalent Chinese-English parallel corpus… See the full description on the dataset page: https://huggingface.co/datasets/shareAI/ShareGPT-Chinese-English-90k.question-answering10K<n<100K288 likes3k downloads9mo agoHugging Face12hiyouga /glaive-function-calling-v2-sharegptThe glaive-function-calling-v2 dataset in sharegpt format. You can use it in LLaMA Factory by specifying --dataset glaive_toolcall_100k. texttext-generation100K<n<1M55 likes3k downloads2y agoHugging Face13ChaoticNeutrals /Creative_Writing-ShareGPTOriginal Dataset Sources: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts, https://huggingface.co/datasets/anthracite-org/nopm_claude_writing_fixed. (Thank the original dataset creators for their work.) (Nopm) Claude / (Grphye) ChatGPT-4o Syntheticly generated creative writing set's combined. Update: Used most up to date version of gryphes, chatGPT-4o set, Rejections/Slop Filtered, Min-hash Deduplication using -… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Creative_Writing-ShareGPT.text1K<n<10K18 likes2.9k downloads2y agoHugging Face14shibing624 /sharegpt_gpt4 Dataset Card Dataset Summary ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。 Languages 数据集是多语言,包括中文、英文、日文等常用语言。 Dataset Structure Data Fields The data fields are the same among all splits. conversations: a List of string . head -n 1 sharegpt_gpt4.jsonl {"conversations":[ {'from': 'human', 'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.texttext-classification100K<n<1M138 likes2.7k downloads3y agoHugging Face15NewEden /xlam-function-calling-60k-shareGPTShareGPT converted version of Salesforce/xlam-function-calling-60k text10K<n<100K0 likes2.5k downloads2y agoHugging Face16philschmid /guanaco-sharegpt-style Dataset Card for "guanaco-sharegpt-style" More Information needed text1K<n<10K49 likes2.3k downloads3y agoHugging Face17NickL77 /Qwen3-8B-BaldEagle-ShareGPT0 likes2k downloads1y agoHugging Face18Lin-Chen /ShareGPT4V News [2024/5/8] We released ShareGPT4Video, a large-scale video-caption dataset, with 40K captions annotated by GPT4V and 4.8M captions annotated by our ShareCaptioner-Video. The total videos last with 300 hours and 3000 hours separately! ShareGPT4V 1.2M Dataset Card Dataset details Dataset type: ShareGPT4V Captions 1.2M is a set of GPT4-Vision-powered multi-modal captions data. It is constructed to enhance modality alignment and fine-grained visual concept… See the full description on the dataset page: https://huggingface.co/datasets/Lin-Chen/ShareGPT4V.textvisual-question-answering1M<n<10M317 likes1.7k downloads2y agoHugging Face19Arun63 /sharegpt-quizz-generation-json-output ShareGPT-Formatted Dataset for Quizz Generation in Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate quizz in structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-quizz-generation-json-output.texttext-generationn<1K1 likes1.5k downloads2y agoHugging Face20ajibawa-2023 /Python-Code-23k-ShareGPTThis dataset is in Vicuna/ShareGPT format. There are 23000+ set of conversations. Each set having 2 conversations. Along with the Python code detailed explanation is provided. This dataset was generated using GPT-3.5, GPT-4 etc. text10K<n<100K42 likes1.2k downloads3y agoHugging Face21LNTANOooo /sharegpt52ktext10K<n<100K3 likes1.1k downloads3y agoHugging Face22RyokoAI /ShareGPT52K Dataset Card for ShareGPT52K90K Dataset Summary This dataset is a collection of approximately 52,00090,000 conversations scraped via the ShareGPT API before it was shut down. These conversations include both user prompts and responses from OpenAI's ChatGPT. This repository now contains the new 90K conversations version. The previous 52K may be found in the old/ directory. Supported Tasks and Leaderboards text-generation Languages This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/RyokoAI/ShareGPT52K.text-generation10K<n<100K364 likes1k downloads3y agoHugging Face23Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes1k downloads2y agoHugging Face24lilacai /glaive-function-calling-v2-sharegpt Dataset Card for "glaive-function-calling-v2-sharegpt" This dataset takes the glaive/glaive-function-calling-v2 dataset and formats it with ShareGPT using Lilac The accompanying notebook can be found here. The original columns "system" and "chat" still exist on the dataset. There are 4 types of roles in the ShareGPT format: system user human function call The original dataset has a column called 'chat' with the following structure: USER: Hi, I need help with calculating a tip. My… See the full description on the dataset page: https://huggingface.co/datasets/lilacai/glaive-function-calling-v2-sharegpt.text100K<n<1M31 likes961 downloads3y agoHugging Face25laion /openthoughts-4-math-qwen3-32b-7k-annotated-sharegpttext1M<n<10M0 likes864 downloads9mo agoHugging Face26openchat /openchat_sharegpt4_datasetThis repository contains cleaned and filtered ShareGPT GPT-4 data used to train OpenChat. Details can be found in the OpenChat repository. text-generation1K<n<10K173 likes748 downloads3y agoHugging Face27open-llm-leaderboard-old /details_Charlie911__MultiLora-drop-sharegpt Dataset Card for Evaluation run of Charlie911/MultiLora-drop-sharegpt Dataset automatically created during the evaluation run of model Charlie911/MultiLora-drop-sharegpt on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Charlie911__MultiLora-drop-sharegpt.0 likes715 downloads3y agoHugging Face28thomas-yanxin /MT-SFT-ShareGPT   MT-SFT-ShareGPT   💻 Github Repo • 🤗 HuggingFace • 🤖 ModelScope Introduction Data has always been an important part of advancing large language models forward. Based on this, we have collected dozens of high-quality open source datasets from the open source community, with a total data volume of 20 M. After some cleaning actions, we have open sourced a set of high-quality datasets for fine-tuning the instructions of the… See the full description on the dataset page: https://huggingface.co/datasets/thomas-yanxin/MT-SFT-ShareGPT.question-answering1M<n<10M12 likes682 downloads2y agoHugging Face29jzsues /ShareGPT4V-Rebuild Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/jzsues/ShareGPT4V-Rebuild.text100K<n<1M0 likes660 downloads2y agoHugging Face30MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1 in ShareGPT Format This dataset is a conversion of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 into the ShareGPT format while preserving the original splits and columns. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "messages": [ {"role": "user", "content": "User message"}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT.text10M<n<100M41 likes598 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.