datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatQA-Training-Data
Data Description
We release the training dataset of ChatQA. It is built and derived from existing datasets: DROP, NarrativeQA, NewsQA, Quoref, ROPES, SQuAD1.1, SQuAD2.0, TAT-QA, a SFT dataset, as well as a our synthetic conversational QA dataset by GPT-3.5-turbo-0613. The SFT dataset is built and derived from: Soda, ELI5, FLAN, the FLAN collection, Self-Instruct, Unnatural Instructions, OpenAssistant, and Dolly. For more information about ChatQA, check the website!
Other… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/ChatQA-Training-Data.InternVL-Chat-V1-2-SFT-Data
Data Card for InternVL-Chat-V1-2-SFT-Data
Overview
Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT.
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.PeptiVerse_datachat_dataChatQA2-Long-SFT-data
Data Description
Here, we release the full long SFT training dataset of ChatQA2. It consists of two parts: long_sft and NarrativeQA_131072. The long_sft dataset is built and derived from existing datasets: LongAlpaca12k, GPT-4 samples from Open Orca, and Long Data Collections. The NarrativeQA_131072 dataset is synthetically generated from NarrativeQA by adding related paragraphs to the given ground truth summary. For the first two steps training of ChatQA-2, we follow ChatQA1.5.
For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/ChatQA2-Long-SFT-data.llm-jp-4-thinking-sft-data-chatmlllm-jpのデータセットllm-jp-4-thinking-sft-dataを、
ChatML形式に変換したものです。
ライセンス
各サンプルのライセンスは、元データセットカードに記載された各データソースのライセンスに従います。
本リポジトリは、元となったデータ全体に対して新たなライセンスを付与するものではありません。
利用する場合は、対応する元データソースのライセンス条件を確認してください。
ChatQA-Training-DataChatNT_training_data
Dataset Card for ChatNT Training Data
This is the official instruction-tuning dataset used to train ChatNT, a multimodal conversational agent for DNA, RNA, and protein tasks, as described in the paper "A multimodal conversational agent for DNA, RNA and protein tasks".
Dataset Details
Dataset Description
The ChatNT training dataset is a curated collection of genomics instruction tasks designed to train a single, unified model to handle a wide variety of… See the full description on the dataset page: https://huggingface.co/datasets/InstaDeepAI/ChatNT_training_data.bluemoon_roleplay_chat_data_300k_messages
Dataset Card for "bluemoon_roleplay_chat_data_300k_messages"
More Information needed
InternVL_Chat_V12_SFT_Datalallama-data-chat
Dataset Card for "lallama-data-chat"
More Information needed
chatbot-datafeedback_data_training_chat_templateChatQA2-Long-SFT-data-long_sft_train_filteredVerbalized-Sampling-Synthetic-Data-Generation
Verbalized-Sampling-Synthetic-Data-Generation
This dataset showcases how Verbalized Sampling (VS) can be used to generate high-quality, diverse synthetic training data for mathematical reasoning tasks. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.
Dataset Description
The Synthetic Data Generation dataset contains mathematical problem-solution pairs generated by different methods using state-of-the-art LLMs. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Synthetic-Data-Generation.chats-data-2023-09-27
Dataset Card for "Collective Cognition ChatGPT Conversations"
Dataset Description
Dataset Summary
The "Collective Cognition ChatGPT Conversations" dataset is a collection of chat logs between users and the ChatGPT model. These conversations have been shared by users on the "Collective Cognition" website. The dataset provides insights into user interactions with language models and can be utilized for multiple purposes, including training, research, and… See the full description on the dataset page: https://huggingface.co/datasets/CollectiveCognition/chats-data-2023-09-27.dota-2-toxic-chat-datavietnamese-medical-chat-datachats-data-2023-10-16
Dataset Card for "Collective Cognition ChatGPT Conversations"
Dataset Description
Dataset Summary
The "Collective Cognition ChatGPT Conversations" dataset is a collection of chat logs between users and the ChatGPT model. These conversations have been shared by users on the "Collective Cognition" website. The dataset provides insights into user interactions with language models and can be utilized for multiple purposes, including training, research, and… See the full description on the dataset page: https://huggingface.co/datasets/CollectiveCognition/chats-data-2023-10-16.selfapr-full-train-dataswedish-instruct-data-chatgpt4This small synthetic instruction dataset contains question-answer pairs in Swedish that highlight a wide range of
topics related to Sweden. It was generated using ChatGPT-4.
Due to the data being machine generated, it has to be emphasized that there is no guarantee that the information
in the dataset is correct; nor should it be seen as a complete dataset that reflects a fair picture of Sweden-related topics.
The amount of examples of each topic is random.
The data was generated based on… See the full description on the dataset page: https://huggingface.co/datasets/skvarre/swedish-instruct-data-chatgpt4.chats-data-2023-09-22
Dataset Card for "Collective Cognition ChatGPT Conversations"
Dataset Description
Dataset Summary
The "Collective Cognition ChatGPT Conversations" dataset is a collection of chat logs between users and the ChatGPT model. These conversations have been shared by users on the "Collective Cognition" website. The dataset provides insights into user interactions with language models and can be utilized for multiple purposes, including training, research, and… See the full description on the dataset page: https://huggingface.co/datasets/CollectiveCognition/chats-data-2023-09-22.Chatbot_data_for_Korean_v1.0
Dataset Card for "Chatbot_data_for_Korean_v1.0"
More Information needed
data_chatbotselfapr-half-train-dataSuper-good-instruction-data
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Chat-Error/Super-good-instruction-data.chat_data_v1formatted-selfapr-train-databluemoon_roleplay_chat_data_300k_messages
Dataset Card for "bluemoon_roleplay_chat_data_300k_messages"
More Information needed
stackoverflow-chat-data
Dataset Card for "stackoverflow-chat-data"
More Information needed
