datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-oss120b-generated-perfectblendgpt-oss120b-generated-magpie-1m-v0.1gpt_generated_10k
Dataset Card for "gpt_generated_10k"
More Information needed
python-docstring-human-gpt-generated-mixgenerated_conversation_data_GPT_oss
ML Engineer Thinking Process Dataset
Dataset Overview
This dataset is a collection of structured, multi-turn dialogues designed to capture the thought process and communication style of an expert Machine Learning Engineer. It was created to fine-tune language models to not only provide technically accurate answers but also to articulate the reasoning behind them, mimicking how a real engineer approaches problems.
The dialogues are formatted in a conversational style… See the full description on the dataset page: https://huggingface.co/datasets/aliRafik/generated_conversation_data_GPT_oss.openrubrics-v2-generated-rubrics-gpt-54-complete-fixed-judgementsentity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-sft-dataset-GPT-4.0-generated-v1.controlled-generated-convos-gpt-4.1-mini
Controlled Generated Conversations: gpt-4.1-mini
Dataset Description
This dataset contains synthetic customer support conversations generated using gpt-4.1-mini as part of research on cross-lingual stability of LLM judges. The conversations are designed for evaluating how well language models maintain consistent performance across different languages, with a focus on Finno-Ugric languages (Estonian, Finnish, Hungarian) and English.
Dataset Summary
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/isaacchung/controlled-generated-convos-gpt-4.1-mini.entity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment… See the full description on the dataset page: https://huggingface.co/datasets/fibonacciai/entity-attribute-sft-dataset-GPT-4.0-generated-v1.vizuara_gpt_generated_textopenrubrics-v2-generated-rubrics-gpt-54entity-attribute-dataset-GPT-3.5-generated-v1
Entity Attribute Dataset 306k (GPT-3.5 generated)
Dataset Summary
The Entity Attribute Dataset 306k (GPT-3.5 generated) is designed for instruction fine-tuning, specifically for the task of generating structured catalogs in JSON format based on product titles. The dataset includes a diverse range of products from various categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and more.
Usage
This dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-dataset-GPT-3.5-generated-v1.Optimizer-gptgeneratedpandasopenrubrics-v2-generated-rubrics-gpt-54-completegenerated_chat_0_4M_sharegpt_system_human_gptGPT_Generated_Dataset_V1gpt-generated-news-sentences
Dataset Card
This dataset was created solely for the purpose of code testing.
This dataset was generated from prompting chatGPT to create sample pieces of news setences according to a topic.
Sample prompt: "generate 50 sentences on the topic of "very recent breaking news on wars and conflicts events" with some sample location names. One example: "a missile struck near a residential building in Kiev last night, Russia denied Ukraine's accusations of attacking non-military targets""… See the full description on the dataset page: https://huggingface.co/datasets/joshuapsa/gpt-generated-news-sentences.wildchat_aqa_with_embedding_and_gpt_generated_querygpt-generated-news-paragraphs-v1.0
Dataset Card for "gpt-generated-news-paragraphs"
More Information needed
GPT_Generated_Dataset_V2_1000GPT_Generated_Dataset_V2_500juraj-juraj-python-docstring-human-gpt-generated-mixGPT_Generated_Dataset_V2_1500GPT_Generated_Dataset_Fold1_2000GPT_Generated_Dataset_Fold2_2000GPT_Generated_Dataset_Fold4_2000GPT_Generated_Dataset_Fold5_2000gpt-generated-news-paragraphs-v1.1
Dataset Card for "gpt-generated-news-paragraphs-v1.1"
This dataset was created solely for the purpose of code testing.
This dataset was generated from prompting chatGPT to create sample pieces of news setences according to a topic.
Sample prompt: "generate 50 paragraphs on the topic of "very recent breaking news on wars and conflicts events" with some sample location names. One example: "a missile struck near a residential building in Kiev last night. Russia denied Ukraine's… See the full description on the dataset page: https://huggingface.co/datasets/joshuapsa/gpt-generated-news-paragraphs-v1.1.GPT_Generated_Dataset_500GPT_Generated_Dataset_1000
