datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human-vs-Ai-generated-datasetgenerated-dataset-for-VLMDynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.stocks_demo_react_agent_generated_train_datasetHarmAug_generated_dataset
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
This dataset contains generated prompts and responses using HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models.This dataset is also used for training our HarmAug Guard Model.The unsafe-score is measured by Llama-Guard-3.For rows without responses, the unsafe-score indicates the unsafeness of the prompt.For rows with responses, the unsafe-score indicates the… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/HarmAug_generated_dataset.keystroke-dataset-raw-generatedentity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment… See the full description on the dataset page: https://huggingface.co/datasets/fibonacciai/entity-attribute-sft-dataset-GPT-4.0-generated-v1.entity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-sft-dataset-GPT-4.0-generated-v1.hallucinated_answer_generated_dataset_cleanedRARE_output_and_generated_datasetsentity-attribute-dataset-GPT-3.5-generated-v1
Entity Attribute Dataset 306k (GPT-3.5 generated)
Dataset Summary
The Entity Attribute Dataset 306k (GPT-3.5 generated) is designed for instruction fine-tuning, specifically for the task of generating structured catalogs in JSON format based on product titles. The dataset includes a diverse range of products from various categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and more.
Usage
This dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-dataset-GPT-3.5-generated-v1.tokenized_generated_ar_en_th_datasets
Dataset Card for "tokenized_generated_ar_en_th_datasets"
More Information needed
ChatGPT-generated_fake_news_datasetgenerated_datasetGeneratedDatasetNEWaugmented_dataset_llm_generated_NER
📚 Augmented LLM-Generated NER Dataset for Scholarly Text
🧠 Dataset Summary
This dataset contains synthetically generated academic text tailored for Named Entity Recognition (NER) in the software engineering domain. The synthetic data augments scholarly writing using large language models (LLMs), with entity consistency maintained via token preservation.
The dataset is generated by merging and rephrasing pairs of annotated sentences from scholarly papers using… See the full description on the dataset page: https://huggingface.co/datasets/psresearch/augmented_dataset_llm_generated_NER.khmer-lesson-dataset-generatedgenerated_ar_en_th_datasets
Dataset Card for "generated_ar_en_th_datasets"
More Information needed
Generated-dataset-by-deepseek-v2.5
概要
このデータセットはnull-instruct-jaとDeepSeek-v2.5のq4を用いて合成されました。
ollamaとA5000*7基を使い2時間7分で作成されました。(使用時VRAMは合計で136GBでした。)
ライセンス
このデータセットはdeepseekのライセンスに基づきます。
deepseekのライセンス → https://github.com/deepseek-ai/DeepSeek-V2/blob/main/LICENSE-MODEL
謝辞
Deepseek-aiとnull-instruct-jaの開発者さんのGooglefanさんに感謝します。
また、機材を貸してくれているMDLの皆様にも感謝を申し上げます
futoshiki_generated_dataset_5x5_6-8GPT_Generated_Dataset_V1gpt2_generated_datasetgenerated-qa-dataset-3
Dataset Card for "generated-qa-dataset-3"
More Information needed
ai-vs-real-text-generated-datasetdataset_generated_by_teacher_meta_llama_Llama_2_13b_hf_00-19-15-04-08-25KB-sLLM-QA-Dataset-GeneratedHarmAug_generated_dataset
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
This dataset contains generated prompts and responses using HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models.This dataset is also used for training our HarmAug Guard Model.The unsafe-score is measured by Llama-Guard-3.For rows without responses, the unsafe-score indicates the unsafeness of the prompt.For rows with responses, the unsafe-score indicates the… See the full description on the dataset page: https://huggingface.co/datasets/AnonHB/HarmAug_generated_dataset.generated_group_chat_dataset_with_summarysynthetic-error-generated-spelling-correction-dataset-100kqa-dataset-generated-21020
Dataset Card for "qa-dataset-generated-21020"
More Information needed
