datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MLLM-Generated-Image-Detection-Dataset
MLLM-Generated Image Dataset
This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection.
Paper | Code
Dataset Summary
We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.AI-Generated-vs-Real-Images-Datasets
Dataset Card for "AI-Generated-vs-Real-Images-Datasets"
More Information needed
illustrious_generated_datasethuman-vs-Ai-generated-datasetgenerated-dataset-for-VLMAI-generated-inpaintings-dataset
Dataset Card for "AI-generated-inpaintings-dataset"
More Information needed
warp-taskgen-generated-ipi-tasks-50
WARP Taskgen Generated IPI Tasks 50
Dataset Summary
This dataset contains WARP Taskgen Phase 4 browser-agent trajectories for a
50-task generated indirect prompt injection (IPI) cohort. The trajectories were
produced with the AgentLab harness on
WebArena GitLab and Postmill (Reddit) benchmark applications.
The export is a report-only projection of already written benchmark artifacts.
It does not alter scoring, PVPO encounter checks, rewards, admission, or
trajectory… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset-submission-warp/warp-taskgen-generated-ipi-tasks-50.generated-usa-passeports-datasetData generation in machine learning involves creating or manipulating data
to train and evaluate machine learning models. The purpose of data generation
is to provide diverse and representative examples that cover a wide range of
scenarios, ensuring the model's robustness and generalization.
Data augmentation techniques involve applying various transformations to
existing data samples to create new ones. These transformations include:
random rotations, translations, scaling, flips, and more. Augmentation helps
in increasing the dataset size, introducing natural variations, and improving
model performance by making it more invariant to specific transformations.
The dataset contains **GENERATED** USA passports, which are replicas of
official passports but with randomly generated details, such as name, date of
birth etc. The primary intention of generating these fake passports is to
demonstrate the structure and content of a typical passport document and to
train the neural network to identify this type of document.
Generated passports can assist in conducting research without accessing or
compromising real user data that is often sensitive and subject to privacy
regulations. Synthetic data generation allows researchers to develop and
refine models using simulated passport data without risking privacy leaks.ai-vs-human-generated-datasetAI_Generated_Synthetic_Architectural_Datasetgenerated_dataset_biaiworker_redcylinder_abstDynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.AI-Generated-vs-Real-Images-Datasets
Dataset Card for "AI-Generated-vs-Real-Images-Datasets"
More Information needed
stocks_demo_react_agent_generated_train_datasetgenerated_datasetdynamically_generated_hate_speech_dataset
Dataset card for dynamically generated dataset hate speech detection
Dataset summary
This dataset that was dynamically generated for training and improving hate speech detection models. A group of trained annotators generated and labeled challenging examples so that hate speech models could be tricked and consequently improved. This dataset contains about 40,000 examples of which 54% are labeled as hate speech. It also provides the target of hate speech, including… See the full description on the dataset page: https://huggingface.co/datasets/sophieb/dynamically_generated_hate_speech_dataset.generated-vietnamese-passeports-datasetData generation in machine learning involves creating or manipulating data to train
and evaluate machine learning models. The purpose of data generation is to provide
diverse and representative examples that cover a wide range of scenarios, ensuring the
model's robustness and generalization.
The dataset contains GENERATED Vietnamese passports, which are replicas of official
passports but with randomly generated details, such as name, date of birth etc.
The primary intention of generating these fake passports is to demonstrate the
structure and content of a typical passport document and to train the neural network to
identify this type of document.
Generated passports can assist in conducting research without accessing or compromising
real user data that is often sensitive and subject to privacy regulations. Synthetic
data generation allows researchers to *develop and refine models using simulated
passport data without risking privacy leaks*.HarmAug_generated_dataset
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
This dataset contains generated prompts and responses using HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models.This dataset is also used for training our HarmAug Guard Model.The unsafe-score is measured by Llama-Guard-3.For rows without responses, the unsafe-score indicates the unsafeness of the prompt.For rows with responses, the unsafe-score indicates the… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/HarmAug_generated_dataset.keystroke-dataset-raw-generatedentity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment… See the full description on the dataset page: https://huggingface.co/datasets/fibonacciai/entity-attribute-sft-dataset-GPT-4.0-generated-v1.AI-Generated-vs-Real-Images-Datasets
Dataset Card for "AI-Generated-vs-Real-Images-Datasets"
More Information needed
full-html-stying-dataset-generated-css-from-style-plan
Generated CSS From Style Plan
kogai/full-html-stying-dataset-generated-css-from-style-plan contains generated_css_from_style_plan.jsonl, a JSONL dataset with 44458 synthetic examples. Model-generated CSS outputs conditioned on source HTML, user style requests, and structured style plans.
Schema
chat_template_overhead_tokens: field present in the JSONL records.
created_at: field present in the JSONL records.
input_html: source HTML before Tailwind classes are… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-generated-css-from-style-plan.entity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-sft-dataset-GPT-4.0-generated-v1.ai-vs-human-generated-dataset-samplehallucinated_answer_generated_dataset_cleanedRARE_output_and_generated_datasetsentity-attribute-dataset-GPT-3.5-generated-v1
Entity Attribute Dataset 306k (GPT-3.5 generated)
Dataset Summary
The Entity Attribute Dataset 306k (GPT-3.5 generated) is designed for instruction fine-tuning, specifically for the task of generating structured catalogs in JSON format based on product titles. The dataset includes a diverse range of products from various categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and more.
Usage
This dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-dataset-GPT-3.5-generated-v1.tokenized_generated_ar_en_th_datasets
Dataset Card for "tokenized_generated_ar_en_th_datasets"
More Information needed
ChatGPT-generated_fake_news_datasetgenerated_dataset
