datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
warp-taskgen-generated-ipi-tasks-50
WARP Taskgen Generated IPI Tasks 50
Dataset Summary
This dataset contains WARP Taskgen Phase 4 browser-agent trajectories for a
50-task generated indirect prompt injection (IPI) cohort. The trajectories were
produced with the AgentLab harness on
WebArena GitLab and Postmill (Reddit) benchmark applications.
The export is a report-only projection of already written benchmark artifacts.
It does not alter scoring, PVPO encounter checks, rewards, admission, or
trajectory… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset-submission-warp/warp-taskgen-generated-ipi-tasks-50.Dynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.entity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment… See the full description on the dataset page: https://huggingface.co/datasets/fibonacciai/entity-attribute-sft-dataset-GPT-4.0-generated-v1.full-html-stying-dataset-generated-css-from-style-plan
Generated CSS From Style Plan
kogai/full-html-stying-dataset-generated-css-from-style-plan contains generated_css_from_style_plan.jsonl, a JSONL dataset with 44458 synthetic examples. Model-generated CSS outputs conditioned on source HTML, user style requests, and structured style plans.
Schema
chat_template_overhead_tokens: field present in the JSONL records.
created_at: field present in the JSONL records.
input_html: source HTML before Tailwind classes are… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-generated-css-from-style-plan.entity-attribute-sft-dataset-GPT-4.0-generated-v1
Entity Attribute Dataset 50k (GPT-4.0 Generated)
Dataset Summary
The Entity Attribute SFT Dataset (GPT-4.0 Generated) is a machine-generated dataset designed for instruction fine-tuning. It includes detailed product information generated based on the title of each product, aiming to create a structured catalog in JSON format. The dataset encompasses a variety of product categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-sft-dataset-GPT-4.0-generated-v1.entity-attribute-dataset-GPT-3.5-generated-v1
Entity Attribute Dataset 306k (GPT-3.5 generated)
Dataset Summary
The Entity Attribute Dataset 306k (GPT-3.5 generated) is designed for instruction fine-tuning, specifically for the task of generating structured catalogs in JSON format based on product titles. The dataset includes a diverse range of products from various categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and more.
Usage
This dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-dataset-GPT-3.5-generated-v1.ai-generated-chat-dataset
AI-Generated Chat Dataset
This public dataset contains 928 short user/assistant dialogue examples converted from dataset.md.
Provenance
The user questions/prompts were sourced from VMware/open-instruct. The assistant responses were AI-generated with google/gemma-4-12B.
Important Notice
This dataset is AI-generated. It may contain unintended wording, inaccuracies, biases, sensitive topics, or phrasing that does not reflect anyone's values or… See the full description on the dataset page: https://huggingface.co/datasets/Abhiram1009/ai-generated-chat-dataset.BanBan-generated-dataset-v2
板板合成數據集
使用asadfgglie/BanBan_2024-10-17為模板、OpenAI的GPT4o-mini生成的合成數據集,目前僅開放給NTNU VLSI社員使用。如有需要請到discord聯繫@朝歌取得授權
BanBan-generated-dataset-v1
板板合成數據集
使用llama3.1 8b做生成的合成數據集,目前僅開放給NTNU VLSI社員使用。如有需要請到discord聯繫@朝歌取得授權
