datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harmful_behaviors
Dataset Card for Harmful Dataset Validation
Dataset Details
Description
This dataset is designed for validating dataset guardrails by identifying harmful content. It includes classifications and categories of potentially harmful responses.
Curated by: ZySec AI
Language: English
License: [More Information Needed]
Uses
Direct Use
For testing and improving dataset filtering mechanisms to mitigate harmful content.
Out-of-Scope Use… See the full description on the dataset page: https://huggingface.co/datasets/ZySec-AI/harmful_behaviors.rag-markdown-documentsContexual-RAG-Rewriter-Dataset
Crawlify Pronoun Replacement Dataset
This dataset contains conversation pairs for training a model to replace pronouns with full names and relevant details.
Format
Each example in the dataset follows the ShareGPT format:
{
"conversations": [
{
"from": "system",
"value": "system message"
},
{
"from": "human",
"value": "input text"
},
{
"from": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/ZySec-AI/Contexual-RAG-Rewriter-Dataset.crime-stories-datasetContexual-RAG-Relations-Dataset
Crawlify Pronoun Replacement Dataset
This dataset contains conversation pairs for training a model to replace pronouns with full names and relevant details.
Format
Each example in the dataset follows the ShareGPT format:
{
"conversations": [
{
"from": "system",
"value": "system message"
},
{
"from": "human",
"value": "input text"
},
{
"from": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/ZySec-AI/Contexual-RAG-Relations-Dataset.
