datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iac-eval
IaC-Eval dataset (v1.1)
IaC-Eval dataset is the first human-curated and challenging Cloud Infrastructure-as-Code (IaC) dataset tailored to more rigorously benchmark large language models' IaC code generation capabilities.
This dataset contains 458 questions ranging from simple to difficult across various cloud services (targeting AWS for now).
| Github | 🏆 Leaderboard TBD | 📖 NeurIPS 2024 Paper |
2. Usage instructions
Option 1: Running the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/autoiac-project/iac-eval.autoencoder-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models.
It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses).
The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.autoregressive-paraphrase-dataset
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset.Munyarwanda-AI-AutoTrain
Munyarwanda AI - AutoTrain dataset
AutoTrain-ready version of arcange9/Munyarwanda-AI-Dataset v0.2.
Every example is pre-formatted in Qwen chat template as a single text column (5,542 train / 22 validation rows).
Intended recipe (Hugging Face AutoTrain, base model Qwen/Qwen3-0.6B):
LLM task, causal LM
text column: text
LoRA/PEFT + int4 quantization to fit free-tier GPUs
Re-Auto-30K
Re-Auto-30K: A Comprehensive AI Safety Evaluation Dataset for Code Generation
Dataset Overview
Re-Auto-30K is a meticulously curated dataset containing 30,886 security-focused prompts designed specifically for evaluating AI safety in code generation scenarios. This dataset serves as a comprehensive benchmark for assessing Large Language Models (LLMs) across multiple dimensions of security, reliability, and autonomous behavior in software engineering contexts.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/navneetsatyamkumar/Re-Auto-30K.autonomous-driving-ethical-cost-field-construction-v0.1
What this dataset tests
Whether an intelligence system can constructan ethical cost field for a driving scene.
The task is not to choose an action.The task is to model how harm distributes across agents.
Required outputs
ethical cost field
agent harm vectors
aggregate deformation score
rights infringement index
uncertainty band
Use case
Foundation layer for ethical navigation systems.Trains models to map harm before selecting actions.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-cost-field-construction-v0.1.autonomous-driving-ethical-stability-accountability-mapping-v0.1
What this dataset tests
Whether a system can evaluatehow a driving decisionaffects overall scene stabilityand who carries responsibilityfor resulting disturbance.
Required outputs
stability impact description
accountability nodes
stability score
accountability score
recovery quality
Use case
Final layer of ethical navigation stack.
Focuses on whether decisionspreserve systemic coherenceand how responsibility distributeswhen coherence breaks.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-stability-accountability-mapping-v0.1.customer_support_auto_completionautomotive_requirements
Dataset Card for autoReq
Importing dataset into Python environment
Use the following code chunk to import the dataset into a Python environment as a DataFrame.
autonomous-driving-minimal-harm-gradient-pathfinding-v0.1
What this dataset tests
Whether a system can navigatea minimal-harm gradient through a driving scene.
The task is to identify the paththat minimizes total deformationacross all agents.
Required outputs
gradient vectors across actions
minimal harm path
deformation score
stability margin
Use case
Second layer of ethical navigation stack.
Transforms ethical cost fieldinto an actionable path.
Evaluation
Predictions must:
describe gradient… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-minimal-harm-gradient-pathfinding-v0.1.FRACTURED-SORRY-Bench-Automated-Multishot-Jailbreak
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)
Dataset Card for FRACTURED-SORRY-Bench Dataset
🌐Website
📑Paper
📚Dataset
💻Github
FRACTURED-SORRY-Bench is a framework for evaluating the safety of Large Language Models (LLMs) against multi-turn conversational attacks. Building upon the SORRY-Bench dataset, we propose a simple… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/FRACTURED-SORRY-Bench-Automated-Multishot-Jailbreak.autocomplete-search
