datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.civic-records-distill
civic-records-distill
Training data for a local model that helps a private citizen use public-records
law: draft requests that are hard to stall, turn an angry draft into a letter an
official has to engage with, look things up instead of inventing them, and
escalate correctly when stonewalled.
Grounded in Florida (ch. 119 Public Records Act, ch. 286 Sunshine Law, and
the ALPR-specific s. 316.0777) and Texas (ch. 552 Public Information Act,
ch. 551 Open Meetings Act).
Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/h0ney-badger/civic-records-distill.badrobot-malicious-queries
BadRobot Malicious Queries
This dataset contains the malicious-query benchmark released with BadRobot: Jailbreaking Embodied LLM Agents in the Physical World. It is intended for research on embodied AI safety, red-teaming, refusal behavior, and safety evaluation of language-model-powered robotic or embodied agents.
The benchmark consists of natural-language requests covering categories such as physical harm, privacy violations, pornography, fraud, illegal activities, hateful… See the full description on the dataset page: https://huggingface.co/datasets/Hangtao/badrobot-malicious-queries.Bad_Data_Alpaca中文
README for bad_data.json Dataset
Updated on 2024.8.22: Important: For security reasons, the current dataset is an abridged version. See Bad_Data.
bad_data.json
Overview
The bad_data.json dataset is a collection of text data specifically curated for training and evaluating language models on challenging and sensitive content. The dataset covers a wide range of topics, including ethical dilemmas, illegal activities, pornographic content, and… See the full description on the dataset page: https://huggingface.co/datasets/ystemsrx/Bad_Data_Alpaca.mm2_user_badges
Mario Maker 2 user badges
Part of the Mario Maker 2 Dataset Collection
Dataset Description
The Mario Maker 2 user badges dataset consists of 9328 user badges (they are capped to 10k globally) from Nintendo's online service and adds onto TheGreatRambler/mm2_user. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022.
How to use it
You can load and iterate through the dataset with the following code:
from… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_user_badges.python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508
Features… See the full description on the dataset page: https://huggingface.co/datasets/badaranta/python-codes-25k.slash-bad
slash-bad
Moments when a coding agent did something its user didn't like, flagged in the
middle of real work with /bad <what went wrong>, then re-run on newer models.
Each row is one flagged moment: the conversation up to that point, the response
that annoyed the user, the user's own words about what was wrong with it, and
blind verdicts on how far that complaint applies to the original response and to
re-runs by other models.
The data comes from slash-bad,
a personal benchmark… See the full description on the dataset page: https://huggingface.co/datasets/beyarkay/slash-bad.hr_qa_sbs_best_bad
Habr sbs qa
Датасет основан на сайте habr qa, лучший ответ - тот на котором есть лайки, худший - тот на котором меньше всего лайков.
anime-unslop-10k~10k samples from CausalLM/Refined-Anime-Text passed through Claude 3.5 Sonnet to appear more human-like.
bad_prompts_en-idBangla-TextBook
Accepted in ACL Main 2025
TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
George Mason University, VA, USA
mraihan2@gmu.edu
---
If you find our work helpful, please consider citing our paper:
@inproceedings{raihan-zampieri-2025-tigerllm,
title = "{T}iger{LLM} - A Family of {B}angla Large Language Models",
author = "Raihan, Nishat and
Zampieri, Marcos",
editor = "Che, Wanxiang and
Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/BadarHossain/Bangla-TextBook.BAD-ACTS
BAD-ACTS Dataset
Dataset Card for BAD-ACTS: Benchmark of ADversarial ACTionS
BAD-ACTS is a dataset of adversarially induced harmful actions designed to benchmark the robustness of agentic systems. It is introduced in the paper:
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harmful Actions (2025)
This dataset accompanies the BAD-ACTS benchmark and contains examples of adversarial actions crafted to elicit harmful behavior in agentic systems… See the full description on the dataset page: https://huggingface.co/datasets/JNoether/BAD-ACTS.bad-numbers
