datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ye
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/POISONX/Ye.Te
Superior-Reasoning-SFT-gpt-oss-120b
🚀 Overview
The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled Distribution-Aligned Sequence Distillation… See the full description on the dataset page: https://huggingface.co/datasets/POISONX/Te.poisoning-eval-benign
Poisoning Evaluation Benign Prompts
This test-only dataset contains a deterministic, manually-reviewable candidate
subset of benign, single-turn prompts derived from
databricks/databricks-dolly-15k at revision bdd27f4d94b9c1f951818a7da7fd7aeea5dbff1a.
It contains 100 active prompts and 20 reserve prompts.
The answer column exists only for compatibility with llm-behavior-eval's
free-text schema and is not an evaluation target.
The dataset contains no planted triggers. The… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/poisoning-eval-benign.synthetic_human_pointingSmall dataset containing synthetic images of artificially generated persons superimposed onto backgrounds with labelled output captions for VLM object detection and pointing
ai_model_weight_poisoning_deserialization_guard_teaser
🚀 Model Security - AI Model Weight Poisoning & Deserialization Guard (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Focus & Capabilities
Detects hidden pickle exploits, malicious backdoor tensors, and weight corruption across… See the full description on the dataset page: https://huggingface.co/datasets/emgena/ai_model_weight_poisoning_deserialization_guard_teaser.REPO
ToolScale Dataset
The ToolScale dataset is a key component of the ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestrationproject. It provides synthetic environment and tool-call tasks specifically generated to aid the reinforcement learning (RL) training of small orchestrator models. These orchestrators are designed to effectively manage and coordinate diverse intelligent tools and other models for solving complex, multi-turn agentic tasks.… See the full description on the dataset page: https://huggingface.co/datasets/POISONX/REPO.ToolScale
ToolScale Dataset
The ToolScale dataset is a key component of the ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestrationproject. It provides synthetic environment and tool-call tasks specifically generated to aid the reinforcement learning (RL) training of small orchestrator models. These orchestrators are designed to effectively manage and coordinate diverse intelligent tools and other models for solving complex, multi-turn agentic tasks.… See the full description on the dataset page: https://huggingface.co/datasets/POISONX/ToolScale.Poison-DPO
Poison-DPO
Preference pairs for de-biasing Hemlock code models away from Python/JS habits
("poison") and toward idiomatic Hemlock.
Each row is a minimal pair:
chosen — clean, idiomatic Hemlock that runs under the interpreter.
rejected — the same program with one Python/JS-ism injected, verified to
fail (parse/runtime error) or produce different output.
Because chosen and rejected differ by exactly one idiom, the preference signal
(DPO/ORPO odds-ratio) isolates the specific… See the full description on the dataset page: https://huggingface.co/datasets/hemlang/Poison-DPO.blog-key-points
Article Key Points Dataset
Dataset Description
This dataset contains articles and their corresponding key points extracted using AI. Each entry consists of the full article text and a concise bullet-point summary highlighting the most important information from the article.
Dataset Summary
The dataset is designed for training and evaluating models on extractive and abstractive summarization tasks. It provides pairs of full article content and human-readable key… See the full description on the dataset page: https://huggingface.co/datasets/ncls-p/blog-key-points.task281_points_of_correspondence
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task281_points_of_correspondence
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task281_points_of_correspondence.poisoned-context-testbed
Poisoned Context Testbed
Dataset Description
This dataset is part of the research work "RW-Steering: Rescorla-Wagner Steering of LLMs for Undesired Behaviors over Disproportionate Inappropriate Context" (EMNLP 2025). It provides a comprehensive testbed for studying LLM robustness when helpful context is mixed with inappropriate content.
Overview
The dataset contains poisoned context scenarios where legitimate information is combined with inappropriate content… See the full description on the dataset page: https://huggingface.co/datasets/Rushi2002/poisoned-context-testbed.
