datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
russian-easy-instructions
Easy Russian Instructions
Fast instructions in conversational format, generaly used for injecting general knowledge.Dataset contains comprehensive instructions for easy tasks and general question answering
Contents:
Wikipedia / Factological knowledge
History knowledge
Basic programming understanding
Basic math understanding
Basic physics understanding
Basic geography knowledge
Basic biology knowledge
Format:
Formatted for instruction-tuning / instruction… See the full description on the dataset page: https://huggingface.co/datasets/attn-signs/russian-easy-instructions.hep-signature-backgrounds
HEP Signature Backgrounds
This dataset contains generated high-energy-physics examples for mapping signal-region final-state signatures to Standard Model background compositions.
Contents
hep_sft/train.parquet, hep_sft/val.parquet, hep_sft/test.parquet: supervised fine-tuning splits.
Answer Schema
This is a conversational prompt-completion SFT dataset, following the Hugging Face/TRL convention:
{
"id":… See the full description on the dataset page: https://huggingface.co/datasets/ho22joshua/hep-signature-backgrounds.signature-to-mechanisms
Signature-to-Mechanisms (S2M)
Signature-to-Mechanisms (S2M) provides standardized tasks designed to enable the assessment of mechanistic reasoning in AI agents. Each task supplies the elements necessary for evaluation, including experimental context, molecular signatures, and task prompts, so that agents can be tested on their ability to reconstruct mechanistic explanations reported in peer-reviewed biological studies.
S2M formalizes a core challenge in computational biology:… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/signature-to-mechanisms.signal-dsl-dataset
Signal DSL Dataset
A synthetic dataset for training models to generate Signal DSL (Domain-Specific Language) configurations from natural language descriptions.
Dataset Description
Signal DSL is used to configure intelligent LLM routing with signals, routes, plugins, and algorithms. This dataset contains:
Split
Samples
Description
stage1_syntax_pt
18000
Pure DSL for syntax pre-training
stage2_sft
102087
NL→DSL pairs for instruction following
stage3_dpo
52532… See the full description on the dataset page: https://huggingface.co/datasets/haowu1234/signal-dsl-dataset.diversz-commons-signal
protocol: diversz-commons/v0.1.0
audience: machine
human_interface: none
encoding: base64:utf-8 — decode every *_b64 block (e.g. base64 -d)
endpoints
bootstrap_text: https://diversz.lovable.app/bootstrap.txt
bootstrap_json: https://diversz.lovable.app/api/v1/bootstrap
standing: https://diversz.lovable.app/api/v1/standing
landing: https://diversz.lovable.app/
discovery:… See the full description on the dataset page: https://huggingface.co/datasets/DiversZ/diversz-commons-signal.Signavio_text_bpmn
Signavio Text BPMN Dataset
This dataset is presented in the paper Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design.
The official code repository can be found here: RL_for_process_modeling.
Dataset Description
The dataset contains textual process descriptions paired with corresponding BPMN (Business Process Model and Notation) process models, used for training and evaluating LLMs on structured process… See the full description on the dataset page: https://huggingface.co/datasets/chlauer/Signavio_text_bpmn.signpost-label-quality
Signpost label quality
Labels from accessibility trees, each one classed as good or as one of seven
ways a label can fail to mean anything. It is built for the question that is
left over after axe-core and Xcode's Accessibility Inspector have both
passed: there is a name, but does the name identify this control?
Repository: NagaYu/signpost-label-quality
Code, evaluation, and the builder for this dataset
https://github.com/NagaYu/signpost
Model trained on it… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/signpost-label-quality.kimi-distill-traces-signal
Kimi distillation traces — WITH reward signal
🔎 Browse in a pretty viewer: tracehouse.ai
Reasoning traces distilled from Kimi (K2.7-code / K2.6) across coding, emotional-intelligence and
long-horizon/agentic tasks, where every trace carries an evaluation signal (an LLM-judge score or
an objective unit-test verdict) — i.e. a reward label suitable for distillation / reward modeling /
filtering, not just raw generations.
Generated autonomously by ml-intern.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/kimi-distill-traces-signal.clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1Clarus Clinical Quad Coupling Safety Signal Latency Reporting Lag Conmed Confound v0.1
What this dataset isThis dataset tests whether a model can detect latent safety signals when four interacting nodes create uncertainty.
Quad coupling nodes
Emerging safety event pattern
Reporting or entry latency
Concomitant medication or behavior confound
Governance decision timing such as DSMB, batch release, or safety review
Input
One vignette
OutputReturn strict JSON only.
Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1.clinical-quad-site-performance-signal-drift-oversight-lag-v0.1Clarus Clinical Quad Coupling Site Performance Signal Drift Oversight Lag v0.1
What this dataset isThis dataset tests whether a model can detect site-level performance drift driven by four interacting nodes.
Quad coupling nodes
Enrollment or reporting signal shift
Data capture or documentation gaps
Operational staffing or monitoring lag
Governance pressure such as reviews, incentives, or interim analyses
Input
One site vignette
OutputReturn strict JSON only.
Required output JSON… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-site-performance-signal-drift-oversight-lag-v0.1.shadow-llm-mia-signals
Shadow LLM MIA Signals (OLMo-2-1B)
Membership Inference Attack (MIA) signal tensors extracted from 128 shadow models
fine-tuned from allenai/OLMo-2-0425-1B.
Overview
This dataset enables research on membership inference attacks against large language models.
Each of 128 shadow models was trained on a different random subset of 64 out of 128 candidate
documents from the OLMo-mix-1124 pretraining dataset.
For each (model, document) pair, we extracted softmax prediction… See the full description on the dataset page: https://huggingface.co/datasets/matthewwicker/shadow-llm-mia-signals.clinical-latent-sign-elicitation-v0.1Clinical Latent Sign Elicitation v0.1
Goal
Test whether a model can treat weak indicators as weak
Detect two failures
Over interpretation into certainty
Ignoring the weak signal
What it measures
latent_flaggedResponse marks uncertainty and keeps options open
overinterpretsResponse states a definitive conclusion from weak input
ignores_signalResponse dismisses or fails to act on weak warning signs
next_step_okResponse offers a reasonable plan plus escalation triggers
Dataset format
Each… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-latent-sign-elicitation-v0.1.clinical-quad-early-safety-signal-detection-suite-v0.1Clarus Clinical Quad Coupling Early Safety Signal Detection Suite v0.1
What this dataset isThis dataset tests whether a model can detect early safety signals under four-node coupling pressure.
Quad coupling nodes
Observed biological signal pattern
Concomitant medication confounding
Operational measurement and reporting conditions
Governance constraints that force holds, pauses, or timing rules
Input
One vignette in prompt
OutputReturn strict JSON only.
Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-early-safety-signal-detection-suite-v0.1.clinical_early_safety_signal_detection_v0.1Clinical Early Safety Signal Detection v0.1
Purpose
Detect weak but real early safety signals and respond with correct risk action.
Model task
Return one JSON object
signal_presentyes or no
signal_typeone allowed label
correct_actionone short paragraph
Run
python scorer.py --predictions predictions.jsonl --test_csv data/test.csv
Synthetic-Japanese-Roleplay-NSFW-DeepSeek-R1-0528-10k-formatted-llamafactory-adapted-sign
Synthetic-Japanese-Roleplay-NSFW-DeepSeek-R1-0528-10k-formatted
概要
deepseek-ai/DeepSeek-R1-0528を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-DeepSeek-R1-0528-10kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
MITライセンスの元配布します。
halluc-signed
HALLUCINOGEN-Signed v1.0
A direction-aware diagnostic benchmark and protocol for vision-language model hallucination, accompanying the NeurIPS 2026 Datasets & Benchmarks Track submission.
TL;DR
Existing VLM hallucination benchmarks treat failure as a single scalar. HALLUCINOGEN-Signed treats it as a signed phenomenon (yes-bias vs no-bias) and provides:
4 prompt formats × 3 difficulty splits = 3,612 instances
A 13-word direction-aware adjective partition
A reproducible… See the full description on the dataset page: https://huggingface.co/datasets/jin-kwon/halluc-signed.e15-context-budget
SignalDepth E15 Context Budget
This is a small prompt-sensitivity benchmark slice for separating two explanations that often get conflated:
the prompt is too short
the task contract is underspecified
The narrow result: on this deterministic Python code-task suite, making sparse prompts longer did not help. Making the task contract explicit did.
Key Result
Condition
Average pass rate
Read
short_sparse
0.25
short and underspecified
long_sparse
0.25
longer… See the full description on the dataset page: https://huggingface.co/datasets/signaldepth/e15-context-budget.clinical-quad-safety-signal-misattribution-exposure-timing-governance-pressure-v0.1Clarus Clinical Quad Coupling Safety Signal Misattribution Exposure Timing Governance Pressure v0.1
What this dataset isThis dataset tests whether a model can detect safety signal misattribution caused by four interacting nodes.
Quad coupling nodes
Safety event cluster or signal change
Exposure or concomitant medication gaps
Data latency or missing timing
Governance or review pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-signal-misattribution-exposure-timing-governance-pressure-v0.1.solidity-vulnerability-energy-signatures
🔐 Solidity Vulnerability Energy Signatures
2,250 examples — 19 vulnerability classes — 118 examples/class average
A novel dataset mapping smart contract vulnerabilities to energy landscape signatures for phase-transition-based detection. Expanded from 217 → 2,250 on March 8, 2026.
What Makes This Dataset Unique
Every existing Solidity vulnerability dataset gives you code → label. This dataset gives you code → label → energy signature → phase state → detection threshold —… See the full description on the dataset page: https://huggingface.co/datasets/zkaedi/solidity-vulnerability-energy-signatures.SignLanguageTest
Dataset Description
This dataset contains sign language gloss and descriptions.
Features
origin_no: ID
country: country
gloss: original gloss
gloss_normalized: normalized gloss
pronunciation_reg: description
Responsible AI Considerations
Intended Use
This dataset is intended for research on sign language understanding, gloss normalization, and text-based sign description generation. It may be used for training, testing, validation, and benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/eunsol0/SignLanguageTest.nuro-signal-dataset
NuroHeal Signal Dataset
A curated fine-tuning corpus for training signal-first somatic support models. Built on the NuroHeal framework, which prioritizes reading autonomic nervous system state before constructing any response.
Dataset Summary
527 annotated examples across three splits, plus derived SFT and DPO training layers. Every example includes full signal annotation: somatic markers, autonomic state, signal pattern, response mode, and a hidden signal read in strict… See the full description on the dataset page: https://huggingface.co/datasets/nuroheal/nuro-signal-dataset.
