datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clawhub-security-signals
ClawHub Security Signals
🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print
ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale.
This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree.
Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals.tradingview-ideas-signals
TradingView Crypto Ideas + Binance 1m OHLCV
51,963 published trading ideas (LONG/SHORT/NEUTRAL) from 2,343 TradingView authors, spanning 2014-06 → 2026-07, paired with 1-minute Binance OHLCV candles (748 symbol-year files, 313 spot symbols, ~40M candles) covering the labeling window around every idea. Built for look-ahead-bias-free backtesting of social trading signals: every idea carries its exact publication timestamp, and popularity counters are snapshotted over time rather… See the full description on the dataset page: https://huggingface.co/datasets/tripolskypetr/tradingview-ideas-signals.clawhub-security-signals-live
ClawHub Security Signals Live
This dataset is the refreshed ClawHub security-signals corpus for scanner testing, prompt regression checks, and operational research against recent public ClawHub skills.
It is a moving dataset, not the fixed paper benchmark. main is expected to change when the ClawHub security dataset snapshot workflow publishes a new sanitized export. Pin a Hugging Face revision or commit when you need reproducibility.
For the frozen research-paper snapshot, use… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals-live.MUCH-signals
[Signal-only dataset] MUCH: A Multilingual Claim Hallucination Benchmark
Jérémie Dentan1, Alexi Canesse1, Davide Buscaldi1, 2, Aymen Shabou3, Sonia Vanier1
1LIX (École Polytechnique, IP Paris, CNSR), 2LIPN (Université Sorbonne Paris Nord), 3Crédit Agricole SA
Important Notice: Signal-Only
This dataset contains only the evaluation signals of the baselines evaluated on the MUCH benchmark. The full benchmark dataset is available at… See the full description on the dataset page: https://huggingface.co/datasets/orailix/MUCH-signals.labs_fr-mopd-signalscve-exploitation-signals
CVE exploitation signals
One row per CVE joining reference data (CVSS, CWE, affected vendors and products) with exploitation signals: CISA KEV listing and due date, whether a public exploit is known, and whether the vulnerability is used by ransomware operators.
Built from the ThreatCluster corpus. 60,879 rows, snapshot generated 2026-09-06.
Fields
Field
Description
cve_id
CVE identifier
description
Vulnerability description
published_date
CVE… See the full description on the dataset page: https://huggingface.co/datasets/threatcluster/cve-exploitation-signals.repro-evaluating-llms-comparative-signals-traces
Agent traces
Agent sessions published from a Trackio Logbook.
clawhub-security-signals
ClawHub Security Signals
🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print
ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale.
This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree.
This Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/sky-meilin/clawhub-security-signals.clawhub-security-signals
ClawHub Security Signals
🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print
ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale.
This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree.
Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/aicreatemo/clawhub-security-signals.clawhub-security-signals
ClawHub Security Signals
🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print
ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale.
This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree.
This Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/rumeshprasanga6/clawhub-security-signals.clawhub-security-signals
ClawHub Security Signals
🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print
ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale.
This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree.
This Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hectorize/clawhub-security-signals.shadow-llm-mia-signals
Shadow LLM MIA Signals (OLMo-2-1B)
Membership Inference Attack (MIA) signal tensors extracted from 128 shadow models
fine-tuned from allenai/OLMo-2-0425-1B.
Overview
This dataset enables research on membership inference attacks against large language models.
Each of 128 shadow models was trained on a different random subset of 64 out of 128 candidate
documents from the OLMo-mix-1124 pretraining dataset.
For each (model, document) pair, we extracted softmax prediction… See the full description on the dataset page: https://huggingface.co/datasets/matthewwicker/shadow-llm-mia-signals.trading-signals
Trading Signals Dataset
Описание
Этот датасет содержит данные о торговых сигналах для криптовалют, включающие взаимодействия между системными промптами, пользовательскими сообщениями и выходными данными языковых моделей для торговли криптовалютой.
Структура данных
Датасет расположен в директории dump/outline/ и содержит 188,975 файлов в формате Markdown, организованных в папки с уникальными идентификаторами сессий.
Структура папок
dump/outline/
├──… See the full description on the dataset page: https://huggingface.co/datasets/tripolskypetr/trading-signals.han-human-cognitive-load-signals-v1
Human Cognitive Load Signals Dataset
This dataset captures indicators of human cognitive load
during interactions with humanoid AI systems.
It helps models adapt responses based on mental effort
and attention levels of humans.
Use Cases
Adaptive interaction pacing
Mental workload estimation
Human-centered AI optimization
Fields
task_type
cognitive_load_level
human_behavior_signal
interaction_duration
Part of
Humanoid Network (HAN)
License… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-human-cognitive-load-signals-v1.clawhub-security-signals
ClawHub Security Signals
🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print
ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale.
This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree.
This Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/Victormart43210/clawhub-security-signals.han-humanoid-perception-anomaly-signals-v1
Humanoid Perception Anomaly Signals (HPAS)
Abstract
HPAS contains structured anomaly signals detected
from humanoid sensor fusion systems during operation.
The dataset supports anomaly detection, sensor
confidence modeling, and environmental uncertainty research.
Fields
timestamp
sensor_type
raw_signal_variance
anomaly_score
confidence_level
environment_context
Intended Research
Sensor anomaly detection
Confidence calibration
Environmental… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-perception-anomaly-signals-v1.han-humanoid-reputation-signals-v1
Humanoid Reputation Signal Dataset
This dataset aggregates signals used to calculate
the reputation of humanoid AI agents over time.
It supports transparent, trust-based evaluation
within decentralized humanoid networks.
Use Cases
Reputation scoring
Trust assessment
Network governance
Fields
agent_id
signal_type
signal_weight
timestamp
Part of
Humanoid Network (HAN)
License
MIT
defendable-honey-signals-v0.1
Defendable Honey Signals · v0.1
Tribunal-graded HONEY signals extracted from the Defendable Federal Demand Intelligence corpus + the Kimi deep-research intake. Every row carries a footnote-anchored citation from the source research file.
Tribunal begins before training. No proof, no honey.
What is a HONEY signal?
A quantified market signal — demand, pricing, or market gap — that survived Tribunal grading at the HIGHEST trust tier: cite-anchored, source-verifiable, ready… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-honey-signals-v0.1.signals_with_completionssignals-n-systemshuman-feedback-signals-json
Human Feedback Signals Dataset
Dataset mapping human non-verbal signals to feedback labels
for social robots.
recall_intelligence_signals
Recall Intelligence Signals v1
Public recall and enforcement records prepared by SingleFoundry. Built for institutional_research buyers, this SingleFoundry data product packages 1 validated records with governed source evidence, quality checks, audit traceability, and ready-to-use delivery metadata.
Dataset Files
data/recall-intelligence-signals-extracted-records-csv.csv: latest validated SingleFoundry CSV package.
singlefoundry-metadata.json: release metadata… See the full description on the dataset page: https://huggingface.co/datasets/singlefoundry/recall_intelligence_signals.queue-management-signals-json
Queue Management Signals Dataset
JSON dataset mapping queue conditions to management actions
for service robots.
ai-signals-2026
