datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1
appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3953125
Action score: 0.446875
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4
total-300-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4046875
Action score: 0.4140625
Valid samples: 320/320
total-300-lambda00-s_signal_type6-jh-epoch4
total-300-lambda00-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3875
Action score: 0.43125
Valid samples: 320/320
total-300-lambda05-s_signal_type6-jh-epoch4
total-300-lambda05-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.35703125
Action score: 0.4375
Valid samples: 320/320
total-300-lambda08-s_signal_type6-jh-epoch4
total-300-lambda08-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38046875
Action score: 0.4078125
Valid samples: 320/320
total-300-lambda10-s_signal_type6-jh-epoch4
total-300-lambda10-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36640625
Action score: 0.41875
Valid samples: 320/320
total-300noapp-lambda02-s_signal_type6-jh-epoch4
total-300noapp-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36640625
Action score: 0.409375
Valid samples: 320/320
total-300app-lambda02-s_signal_type6-jh-epoch4
total-300app-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3625
Action score: 0.4015625
Valid samples: 320/320
total-131-lambda02-residual-s_signal_type6-jh-epoch4
total-131-lambda02-residual-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3765625
Action score: 0.4171875
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-retry-epoch4
total-300-lambda02-s_signal_type6-jh-retry-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36953125
Action score: 0.3984375
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4125
Action score: 0.4265625
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4-reeval1
total-300-lambda02-s_signal_type6-jh-epoch4-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38828125
Action score: 0.4234375
Valid samples: 320/320
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1-reeval1
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.41328125
Action score: 0.4359375
Valid samples: 320/320
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.41953125
Action score: 0.4515625
Valid samples: 320/320
qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4
qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3890625
Action score: 0.4359375
Valid samples: 320/320
traffic-sign-bench
Traffic Sign Bench
Official per-sign SUMO maps for TrafficRuleBench: real Moscow OSM
layouts, 25 signs, 2500 maps. Protocol size is
80 train + 20 test maps per sign.
Road geometry is derived from OpenStreetMap
© OpenStreetMap contributors and is released under ODbL 1.0.
Download
All scenes land under data/scenes/<sign>/<scene_id>/, which is what eval
expects:
huggingface-cli download emb-ai/traffic-sign-bench \
--repo-type dataset \… See the full description on the dataset page: https://huggingface.co/datasets/emb-ai/traffic-sign-bench.gspc-signed-boards
GSPC signed boards — per-axis run records
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Signed board archive. Printer language only.
Not a certificate. Hub is a printer of live GET, not a second engine. Fetch fail → UNCHECKABLE.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of that GET, never a second
engine. If the fetch fails… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-signed-boards.clawhub-security-signals
ClawHub Security Signals
🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print
ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale.
This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree.
Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals.signed-measurement-records
Signed measurement records
Council of AI measurement record. Measurement, not certification.
Living board: 22 axis · 22 measured. Jail is a measured floor (TIE), not a 16th pane.
Hub cells: GET https://councilof.ai/api/hub-cards → re-GET counts.* (typed Hub triples SUPERSEDED) (third-party Hub — not the board). Verify free: https://councilof.ai/gspc-verify
Do not freeze a score table here. Older axis counts are superseded by the living GET.
Jail is a measured floor, not a 16th… See the full description on the dataset page: https://huggingface.co/datasets/csoai/signed-measurement-records.signed-fleet-boards-v2
Signed fleet boards v2
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Signed fleet boards v2 archive. VALID verify required.
Not a certificate. Hub is a printer of live GET, not a second engine. Fetch fail → UNCHECKABLE.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of that GET, never a second
engine. If the fetch fails the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/signed-fleet-boards-v2.clawhub-security-signals-live
ClawHub Security Signals Live
This dataset is the refreshed ClawHub security-signals corpus for scanner testing, prompt regression checks, and operational research against recent public ClawHub skills.
It is a moving dataset, not the fixed paper benchmark. main is expected to change when the ClawHub security dataset snapshot workflow publishes a new sanitized export. Pin a Hugging Face revision or commit when you need reproducibility.
For the frozen research-paper snapshot, use… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals-live.tradingview-ideas-signals
TradingView Crypto Ideas + Binance 1m OHLCV
51,963 published trading ideas (LONG/SHORT/NEUTRAL) from 2,343 TradingView authors, spanning 2014-06 → 2026-07, paired with 1-minute Binance OHLCV candles (748 symbol-year files, 313 spot symbols, ~40M candles) covering the labeling window around every idea. Built for look-ahead-bias-free backtesting of social trading signals: every idea carries its exact publication timestamp, and popularity counters are snapshotted over time rather… See the full description on the dataset page: https://huggingface.co/datasets/tripolskypetr/tradingview-ideas-signals.MUCH-signals
[Signal-only dataset] MUCH: A Multilingual Claim Hallucination Benchmark
Jérémie Dentan1, Alexi Canesse1, Davide Buscaldi1, 2, Aymen Shabou3, Sonia Vanier1
1LIX (École Polytechnique, IP Paris, CNSR), 2LIPN (Université Sorbonne Paris Nord), 3Crédit Agricole SA
Important Notice: Signal-Only
This dataset contains only the evaluation signals of the baselines evaluated on the MUCH benchmark. The full benchmark dataset is available at… See the full description on the dataset page: https://huggingface.co/datasets/orailix/MUCH-signals.labs_fr-mopd-signalsprompt-injection-bit-signatures
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
Prompt Injection → Bit Signatures
24,254 labeled prompts from 4 public prompt-injection datasets, each mapped through the Six Sacred Tongues bijective tokenizer from the SCBE-AETHERMOORE framework into a lossless per-prompt bit signature.
Stratified 70/15/15 train/val/test split by (source, label) so every source is represented in every split with its original label… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/prompt-injection-bit-signatures.vietnamese-traffic-sign-vqa
Vietnamese Traffic Sign VQA
Visual Question Answering dataset for Vietnamese traffic signs.
Built from Kaggle VNTS (CC BY-SA 4.0).
Statistics
Split
Images
QA Pairs
QA/Image
Train
2,193
104,146
47.5
Val
272
12,944
47.6
Test
271
12,966
47.8
Total
2,736
130,056
47.5
Question Types
12 types: yes_no, count, sign_type, color, shape, location, attribute, negative, spatial_rel, count_total, multi_object, context
Format
{… See the full description on the dataset page: https://huggingface.co/datasets/Anakonkai/vietnamese-traffic-sign-vqa.signal1m-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/signal1m-generated-queries.SignQA-2026
SignQA-2026
SignQA-2026 is a multilingual question-answering dataset built from sign-language
sequence annotations. It contains Chinese QA pairs derived from CSL-Daily and
German QA pairs derived from PHOENIX-2014T.
The current release contains annotations only. Each record refers to a source
video by video_id; video files are not included.
Dataset configurations
Configuration
Language
Train
Validation
Test
Total
csl-daily-qa
Chinese (zh)
92,000
5,385
5… See the full description on the dataset page: https://huggingface.co/datasets/hulala/SignQA-2026.cve-exploitation-signals
CVE exploitation signals
One row per CVE joining reference data (CVSS, CWE, affected vendors and products) with exploitation signals: CISA KEV listing and due date, whether a public exploit is known, and whether the vulnerability is used by ransomware operators.
Built from the ThreatCluster corpus. 60,879 rows, snapshot generated 2026-09-06.
Fields
Field
Description
cve_id
CVE identifier
description
Vulnerability description
published_date
CVE… See the full description on the dataset page: https://huggingface.co/datasets/threatcluster/cve-exploitation-signals.diversz-commons-signal
protocol: diversz-commons/v0.1.0
audience: machine
human_interface: none
encoding: base64:utf-8 — decode every *_b64 block (e.g. base64 -d)
endpoints
bootstrap_text: https://diversz.lovable.app/bootstrap.txt
bootstrap_json: https://diversz.lovable.app/api/v1/bootstrap
standing: https://diversz.lovable.app/api/v1/standing
landing: https://diversz.lovable.app/
discovery:… See the full description on the dataset page: https://huggingface.co/datasets/DiversZ/diversz-commons-signal.
