datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
german-canary-asr-0324
Dataset Beschreibung
Allgemeine Informationen
Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 16.1, Voxpopuli und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert.
Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-canary-asr-0324.polymarket-canary-tape
polymarket-canary-tape
Continuous tape from Scribe (Bot E / bot_e_recorder), a single always-on VPS node. Captures co-located CEX trades and Polymarket market-channel WebSocket events over a fixed UTC window for microstructure and lead-lag research.
Where this came from: released alongside polymarket-bot-lab
(11 open-source Polymarket trading bot candidates, Apache-2.0) by the team behind
OracleMangle, which builds dispute-risk scoring for
prediction-market questions. Both… See the full description on the dataset page: https://huggingface.co/datasets/oraclemangle/polymarket-canary-tape.CanaryAura
Dataset Card for "Canary Aura"
This is a dataset for...
MIRAGE-CanaryDocs
MIRAGE CanaryDocs
MIRAGE CanaryDocs is an English synthetic enterprise-document dataset for structured privacy-unit,
canary, and ordered multi-chunk evaluation. It is the companion dataset for the EMNLP 2026 paper
When Metadata Remembers: Ordered Provenance Enables Document-Level Embedding Inversion.
Project documentation and schemas are also available in the
MIRAGE GitHub repository.
Dataset summary
The dataset contains complete synthetic documents, ordered token… See the full description on the dataset page: https://huggingface.co/datasets/LevenKoko/MIRAGE-CanaryDocs.terminal_bench_2_Qwen3_32B_canary_ghdevenron_canary
CanaryBench-Enron
Frequency-aware canary injection benchmark for auditing memorization
in finetuned language models, built on the Enron email corpus.
Dataset Description
This dataset is part of CanaryBench, a benchmark for evaluating
memorization in finetuned language models across repetition tiers
and privacy regimes.
Frequency tiers: 1×, 10×, 50×
Domain: Email (Enron corpus)
Member canaries: 770
Reference canaries: 1000
Files… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/enron_canary.browsecomp-ctxgraph-30b-rl-fusedent-canary-v1
browsecomp-ctxgraph-30b-rl-fusedent-canary-v1
Fused-kernels entropy canary (job vista:820770, 2026-07-10). First run ever with a working entropy bonus: entropy_coeff=0.005 via use_fused_kernels=True (FusedLinearForPPO). Step-1 actor update survived (787422 OOMed here). actor/entropy=0.346, actor/entropy_loss=0.00173, step time 50.5 min, mem peak 101.4GB, val_before_train task_reward=0.453 (max_turn=100/max_session=10), train step-1 task_reward=0.553, aborted_ratio 19%. Judge… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-fusedent-canary-v1.qwen3p6-27b-canary-rolloutsslimder-qwen38-s2-canary-20260831ds-canary-statbelief-state-probes-canary-v1
belief-state-probes-canary-v1
COMPLETE E2E canary (73 min, A40): wandb OK; mess3 paper-exact 50k steps at Bayes floor, probe R2 0.986 (held-out-beliefs 0.985), simplex PNG; rrxor concat R2 0.72 > final-layer 0.47 (distributed-representation pattern); mini-patching exercised all 5 conditions incl. matched-norm additive. Canary numbers are pipeline validation, not paper-scale results.
Dataset Info
Rows: 8
Columns: 2
Columns
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/latkes/belief-state-probes-canary-v1.ds-canary-9xjinyang-gse138866-rmats-psi-canary-v1
jinyang-gse138866-rmats-psi-canary-v1
Canary per-sample rMATS-turbo PSI matrix (single-group, --statoff) for 2 GSE138866 samples (GSM4120625, GSM4120690 -- same 2 GSMs used as the jinyang-gse138866-rseqc canary). 158620 events x 2 samples.
Dataset Info
Rows: 158620
Columns: 9
Columns
Column
Type
Description
event_id
Value('large_string')
rMATS event type + numeric ID, e.g. 'SE_4038' (unique within this table -- use this, not 'coords'… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-gse138866-rmats-psi-canary-v1.daily-paper-2026-07-20-canary-gated-skill-rollout
Canary-Gated Autonomous Skill Rollout: Trace-Triggered Regression Detection and Automatic Rollback in Self-Evolving Agent Harnesses
TL;DR — Canary-gated rollout with a binomial-calibrated rolling-window trace gate reduces self-evolving skill artifact blast radius by 82.5-89.5% at a severity-dependent MTTR cost (74.5% slower for subtle, under 9% for severe regressions).
ThakiCloud AI Research · 2026-07-20 · 📝 Tech blog (KO)
Problem
Self-evolving agent harnesses… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-20-canary-gated-skill-rollout.details_CausalLM__72B-preview-canary-llamafied-qwen-llamafy-unbias-qkv
Dataset Card for Evaluation run of CausalLM/72B-preview-canary-llamafied-qwen-llamafy-unbias-qkv
Dataset automatically created during the evaluation run of model CausalLM/72B-preview-canary-llamafied-qwen-llamafy-unbias-qkv on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CausalLM__72B-preview-canary-llamafied-qwen-llamafy-unbias-qkv.jinyang-omentum-subtype-artifact-mechanism-canary-expression-matrix-v1
jinyang-omentum-subtype-artifact-mechanism-canary-expression-matrix-v1
Canary QC artifact for the expression arm of jinyang-omentum-subtype-artifact-mechanism.
Shiba v0.8.2 expression matrices for 40 canary samples (20 NovaSeq / 20 non-NovaSeq)
over 78,724 genes, from STAR 2nd-pass BAMs against the Ensembl 113 annotation.
This is canary-scale QC data, not a result. 40 of the cohort's 160 samples,
on a 78,724-gene matrix. It exists to prove the pipeline runs, to measure what it… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-expression-matrix-v1.jinyang-gse138866-rseqc-canary-v1
jinyang-gse138866-rseqc-canary-v1
CANARY (2 of 130) per-sample QC table for GSE138866 FFPE omental metastatic HGSOC bulk RNA-seq. Pipeline: STAR 2-pass alignment (split into pass1-only + pass2-with-on-the-fly-junction-insertion sbatch steps, GRCh38 Ensembl-113) -> samtools markdup -> RustQC rna (all QC modules in one BAM pass) -> per-sample JSON -> this aggregate table. Validates the full pipeline E2E (including the Stage A pass1/pass2 split and a Stage B markdup OOM fix… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-gse138866-rseqc-canary-v1.ds-canary-stat2jinyang-omentum-subtype-artifact-mechanism-canary-splice-enrichment-v1
jinyang-omentum-subtype-artifact-mechanism -- canary splicing-arm, enrichment summary
One row per rMATS event type. Every value is read directly out of the job's own
splicing_gates.json; nothing here is retyped by hand.
Read this before quoting any number
None of the enrichment results below is a scientific finding. This is a
canary: chr21+chr22 only, 40 of 1048 samples. Its job was to prove the
enrichment code path runs and emits a well-formed null. It did. The… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-splice-enrichment-v1.jinyang-omentum-subtype-artifact-mechanism-canary-strandedness-control-v1
omentum-subtype-artifact-mechanism -- canary strandedness control
Answers one question before any TPM is trusted: is the pinned featureCounts
strandedness (-s 2) correct for this cohort?
experiment.yaml pins stranded counting. Upstream Shiba passes no -s at all,
i.e. unstranded, so this is the one deliberate local deviation from upstream,
and the run is only interpretable if the deviation is right. Nothing in the
project had ever measured it, so the canary did.… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-strandedness-control-v1.jinyang-omentum-subtype-artifact-mechanism-canary-splice-events-v1
jinyang-omentum-subtype-artifact-mechanism -- canary splicing-arm, every event
The complete per-event output of the canary differential test: all
38,750 alternative-splicing events on chr21+chr22 across 40 canary
samples, all five rMATS event types. Nothing is filtered out of this table --
events that failed the missingness filter are present with keep=False and
empty p/q, so the exclusions are auditable rather than invisible.
What this is
A CANARY. 40 samples… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-splice-events-v1.jinyang-omentum-subtype-artifact-mechanism-canary-expression-arm-test-v1
jinyang-omentum-subtype-artifact-mechanism-canary-expression-arm-test-v1
Per-gene differential expression between the two canary arms (20 NovaSeq vs 20
non-NovaSeq), from the TPM matrix in jinyang-omentum-subtype-artifact-mechanism-canary-expression-matrix-v1.
78,724 genes, Mann-Whitney U (two-sided, asymptotic), BH-FDR adjusted.
Read the composition warning before using any gene from this table. The arm
contrast is not a clean platform contrast — see below.
Headline… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-expression-arm-test-v1.us-warn-layoffs
US WARN Act Layoff and Closure Filings
Every US Worker Adjustment and Retraining Notification (WARN) Act filing that has already taken effect, collected from 34 state labor department registries, normalised into a single schema, and resolved so that one employer's filings across many states read as one company.
17,000+ filings · ~2.2 million affected workers · 2005 to present.
Maintained and updated continuously by CanaryWhistle.
What WARN filings are
US employers… See the full description on the dataset page: https://huggingface.co/datasets/CanaryWhistle/us-warn-layoffs.seamless-interact-canary-transcripts
Seamless Interact - Canary Transcripts
Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries.
Dataset Description
This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.dv-media-ref-canary-1785327928ai-canary-2026codedp-bench-canarycanary_1b_SpNT_infer_result_1Minside-out-replication-canary-v1
inside-out-replication-canary-v1
Canary run: 5 questions per relation, 50 samples, Llama-3-8B. Full pipeline E2E test.
Dataset Info
Rows: 482
Columns: 11
Columns
Column
Type
Description
relation
Value('string')
Wikidata relation (P26=spouse, P264=label, P176=manufacturer, P50=author)
question_id
Value('string')
Unique question identifier
question
Value('string')
Entity-centric question text
gold_answer
Value('string')
Ground truth answer from… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-canary-v1.dataset-java-deprecation-incode-canary
