datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polymarket-canary-tape
polymarket-canary-tape
Continuous tape from Scribe (Bot E / bot_e_recorder), a single always-on VPS node. Captures co-located CEX trades and Polymarket market-channel WebSocket events over a fixed UTC window for microstructure and lead-lag research.
Where this came from: released alongside polymarket-bot-lab
(11 open-source Polymarket trading bot candidates, Apache-2.0) by the team behind
OracleMangle, which builds dispute-risk scoring for
prediction-market questions. Both… See the full description on the dataset page: https://huggingface.co/datasets/oraclemangle/polymarket-canary-tape.MIRAGE-CanaryDocs
MIRAGE CanaryDocs
MIRAGE CanaryDocs is an English synthetic enterprise-document dataset for structured privacy-unit,
canary, and ordered multi-chunk evaluation. It is the companion dataset for the EMNLP 2026 paper
When Metadata Remembers: Ordered Provenance Enables Document-Level Embedding Inversion.
Project documentation and schemas are also available in the
MIRAGE GitHub repository.
Dataset summary
The dataset contains complete synthetic documents, ordered token… See the full description on the dataset page: https://huggingface.co/datasets/LevenKoko/MIRAGE-CanaryDocs.qwen3p6-27b-canary-rolloutsjinyang-gse138866-rmats-psi-canary-v1
jinyang-gse138866-rmats-psi-canary-v1
Canary per-sample rMATS-turbo PSI matrix (single-group, --statoff) for 2 GSE138866 samples (GSM4120625, GSM4120690 -- same 2 GSMs used as the jinyang-gse138866-rseqc canary). 158620 events x 2 samples.
Dataset Info
Rows: 158620
Columns: 9
Columns
Column
Type
Description
event_id
Value('large_string')
rMATS event type + numeric ID, e.g. 'SE_4038' (unique within this table -- use this, not 'coords'… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-gse138866-rmats-psi-canary-v1.jinyang-omentum-subtype-artifact-mechanism-canary-expression-matrix-v1
jinyang-omentum-subtype-artifact-mechanism-canary-expression-matrix-v1
Canary QC artifact for the expression arm of jinyang-omentum-subtype-artifact-mechanism.
Shiba v0.8.2 expression matrices for 40 canary samples (20 NovaSeq / 20 non-NovaSeq)
over 78,724 genes, from STAR 2nd-pass BAMs against the Ensembl 113 annotation.
This is canary-scale QC data, not a result. 40 of the cohort's 160 samples,
on a 78,724-gene matrix. It exists to prove the pipeline runs, to measure what it… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-expression-matrix-v1.jinyang-gse138866-rseqc-canary-v1
jinyang-gse138866-rseqc-canary-v1
CANARY (2 of 130) per-sample QC table for GSE138866 FFPE omental metastatic HGSOC bulk RNA-seq. Pipeline: STAR 2-pass alignment (split into pass1-only + pass2-with-on-the-fly-junction-insertion sbatch steps, GRCh38 Ensembl-113) -> samtools markdup -> RustQC rna (all QC modules in one BAM pass) -> per-sample JSON -> this aggregate table. Validates the full pipeline E2E (including the Stage A pass1/pass2 split and a Stage B markdup OOM fix… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-gse138866-rseqc-canary-v1.jinyang-omentum-subtype-artifact-mechanism-canary-splice-enrichment-v1
jinyang-omentum-subtype-artifact-mechanism -- canary splicing-arm, enrichment summary
One row per rMATS event type. Every value is read directly out of the job's own
splicing_gates.json; nothing here is retyped by hand.
Read this before quoting any number
None of the enrichment results below is a scientific finding. This is a
canary: chr21+chr22 only, 40 of 1048 samples. Its job was to prove the
enrichment code path runs and emits a well-formed null. It did. The… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-splice-enrichment-v1.jinyang-omentum-subtype-artifact-mechanism-canary-strandedness-control-v1
omentum-subtype-artifact-mechanism -- canary strandedness control
Answers one question before any TPM is trusted: is the pinned featureCounts
strandedness (-s 2) correct for this cohort?
experiment.yaml pins stranded counting. Upstream Shiba passes no -s at all,
i.e. unstranded, so this is the one deliberate local deviation from upstream,
and the run is only interpretable if the deviation is right. Nothing in the
project had ever measured it, so the canary did.… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-strandedness-control-v1.jinyang-omentum-subtype-artifact-mechanism-canary-splice-events-v1
jinyang-omentum-subtype-artifact-mechanism -- canary splicing-arm, every event
The complete per-event output of the canary differential test: all
38,750 alternative-splicing events on chr21+chr22 across 40 canary
samples, all five rMATS event types. Nothing is filtered out of this table --
events that failed the missingness filter are present with keep=False and
empty p/q, so the exclusions are auditable rather than invisible.
What this is
A CANARY. 40 samples… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-splice-events-v1.jinyang-omentum-subtype-artifact-mechanism-canary-expression-arm-test-v1
jinyang-omentum-subtype-artifact-mechanism-canary-expression-arm-test-v1
Per-gene differential expression between the two canary arms (20 NovaSeq vs 20
non-NovaSeq), from the TPM matrix in jinyang-omentum-subtype-artifact-mechanism-canary-expression-matrix-v1.
78,724 genes, Mann-Whitney U (two-sided, asymptotic), BH-FDR adjusted.
Read the composition warning before using any gene from this table. The arm
contrast is not a clean platform contrast — see below.
Headline… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-subtype-artifact-mechanism-canary-expression-arm-test-v1.seamless-interact-canary-transcripts
Seamless Interact - Canary Transcripts
Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries.
Dataset Description
This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.codedp-bench-canaryinside-out-replication-canary-v1
inside-out-replication-canary-v1
Canary run: 5 questions per relation, 50 samples, Llama-3-8B. Full pipeline E2E test.
Dataset Info
Rows: 482
Columns: 11
Columns
Column
Type
Description
relation
Value('string')
Wikidata relation (P26=spouse, P264=label, P176=manufacturer, P50=author)
question_id
Value('string')
Unique question identifier
question
Value('string')
Entity-centric question text
gold_answer
Value('string')
Ground truth answer from… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-canary-v1.codedp-bench-canary-miafnbm-current-simulation-identifiability-canary-v1
fnbm-current-simulation-identifiability-canary-v1
Canary validating the new training-time simulation_eval hook: per-sample model contributions (pred, z1_f{k}, z2_f{i}f{j}, model_z_abs*) extracted via forward_interpret on the 20260706_simulation valid split, at the final epoch (9) of a 10-epoch canary run. Baseline (all-zero) antiabsorption penalties -- validates plumbing on real data/GPU, not a scientific result.
Dataset Info
Rows: 10000
Columns: 277… See the full description on the dataset page: https://huggingface.co/datasets/arushram/fnbm-current-simulation-identifiability-canary-v1.precommittal-canary-results-v1
precommittal-canary-results-v1
Canary shard results for precommittal experiment. Contains signal check metrics, layer sweep results (layers 8/17/26/35, alpha 0.8-0.95), and answer extraction check samples. Canary FAILED: rho≈0 at all layers, cosine@25%<0.7.
Dataset Info
Rows: 41
Columns: 17
Columns
Column
Type
Description
artifact_type
Value('string')
Type: signal_check, layer_sweep, layer_sweep_alpha, or answer_extraction_check
key… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/precommittal-canary-results-v1.sdc-canary-v1
sdc-canary-v1
Canary shard outputs for semantic distance coding experiment.
Dataset Info
Rows: 6
Columns: 15
Columns
Column
Type
Description
problem_id
Value('string')
Problem identifier from EsoLang-Bench (E01-E20)
language
Value('string')
Target programming language name
tiobe_rank
Value('int64')
TIOBE index rank (1=Python, 47=OCaml)
tiobe_pct
Value('float64')
TIOBE index percentage share
condition
Value('string')
Prompting strategy:… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/sdc-canary-v1.editable-sketch-repro-canaryinside-out-replication-v2-pamqfix-canary-v1
Inside-Out Replication V2 — P(a|q) A.8.1-fix CANARY
Corrected externals (04_external_scores.py with the A.8.1 in-context
tokenization fix) over a deterministic 25-question-id subset of
Mistral-7B-Instruct-v0.3 / P264 / test (5832 candidate (q,a) rows).
Bounded subset used to validate the fix end-to-end on MLL before the full
3×4 re-score (per preflight rev2 spec; see red_team_brief.md).
Canary self-validation
C2 rows scored: 5832 (no build_qa_sequence ValueError;… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-v2-pamqfix-canary-v1.test-rewrite-canary
test-rewrite-canary
Subset: tiny
5-row test parquet, columns id (int) and value (float).
self-consistency-correction-exp0-sanity-canary
self-consistency-correction — Experiment 0 (sanity check)
Paper: Correcting Generator Scores via Self-Consistency: Theory, Practical Approximations, and Experiments.
§7.1 no-paraphrase baseline. 8 hand-crafted cases (L1-L4 low paraphrase richness, H1-H4 high paraphrase richness). For each case, all candidate answers are scored with:
G_raw = log P(y | X) via teacher-forced scoring (longest-common-prefix boundary handling to avoid whitespace retokenization bugs)
V' = log P(Yes |… See the full description on the dataset page: https://huggingface.co/datasets/latkes/self-consistency-correction-exp0-sanity-canary.
