canary
Datasets
All datasets matching “canary”german-canary-asr-0324
Dataset Beschreibung
Allgemeine Informationen
Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 16.1, Voxpopuli und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert.
Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-canary-asr-0324.polymarket-canary-tape
polymarket-canary-tape
Continuous tape from Scribe (Bot E / bot_e_recorder), a single always-on VPS node. Captures co-located CEX trades and Polymarket market-channel WebSocket events over a fixed UTC window for microstructure and lead-lag research.
Where this came from: released alongside polymarket-bot-lab
(11 open-source Polymarket trading bot candidates, Apache-2.0) by the team behind
OracleMangle, which builds dispute-risk scoring for
prediction-market questions. Both… See the full description on the dataset page: https://huggingface.co/datasets/oraclemangle/polymarket-canary-tape.MIRAGE-CanaryDocs
MIRAGE CanaryDocs
MIRAGE CanaryDocs is an English synthetic enterprise-document dataset for structured privacy-unit,
canary, and ordered multi-chunk evaluation. It is the companion dataset for the EMNLP 2026 paper
When Metadata Remembers: Ordered Provenance Enables Document-Level Embedding Inversion.
Project documentation and schemas are also available in the
MIRAGE GitHub repository.
Dataset summary
The dataset contains complete synthetic documents, ordered token… See the full description on the dataset page: https://huggingface.co/datasets/LevenKoko/MIRAGE-CanaryDocs.CanaryAura
Dataset Card for "Canary Aura"
This is a dataset for...
terminal_bench_2_Qwen3_32B_canary_ghdevenron_canary
CanaryBench-Enron
Frequency-aware canary injection benchmark for auditing memorization
in finetuned language models, built on the Enron email corpus.
Dataset Description
This dataset is part of CanaryBench, a benchmark for evaluating
memorization in finetuned language models across repetition tiers
and privacy regimes.
Frequency tiers: 1×, 10×, 50×
Domain: Email (Enron corpus)
Member canaries: 770
Reference canaries: 1000
Files… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/enron_canary.
