datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Instruction_recall_dataset
CanaryBench-PII
Frequency-aware canary injection benchmark for auditing memorization
in finetuned language models, built on the AI4Privacy PII reconstruction
task.
Dataset Description
This dataset is part of CanaryBench, a benchmark for evaluating
memorization in finetuned language models across repetition tiers
and privacy regimes.
Frequency tiers: 1×, 10×, 50×
PII types: EMAIL, PHONE
Member canaries: 770
Reference canaries: 1000
Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.enron_canary
CanaryBench-Enron
Frequency-aware canary injection benchmark for auditing memorization
in finetuned language models, built on the Enron email corpus.
Dataset Description
This dataset is part of CanaryBench, a benchmark for evaluating
memorization in finetuned language models across repetition tiers
and privacy regimes.
Frequency tiers: 1×, 10×, 50×
Domain: Email (Enron corpus)
Member canaries: 770
Reference canaries: 1000
Files… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/enron_canary.gencode-mouse
GENCODE
GENCODE is a comprehensive annotation project that aims to provide high-quality annotations of the human and mouse genomes.
The project is part of the ENCODE (ENCyclopedia Of DNA Elements) scale-up project, which seeks to identify all functional elements in the human genome.
Disclaimer
This is an UNOFFICIAL release of the GENCODE by Paul Flicek, Roderic Guigo, Manolis Kellis, Mark Gerstein, Benedict Paten, Michael Tress, Jyoti Choudhary, et al.
The team… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/gencode-mouse.
