datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
czech_encryptedenglish_encrypted_HistCiph
Dataset Card for HistCiph — English
Dataset Description
Dataset Summary
The English subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical English plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/english_encrypted_HistCiph.dutch_encryptedspanish_encryptedczech_encrypted_HistCiph
Dataset Card for HistCiph — Czech
Dataset Description
Dataset Summary
The Czech subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Czech plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants per… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/czech_encrypted_HistCiph.italian_encryptedswedish_encryptedspanish_encrypted_HistCiph
Dataset Card for HistCiph — Spanish
Dataset Description
Dataset Summary
The Spanish subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Spanish plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/spanish_encrypted_HistCiph.polish_encryptedenglish_encrypteddutch_encrypted_HistCiph
Dataset Card for HistCiph — Dutch
Dataset Description
Dataset Summary
The Dutch subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Dutch plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants per… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/dutch_encrypted_HistCiph.swedish_encrypted_HistCiph
Dataset Card for HistCiph — Swedish
Dataset Description
Dataset Summary
The Swedish subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Swedish plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/swedish_encrypted_HistCiph.hungarian_encryptedtlsflow-encrypted-traffic-dataset
TLSFlow Encrypted Traffic Flow Dataset
Synthetic encrypted traffic flow metadata for training and evaluating metadata-only classifiers.
Classes (9)
Label
Description
normal
Benign web/application traffic
vpn
VPN tunnel traffic
tor
Tor anonymity network
botnet
Botnet beaconing
c2
Command-and-control channels
file_transfer
Large file uploads/downloads
streaming
Media streaming
scanning
Port/service scanning
brute_force
Credential… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/tlsflow-encrypted-traffic-dataset.french_encryptedpolish_encrypted_HistCiph
Dataset Card for HistCiph — Polish
Dataset Description
Dataset Summary
The Polish subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Polish plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/polish_encrypted_HistCiph.icelandic_encrypteditalian_encrypted_HistCiph
Dataset Card for HistCiph — Italian
Dataset Description
Dataset Summary
The Italian subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Italian plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/italian_encrypted_HistCiph.french_encrypted_HistCiph
Dataset Card for HistCiph — French
Dataset Description
Dataset Summary
The French subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical French plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/french_encrypted_HistCiph.hungarian_encrypted_HistCiph
Dataset Card for HistCiph — Hungarian
Dataset Description
Dataset Summary
The Hungarian subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Hungarian plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/hungarian_encrypted_HistCiph.icelandic_encrypted_HistCiph
Dataset Card for HistCiph — Icelandic
Dataset Description
Dataset Summary
The Icelandic subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Icelandic plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/icelandic_encrypted_HistCiph.
