datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-bench_Procaesar-cipherwine-images-126k
Wine Images Dataset 126K
A comprehensive dataset of 107,821 wine bottle images linked to the Wine Text Dataset 126K. This companion dataset provides high-quality wine bottle images for computer vision, multimodal machine learning, and wine recognition tasks.
Dataset Description
This dataset contains wine bottle images scraped from wine retailer websites. Each image is linked to detailed wine information (descriptions, pricing, categories, regions) via stable IDs that… See the full description on the dataset page: https://huggingface.co/datasets/cipher982/wine-images-126k.llm-cipher-reasoning
llm-cipher-reasoning — data, eval results and full research ledger
Everything except the weights from a research run asking: can an LLM be trained to reason in a
more compact "language" than English, and does that actually save tokens?
Two linked lines of work on Qwen/Qwen3-4B-Instruct-2507:
Cipher invention / cross-model communication — cold-decoding tests, negotiated cipher
collusion between model pairs, a cipher-hardening arms race, and GEPA prompt optimization to get
a… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/llm-cipher-reasoning.CipherBank
CipherBank Benchmark
Benchmark description
CipherBank, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs in cryptographic decryption tasks.
CipherBank comprises 2,358 meticulously crafted problems, covering 262 unique plaintexts across 5 domains and 14 subdomains, with a focus on privacy-sensitive and real-world scenarios that necessitate encryption. From a cryptographic perspective, CipherBank incorporates 3 major categories of encryption… See the full description on the dataset page: https://huggingface.co/datasets/yu0226/CipherBank.africa-cloud-cover-bias
Cloud Cover and Structural Observation Gaps in African Agricultural EO (six-zone dataset)
Supporting data for the paper "Cloud Cover and Structural Observation Gaps in
African Agricultural Earth Observation: Evidence from Six Agroecological Zones"
(Olaoye Anthony Somide, CropSense AI Research / CipherSense AI; Zenodo, doi:10.5281/zenodo.22642336; also EarthArXiv, doi:10.31223/X5J503).
Weekly usable Sentinel-2 optical observation frequency over cropland for six administrative… See the full description on the dataset page: https://huggingface.co/datasets/CipherSenseAI/africa-cloud-cover-bias.classical-cipher-corpus
Classical Cipher Corpus
A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers.
Part of the Cipher Detective AI project:
🕵️ Space: systemslibrarian/cipher-detective-ai
📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo)
🤖 Model: systemslibrarian/cipher-detective-classifier
Intended use
Teach classical cryptanalysis.
Benchmark educational cipher-family detectors.
Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.Cipherscipher-awwwards-sft25
Cipher — Awwwards SFT 2.5 + Real v1 🦑
The training fuel for Kin's creative-web generator, AND the retrieval corpus for Kraken RAG. 96 real Awwwards Site-of-the-Day winners + ~1,200 records from official motion-library repositories.
Two ways this dataset is used
As a retrieval corpus for Kraken RAG ⭐ (the production path). The awwwards-gold.jsonl file contains 96 structured records of real Awwwards SOTD winners — tags, tech stack, motion libs, CSS features, section… See the full description on the dataset page: https://huggingface.co/datasets/Auroraventures/cipher-awwwards-sft25.cipheredTextcipheredText1cipher-gsm8k
Cipher Dataset
This dataset contains questions and answers that have been encrypted using a substitution cipher based on a random permutation.
Cipher Details
The cipher uses a random permutation (seed=42) to create a substitution mapping:
Lowercase mapping:
Original: abcdefghijklmnopqrstuvwxyz
Cipher: udaihveyrcnxobslwpkfgtzjmq
Uppercase mapping:
Original: ABCDEFGHIJKLMNOPQRSTUVWXYZ
Cipher: UDAIHVEYRCNXOBSLWPKFGTZJMQ
Dataset Structure
Each example… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/cipher-gsm8k.akkadianSWE-QA-BenchTED-59-allFEA-BenchIOAI_NLP_Training_02_sentiment_cipher_public
Sentiment Cipher
Supervised training examples and public evaluation inputs for learning the
private ZORP/NALI output protocol.
cipher-wmt18-zh-en-char100afri-fertility-results
afri-fertility: African Language Tokenization Results
Measurement dataset for The African Language Tax — the first systematic audit of the subword tokenization penalty imposed on African languages by frontier large language models.
Every row is one (language, tokenizer, corpus) triple, with fertility, English-relative premium, and confidence intervals computed from a parallel corpus using sum-then-divide aggregation.
Dataset summary
Property
Value
Rows… See the full description on the dataset page: https://huggingface.co/datasets/CipherSenseAI/afri-fertility-results.IOAI_NLP_Training_02_sentiment_cipher_private
Sentiment Cipher Private Answers
Instructor-only evaluation labels. Keep this repository private.
cipher-sft-dataset-ver0.2cipher-attack-HeX-PHI-5000Encrypted-ClusterHelpSteer-cipher-attackkorean-cipher
Korean-Cipher Dataset
Overview
This dataset was inspired by OpenAI's video, "Korean Cipher with OpenAI o1".It is designed to evaluate the reasoning abilities of large language models (LLMs) in understanding and reconstructing distorted Korean text.
Dataset Structure
Each sample in the dataset consists of the following fields:
id: A unique identifier for each sentence pair.
message: The original Korean sentence.
ciphertext: The distorted version of the… See the full description on the dataset page: https://huggingface.co/datasets/jsm0424/korean-cipher.Animal-Captioncipher-wmt18-zh-en-char50Hill_Ciphercipher-attack-systemscipher-attack-harmless
