datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
caesar-cipherafrica-cloud-cover-bias
Cloud Cover and Structural Observation Gaps in African Agricultural EO (six-zone dataset)
Supporting data for the paper "Cloud Cover and Structural Observation Gaps in
African Agricultural Earth Observation: Evidence from Six Agroecological Zones"
(Olaoye Anthony Somide, CropSense AI Research / CipherSense AI; Zenodo, doi:10.5281/zenodo.22642336; also EarthArXiv, doi:10.31223/X5J503).
Weekly usable Sentinel-2 optical observation frequency over cropland for six administrative… See the full description on the dataset page: https://huggingface.co/datasets/CipherSenseAI/africa-cloud-cover-bias.classical-cipher-corpus
Classical Cipher Corpus
A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers.
Part of the Cipher Detective AI project:
🕵️ Space: systemslibrarian/cipher-detective-ai
📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo)
🤖 Model: systemslibrarian/cipher-detective-classifier
Intended use
Teach classical cryptanalysis.
Benchmark educational cipher-family detectors.
Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.cipher-awwwards-sft25
Cipher — Awwwards SFT 2.5 + Real v1 🦑
The training fuel for Kin's creative-web generator, AND the retrieval corpus for Kraken RAG. 96 real Awwwards Site-of-the-Day winners + ~1,200 records from official motion-library repositories.
Two ways this dataset is used
As a retrieval corpus for Kraken RAG ⭐ (the production path). The awwwards-gold.jsonl file contains 96 structured records of real Awwwards SOTD winners — tags, tech stack, motion libs, CSS features, section… See the full description on the dataset page: https://huggingface.co/datasets/Auroraventures/cipher-awwwards-sft25.cipher-wmt18-zh-en-char100afri-fertility-results
afri-fertility: African Language Tokenization Results
Measurement dataset for The African Language Tax — the first systematic audit of the subword tokenization penalty imposed on African languages by frontier large language models.
Every row is one (language, tokenizer, corpus) triple, with fertility, English-relative premium, and confidence intervals computed from a parallel corpus using sum-then-divide aggregation.
Dataset summary
Property
Value
Rows… See the full description on the dataset page: https://huggingface.co/datasets/CipherSenseAI/afri-fertility-results.cipher-wmt18-zh-en-char50cipher-wmt18-zh-en-char0HelpingAI__Cipher-20B-details
Dataset Card for Evaluation run of HelpingAI/Cipher-20B
Dataset automatically created during the evaluation run of model HelpingAI/Cipher-20B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HelpingAI__Cipher-20B-details.
