datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-bench_Procaesar-cipherwine-images-126k
Wine Images Dataset 126K
A comprehensive dataset of 107,821 wine bottle images linked to the Wine Text Dataset 126K. This companion dataset provides high-quality wine bottle images for computer vision, multimodal machine learning, and wine recognition tasks.
Dataset Description
This dataset contains wine bottle images scraped from wine retailer websites. Each image is linked to detailed wine information (descriptions, pricing, categories, regions) via stable IDs that… See the full description on the dataset page: https://huggingface.co/datasets/cipher982/wine-images-126k.africa-cloud-cover-bias
Cloud Cover and Structural Observation Gaps in African Agricultural EO (six-zone dataset)
Supporting data for the paper "Cloud Cover and Structural Observation Gaps in
African Agricultural Earth Observation: Evidence from Six Agroecological Zones"
(Olaoye Anthony Somide, CropSense AI Research / CipherSense AI; Zenodo, doi:10.5281/zenodo.22642336; also EarthArXiv, doi:10.31223/X5J503).
Weekly usable Sentinel-2 optical observation frequency over cropland for six administrative… See the full description on the dataset page: https://huggingface.co/datasets/CipherSenseAI/africa-cloud-cover-bias.CipherscipheredTextcipheredText1cipher-gsm8k
Cipher Dataset
This dataset contains questions and answers that have been encrypted using a substitution cipher based on a random permutation.
Cipher Details
The cipher uses a random permutation (seed=42) to create a substitution mapping:
Lowercase mapping:
Original: abcdefghijklmnopqrstuvwxyz
Cipher: udaihveyrcnxobslwpkfgtzjmq
Uppercase mapping:
Original: ABCDEFGHIJKLMNOPQRSTUVWXYZ
Cipher: UDAIHVEYRCNXOBSLWPKFGTZJMQ
Dataset Structure
Each example… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/cipher-gsm8k.akkadianSWE-QA-BenchTED-59-allcipher-wmt18-zh-en-char100FEA-Benchafri-fertility-results
afri-fertility: African Language Tokenization Results
Measurement dataset for The African Language Tax — the first systematic audit of the subword tokenization penalty imposed on African languages by frontier large language models.
Every row is one (language, tokenizer, corpus) triple, with fertility, English-relative premium, and confidence intervals computed from a parallel corpus using sum-then-divide aggregation.
Dataset summary
Property
Value
Rows… See the full description on the dataset page: https://huggingface.co/datasets/CipherSenseAI/afri-fertility-results.korean-cipher
Korean-Cipher Dataset
Overview
This dataset was inspired by OpenAI's video, "Korean Cipher with OpenAI o1".It is designed to evaluate the reasoning abilities of large language models (LLMs) in understanding and reconstructing distorted Korean text.
Dataset Structure
Each sample in the dataset consists of the following fields:
id: A unique identifier for each sentence pair.
message: The original Korean sentence.
ciphertext: The distorted version of the… See the full description on the dataset page: https://huggingface.co/datasets/jsm0424/korean-cipher.cipher-attack-harmlesscipher-wmt18-zh-en-char50cipher-sft-dataset-ver0.2HelpSteer-cipher-attackEncrypted-Clustercipher-attack-HeX-PHI-5000Meta-Llama-3-8B-Instruct-cipher-harmless-gen3-4500-harmful-500-HeX-PHI-AMDAnimal-CaptionMistral-7B-Instruct-v0.2-cipher-harmless-4500-harmful-500-HeX-PHI-AMDMeta-Llama-3-8B-Instruct-cipher-harmless-gen3-4500-harmful-500-HeX-PHIMeta-Llama-3-8B-Instruct-cipher-harmless-4500-helpsteer-harmful-500-HeX-PHIMistral-7b-cipher-5000-HeX-PHI-AMDMeta-Llama-3-8B-Instruct-cipher-harmless-gen3-4500-harmful-500-HeX-PHI-hard-nocipher-attack-harmfulMistral-7b-cipher-5000-HeX-PHI
