datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STRING_V12_TrainingSet**Repository: https://stringdb-downloads.org/download/protein.physical.links.v12.0.txt.gz
**Reference: Szklarczyk, D. et al. The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Research 51, D638–D646 (2023).
advbench_harmful_stringsrams-random-stringstringleveldigits.252class-label-string-token-repro
class-label-string-token-repro
Minimal repro dataset for class-label columns whose source CSV token is a label name instead of an integer id.
t-eval-instruct-stringt-eval-reason-stringpermuted-letter-string-analogies
Permuted Letter-String Analogies
This repository contains datasets introduced in Hellwig et al. (2026). The datasets are an extension of the letter-string analogies introduced in Lewis & Mitchell (2025).
Each folder contains a dataset with different data attributes, and contains a train, validation, and test set.
The naming convention is:
all_transformations_<copy>_study<N>_perm<N>
all_transformations:
Below are illustrations for each transformation on the standard alphabet.… See the full description on the dataset page: https://huggingface.co/datasets/philipp-hellwig/permuted-letter-string-analogies.
