CoolFace
Datasetpublic

s4um1l/tiny-aya-medical-concept-probes

Tiny Aya Cross-Lingual Medical Concept Probes Dataset Description 20 medical concepts expressed as full sentences in 10 languages, designed for probing cross-lingual concept representations in multilingual LLMs. Each concept is a complete declarative sentence preserving the same semantic structure across all languages. Purpose These probe sentences serve as stimuli for mechanistic interpretability analysis -- specifically, extracting residual stream… See the full description on the dataset page: https://huggingface.co/datasets/s4um1l/tiny-aya-medical-concept-probes.

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes9downloads
2 commits on main
1ab87487mo ago

initial dataset: 20 medical concepts x 10 languages

s4um1l
a82276e7mo ago

initial commit

s4um1l