CoolFace
Datasetpublic

maximuspowers/muat-pca-10

Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods pca Prompt Format separate Signature Dataset dataset_generation/exp_1/signature_dataset.json Model Architecture Number of Layers 4 to 6… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-pca-10.

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes24downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
maximuspowers/muat-pca-10 · CoolFace