datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dr-saeid-ghezelbaash-entity-data
Dr. Saeed Ghezelbash Public Knowledge Graph
A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval.
The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m
SAE Feature Activations — Llama 3.1 8B Instruct, Layer 19 (LMSYS-Chat-1M)
This dataset contains Sparse Autoencoder (SAE) feature activations extracted from layer 19 of Meta's Llama 3.1 8B Instruct on conversations from LMSYS-Chat-1M.
It also has natural language explainations of features generated by GPT OSS 120B. See subset 4 for details.
The SAE used is Goodfire/Llama-3.1-8B-Instruct-SAE-l19, which decomposes layer-19 residual stream activations into interpretable sparse features.… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m.saenaiheroinenosodatekata
Bangumi Image Base of Saenai Heroine No Sodatekata
This is the image base of bangumi Saenai Heroine no Sodatekata, we detected 26 characters, 3436 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saenaiheroinenosodatekata.sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playcommon-voice-17-en-age-gender-accentsae-jailbreaks-resultsllama-3.1-8b-Instruct_sae_repscommon-voice-17-en-age-genderllama_3.1-sae-23-29-code-activationsfa-en-ar-handwritten-ocr-v1
Multi-script Synthetic Handwritten OCR — fa / ar / en
A large, clean, augmentation-rich synthetic handwriting dataset for
training and benchmarking OCR / HTR models on Persian (fa), Arabic
(ar) and English (en). Every line image ships with an exact Unicode
transcription plus rich provenance metadata (writer style, font, ink, script
direction, digit system). Page-level PAGE-XML and COCO ground truth support
layout-aware training and evaluation out of the box.
1,000 rendered… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/fa-en-ar-handwritten-ocr-v1.qwen3-8b-base-atlas-SAE
Qwen3-8B-Base Feature Atlas
A single queryable SQLite database (atlas.sqlite, ~570 MB) that maps the internals of
Qwen/Qwen3-8B-Base — every weight channel and
every sparse-autoencoder feature scored for what it selects for, across a register-diverse
corpus of 4,946 prompts.
It is not a text dataset. There are no training rows. It is an index of model internals
— the kind of thing you query to find "which channels in layer 23 discriminate compliance from
authentic-personality… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/qwen3-8b-base-atlas-SAE.SAE_activations_modal_sentencesvctk-16khz
Dataset Card for VCTK (16kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.vctk-48khz
Dataset Card for VCTK (48kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.llava-sae-explanations-5kThis is the explanation generated for the first 5k features for sae 131k trained on llava-next-llama3-8B.
The revised one is the explanation using revised prompt and cached on the lmms-lab/sae-sample-cache-dataset. The legacy one is the old explanations with an old prompt and use the first 15% of the LLaVA-NeXT-Data.
In our paper, the reported evaluation result are based on the revised ones and feature probing is done on both of the versions
majestrino-1.00-16xk5-sae-features
Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5)
Top-2000 activating audio samples for each feature in the
Majestrino 1.00 SAE.
Overview
Metric
Value
SAE Architecture
16x expansion, k=5, d_model=768
Total Features
12,288
Alive Features
10,684
Audio per Feature
Up to 2,000 highest-activating
Audio Format
Opus (24 kbps OGG container)
Total TAR Files
1069
Source Dataset
laion/majestrino-data
File Structure
Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.mmlu-pro-pluspersian-poetics-kb
Persian Poetics Knowledge Base — پایگاه دانش شعر و رپ فارسی
مجموعهٔ بازیابی برای ساخت و ارزیابی شعر/رپ فارسی بهمثابهٔ پرسشوپاسخِ
محدودیتدار و مبتنی بر شواهد (PersianPoet-RAG v2). هر واحد هم حاشیهنویسی
معنایی (معنا، دامنه، تصویر) دارد و هم آوایی (واجها، ساخت هجا، تکیه،
کلید قافیهٔ سختگیرانه/آسانگیر، زنجیرهٔ واکهها) — چیزی که بازیاب را قادر
میکند وزن و قافیه را قبل از فراخوانی مدل زبانی تأمین کند.
چه چیزهایی داخل این دیتاست است؟ (What's inside)
کانفیگ… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/persian-poetics-kb.ESMC-SAE-Features
ESMC Sparse Autoencoder Features Table
This dataset contains a Parquet table of the 16,384 features from the ESMC-6B-sae-layer60-k64-codebook16384, that was used for analysis in the ESMC paper and to construct the ESM Atlas. This table provides descriptions of the precomputed features that can be activated through the spotlight SAE model, assisting users for downstream interpretation of the insights revealed by ESMC.
Download the table here.
The features descriptions are in the… See the full description on the dataset page: https://huggingface.co/datasets/biohub/ESMC-SAE-Features.JP-HomophoneBench
JP-HomophoneBench
A deterministic Japanese ASR benchmark index for separating eight error/disambiguation classes:
exact_homophone
near_homophone
voicing
long_vowel
geminate
moraic_nasal
pitch_accent
semantic_only
Important design rule
This repository is metadata-first. Source audio is not redistributed by default. Each row stores source repository/config/split/row identifiers so audio can be rehydrated under the original source license.
exact_homophone and… See the full description on the dataset page: https://huggingface.co/datasets/saeeew/JP-HomophoneBench.iemocap-original-wavlm-large-layer-9-temporaliemocap-vc-wavlm-large-layer-9-temporalsmollm2-135m-instruct-SAE
Layer
EV
Mean L0
Recon Loss
Dead %
0
0.9480
48.74
0.2074
0.0
1
0.9599
43.65
0.3298
0.0
2
0.9631
46.81
0.5021
0.0
3
0.9508
46.56
0.7462
0.0
4
0.9463
46.23
0.8936
0.0
5
0.9350
47.57
1.1605
0.0
6
0.9306
48.44
1.3838
0.0
7
0.9318
49.51
1.5446
0.0
8
0.9432
46.52
1.6598
0.0
9
0.9373
47.15
2.0706
0.0
10
0.9348
45.53
2.2983
0.0
11
0.9905
48.58
5.8113
0.0
12
0.9901
48.42
6.1039
0.0
13
0.9891
46.15
6.9692
0.0
14
0.9884
44.76
7.1844
0.0
15
0.9863
47.63
8.6521
0.0… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/smollm2-135m-instruct-SAE.mfa-vs-sae-2026-webapp-dataMolDeBERTa
PubChem Dataset for MolDeBERTa Pretraining
MolDeBERTa: Foundational Model for Physicochemical and Substructure-Informed Molecular Representation Learning
[Paper] | [Github Repo] | [Model Collection] | [Cite]
Abstract
Foundational models that learn the "language" of molecules are essential for accelerating material and drug discovery. These self-learning models can be trained on large collections of unlabelled molecules, enabling applications such as property… See the full description on the dataset page: https://huggingface.co/datasets/SaeedLab/MolDeBERTa.e-daic-ai-controlledkhattat
Khattat
Synthetic multilingual handwritten OCR dataset by saeidseyfi.
Handwritten samples in Persian (fa), Arabic (ar) and English (en) rendered with a dynamic pen (variable stroke width and pressure along the writing path), with three configs:
Configs
1. default — line crops (train 21,726 / validation 2,062 / test 924)
Handwritten line images with ground-truth transcriptions.
Column
Type
Description
image
image
RGB handwritten line crop… See the full description on the dataset page: https://huggingface.co/datasets/saeidseyfi/khattat.iemocap-original-wavlm-layer-6-temporaliemocap-vc-wavlm-layer-6-temporalchess_sae_individual_games_filtered
