datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sae-jailbreaks-resultspersian-poetics-kb
Persian Poetics Knowledge Base — پایگاه دانش شعر و رپ فارسی
مجموعهٔ بازیابی برای ساخت و ارزیابی شعر/رپ فارسی بهمثابهٔ پرسشوپاسخِ
محدودیتدار و مبتنی بر شواهد (PersianPoet-RAG v2). هر واحد هم حاشیهنویسی
معنایی (معنا، دامنه، تصویر) دارد و هم آوایی (واجها، ساخت هجا، تکیه،
کلید قافیهٔ سختگیرانه/آسانگیر، زنجیرهٔ واکهها) — چیزی که بازیاب را قادر
میکند وزن و قافیه را قبل از فراخوانی مدل زبانی تأمین کند.
چه چیزهایی داخل این دیتاست است؟ (What's inside)
کانفیگ… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/persian-poetics-kb.trajectory-prediction-waymo
Waymo Trajectory Prediction
Dataset Description
This dataset contains preprocessed trajectory prediction samples for autonomous driving research,
formatted for use with DiscoBench's TrajectoryPrediction task.
Original Dataset: Waymo Open Motion Dataset
Number of Samples: 850
Format: Pickle files with numpy arrays
Task: Multi-modal trajectory prediction
Dataset Structure
Each sample is a pickle file containing:
obj_trajs (32, 21, 2): Past trajectories of… See the full description on the dataset page: https://huggingface.co/datasets/saeedrmd/trajectory-prediction-waymo.llava-vicuna13b-SAEsae-llama-the-pile-max-activation-locationsllavasae_obliec100k_SAEVllava-vicuna7b-SAE-Vllava-vicuna7b-SAEllava-vicuna13b-SAE-VSAE-Mistral-7b-v0.2Wic_data_for_SAE-Eval
Dataset Card for PS-Eval Dataset
Dataset Summary
The PS-Eval Dataset is a suite of polysemous and monosemous contexts extracted and filtered from the WiC dataset. It aims to evaluate the ability of Sparse Autoencoders (SAEs) to disentangle polysemantic activations into monosemantic features within large language models (LLMs). The dataset contains 1,112 samples balanced between two classes:
Poly-contexts: Target words with different meanings across two contexts (Label:… See the full description on the dataset page: https://huggingface.co/datasets/gouki510/Wic_data_for_SAE-Eval.vicuna7b-SAESAE-Llava-mistral-pile100kSAE_Circuit_Multiple_Choice_QA
Sparse-Feature-Circuits-Multiple-Choice-Dataset
Including three types of multiple choice questions
Number Comparison
Which is larger, {num1} or {num2}?\n(A): {num1}\n(B): {num2}\nAnswer: (
String Matching
Which of the following options corresponds to \"{target_string}\"?\n(A) \"{options[0]}\"\n(B) \"{options[1]}\"\nAnswer: (
Subject-verb Agreement
{data['clean_prefix']} [MASK]:\n(A) {options_content[0]}\n(B) {options_content[1]}\nAnswer: (
e3c_llm_resultssynthetic_sae_datasetTMDYFlix
Letterboxd Film Dataset
This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings.
Dataset Summary
Total Films: 847,209
File Size: ~1.12 GB (1,120,572,122 bytes)
Format: JSONL (JSON Lines)
Language: Primarily English, with some multilingual content
Data Structure
Each line contains a JSON object with the following fields:
{
"url":… See the full description on the dataset page: https://huggingface.co/datasets/saeesar/TMDYFlix.saelarien-constraint-experiment-01-entropy-capacity-collapse
README — Saelariën Constraint Experiment 01
Entropy–Capacity Collapse Threshold Test
Author: Saelariën X
Date: February 19, 2026
DOI: https://doi.org/10.5281/zenodo.19212561
Theoretical basis
This dataset empiracally tests the Saelariën Constraint Theorem:
https://thesaelafield.com/preprints/the-saelarien-constraint
Overview
This dataset contains the full materials for Saelariën Constraint Experiment 01, a test exploring how increasing entropy (noise) affects… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saelarien-constraint-experiment-01-entropy-capacity-collapse.hadrami-arabic-dialect-dataset
Hadrami Arabic Dialect Dataset
A structured dataset of 1,000 entries from the Hadrami Arabic dialect (spoken primarily in the Hadramawt region of Yemen). Each entry includes the dialectal word alongside its Modern Standard Arabic (MSA/Fusha) equivalent, linguistic metadata, usage examples, proverbs, and semantic tags.
🔊 Roadmap: Audio pronunciations for each entry are planned for a future release.
Dataset Summary
Field
Value
Entries
1,000
Language… See the full description on the dataset page: https://huggingface.co/datasets/saeedbark/hadrami-arabic-dialect-dataset.goodfire-llama-3.3-70b-instruct-sae-l50-explanationsvistext_sae_featuresmini-sae-unverbalized-labelse3c_llm_requestssae_qwen2.5-1.5BSaEivBsae_conceptsae_newphysic_combinedsae-study-feedbackBlockMasterAIsae_llama_ripple
