datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dr-saeid-ghezelbaash-entity-data
Dr. Saeed Ghezelbash Public Knowledge Graph
A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval.
The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.persian-poetics-kb
Persian Poetics Knowledge Base — پایگاه دانش شعر و رپ فارسی
مجموعهٔ بازیابی برای ساخت و ارزیابی شعر/رپ فارسی بهمثابهٔ پرسشوپاسخِ
محدودیتدار و مبتنی بر شواهد (PersianPoet-RAG v2). هر واحد هم حاشیهنویسی
معنایی (معنا، دامنه، تصویر) دارد و هم آوایی (واجها، ساخت هجا، تکیه،
کلید قافیهٔ سختگیرانه/آسانگیر، زنجیرهٔ واکهها) — چیزی که بازیاب را قادر
میکند وزن و قافیه را قبل از فراخوانی مدل زبانی تأمین کند.
چه چیزهایی داخل این دیتاست است؟ (What's inside)
کانفیگ… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/persian-poetics-kb.SAEVerbalizer-Data
SAEVerbalizer Data
Evaluation data for SAEVerbalizer: Generating Explanations for Sparse
Autoencoder Features via Representation Verbalization.
This release contains the paper's three evaluation sets for both the 27B
verbalizer and the 1B-to-27B adapter-based verbalization system. Training data
may be added in a later release.
Inference and Reference Agreement evaluation code is available in
THU-KEG/SAEVerbalizer.
Data collections
Collection
Feature space… See the full description on the dataset page: https://huggingface.co/datasets/THU-KEG/SAEVerbalizer-Data.dialectic-preferences-bias-aae-sae-parallel
Dialectic Preferences Bias Dataset
Dataset Description
Overview
This dataset is part of a research study examining dialectic preference bias in Large Language Models (LLMs). It contains paired sentences in African American English (AAE) and Standard American English (SAE), used to analyze potential biases in language models' treatment of different dialects.
The dataset contains two columns:
african_american_english: Text samples in African American English… See the full description on the dataset page: https://huggingface.co/datasets/furquan/dialectic-preferences-bias-aae-sae-parallel.TMDYFlix
Letterboxd Film Dataset
This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings.
Dataset Summary
Total Films: 847,209
File Size: ~1.12 GB (1,120,572,122 bytes)
Format: JSONL (JSON Lines)
Language: Primarily English, with some multilingual content
Data Structure
Each line contains a JSON object with the following fields:
{
"url":… See the full description on the dataset page: https://huggingface.co/datasets/saeesar/TMDYFlix.saelarien-constraint-experiment-01-entropy-capacity-collapse
README — Saelariën Constraint Experiment 01
Entropy–Capacity Collapse Threshold Test
Author: Saelariën X
Date: February 19, 2026
DOI: https://doi.org/10.5281/zenodo.19212561
Theoretical basis
This dataset empiracally tests the Saelariën Constraint Theorem:
https://thesaelafield.com/preprints/the-saelarien-constraint
Overview
This dataset contains the full materials for Saelariën Constraint Experiment 01, a test exploring how increasing entropy (noise) affects… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saelarien-constraint-experiment-01-entropy-capacity-collapse.
