CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01doctor-ghezelbaash /dr-saeid-ghezelbaash-entity-data Dr. Saeed Ghezelbash Public Knowledge Graph A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval. The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.textquestion-answering1K<n<10K1 likes2k downloads3d agoHugging Face02scaleinvariant /sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m SAE Feature Activations — Llama 3.1 8B Instruct, Layer 19 (LMSYS-Chat-1M) This dataset contains Sparse Autoencoder (SAE) feature activations extracted from layer 19 of Meta's Llama 3.1 8B Instruct on conversations from LMSYS-Chat-1M. It also has natural language explainations of features generated by GPT OSS 120B. See subset 4 for details. The SAE used is Goodfire/Llama-3.1-8B-Instruct-SAE-l19, which decomposes layer-19 residual stream activations into interpretable sparse features.… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m.tabularfeature-extraction100M<n<1B0 likes1.3k downloads7mo agoHugging Face03BangumiBase /saenaiheroinenosodatekata Bangumi Image Base of Saenai Heroine No Sodatekata This is the image base of bangumi Saenai Heroine no Sodatekata, we detected 26 characters, 3436 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saenaiheroinenosodatekata.image1K<n<10K0 likes758 downloads3y agoHugging Face04apollo-research /sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playtext10K<n<100K0 likes747 downloads3y agoHugging Face05saeedzou /common-voice-17-en-age-gender-accentaudio100K<n<1M0 likes707 downloads2mo agoHugging Face06djain95 /sae-jailbreaks-resultsimagen<1K1 likes664 downloads5mo agoHugging Face07mksethi /llama-3.1-8b-Instruct_sae_repstext100K<n<1M0 likes571 downloads1y agoHugging Face08saeedzou /common-voice-17-en-age-genderaudio100K<n<1M0 likes503 downloads2mo agoHugging Face09quintic /llama_3.1-sae-23-29-code-activationstext10K<n<100K1 likes425 downloads2y agoHugging Face10saeid1999 /fa-en-ar-handwritten-ocr-v1 Multi-script Synthetic Handwritten OCR — fa / ar / en A large, clean, augmentation-rich synthetic handwriting dataset for training and benchmarking OCR / HTR models on Persian (fa), Arabic (ar) and English (en). Every line image ships with an exact Unicode transcription plus rich provenance metadata (writer style, font, ink, script direction, digit system). Page-level PAGE-XML and COCO ground truth support layout-aware training and evaluation out of the box. 1,000 rendered… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/fa-en-ar-handwritten-ocr-v1.imageimage-to-text10K<n<100K0 likes412 downloads9d agoHugging Face11juiceb0xc0de /qwen3-8b-base-atlas-SAE Qwen3-8B-Base Feature Atlas A single queryable SQLite database (atlas.sqlite, ~570 MB) that maps the internals of Qwen/Qwen3-8B-Base — every weight channel and every sparse-autoencoder feature scored for what it selects for, across a register-diverse corpus of 4,946 prompts. It is not a text dataset. There are no training rows. It is an index of model internals — the kind of thing you query to find "which channels in layer 23 discriminate compliance from authentic-personality… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/qwen3-8b-base-atlas-SAE.tabular1M<n<10M1 likes373 downloads11d agoHugging Face12yoheikobashi /SAE_activations_modal_sentencestabular100K<n<1M0 likes333 downloads1y agoHugging Face13saeedzou /vctk-16khz Dataset Card for VCTK (16kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.audioautomatic-speech-recognition10K<n<100K0 likes304 downloads2mo agoHugging Face14saeedzou /vctk-48khzgated Dataset Card for VCTK (48kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.audioautomatic-speech-recognition10K<n<100K1 likes301 downloads2mo agoHugging Face15lmms-lab /llava-sae-explanations-5kThis is the explanation generated for the first 5k features for sae 131k trained on llava-next-llama3-8B. The revised one is the explanation using revised prompt and cached on the lmms-lab/sae-sample-cache-dataset. The legacy one is the old explanations with an old prompt and use the first 15% of the LLaVA-NeXT-Data. In our paper, the reported evaluation result are based on the revised ones and feature probing is done on both of the versions image1K<n<10K6 likes295 downloads2y agoHugging Face16laion /majestrino-1.00-16xk5-sae-features Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5) Top-2000 activating audio samples for each feature in the Majestrino 1.00 SAE. Overview Metric Value SAE Architecture 16x expansion, k=5, d_model=768 Total Features 12,288 Alive Features 10,684 Audio per Feature Up to 2,000 highest-activating Audio Format Opus (24 kbps OGG container) Total TAR Files 1069 Source Dataset laion/majestrino-data File Structure Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.audioaudio-classification10M<n<100M0 likes271 downloads6mo agoHugging Face17saeidasgari /mmlu-pro-plustabular10K<n<100K2 likes268 downloads2y agoHugging Face18saeid1999 /persian-poetics-kb Persian Poetics Knowledge Base — پایگاه دانش شعر و رپ فارسی مجموعهٔ بازیابی برای ساخت و ارزیابی شعر/رپ فارسی به‌مثابهٔ پرسش‌وپاسخِ محدودیت‌دار و مبتنی بر شواهد (PersianPoet-RAG v2). هر واحد هم حاشیه‌نویسی معنایی (معنا، دامنه، تصویر) دارد و هم آوایی (واج‌ها، ساخت هجا، تکیه، کلید قافیهٔ سخت‌گیرانه/آسان‌گیر، زنجیرهٔ واکه‌ها) — چیزی که بازیاب را قادر می‌کند وزن و قافیه را قبل از فراخوانی مدل زبانی تأمین کند. چه چیزهایی داخل این دیتاست است؟ (What's inside) کانفیگ… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/persian-poetics-kb.tabularquestion-answering10K<n<100K0 likes248 downloads10d agoHugging Face19biohub /ESMC-SAE-Features ESMC Sparse Autoencoder Features Table This dataset contains a Parquet table of the 16,384 features from the ESMC-6B-sae-layer60-k64-codebook16384, that was used for analysis in the ESMC paper and to construct the ESM Atlas. This table provides descriptions of the precomputed features that can be activated through the spotlight SAE model, assisting users for downstream interpretation of the insights revealed by ESMC. Download the table here. The features descriptions are in the… See the full description on the dataset page: https://huggingface.co/datasets/biohub/ESMC-SAE-Features.tabular10K<n<100K5 likes214 downloads4mo agoHugging Face20saeeew /JP-HomophoneBench JP-HomophoneBench A deterministic Japanese ASR benchmark index for separating eight error/disambiguation classes: exact_homophone near_homophone voicing long_vowel geminate moraic_nasal pitch_accent semantic_only Important design rule This repository is metadata-first. Source audio is not redistributed by default. Each row stores source repository/config/split/row identifiers so audio can be rehydrated under the original source license. exact_homophone and… See the full description on the dataset page: https://huggingface.co/datasets/saeeew/JP-HomophoneBench.audioautomatic-speech-recognitionn<1K0 likes194 downloads25d agoHugging Face21saeedzou /iemocap-original-wavlm-large-layer-9-temporalaudio1K<n<10K0 likes175 downloads2mo agoHugging Face22saeedzou /iemocap-vc-wavlm-large-layer-9-temporalaudio1K<n<10K0 likes175 downloads2mo agoHugging Face23juiceb0xc0de /smollm2-135m-instruct-SAE Layer EV Mean L0 Recon Loss Dead % 0 0.9480 48.74 0.2074 0.0 1 0.9599 43.65 0.3298 0.0 2 0.9631 46.81 0.5021 0.0 3 0.9508 46.56 0.7462 0.0 4 0.9463 46.23 0.8936 0.0 5 0.9350 47.57 1.1605 0.0 6 0.9306 48.44 1.3838 0.0 7 0.9318 49.51 1.5446 0.0 8 0.9432 46.52 1.6598 0.0 9 0.9373 47.15 2.0706 0.0 10 0.9348 45.53 2.2983 0.0 11 0.9905 48.58 5.8113 0.0 12 0.9901 48.42 6.1039 0.0 13 0.9891 46.15 6.9692 0.0 14 0.9884 44.76 7.1844 0.0 15 0.9863 47.63 8.6521 0.0… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/smollm2-135m-instruct-SAE.tabularfeature-extractionn<1K0 likes158 downloads3mo agoHugging Face24omrifahn /mfa-vs-sae-2026-webapp-datatabularn<1K0 likes153 downloads8mo agoHugging Face25SaeedLab /MolDeBERTa PubChem Dataset for MolDeBERTa Pretraining MolDeBERTa: Foundational Model for Physicochemical and Substructure-Informed Molecular Representation Learning [Paper] | [Github Repo] | [Model Collection] | [Cite] Abstract Foundational models that learn the "language" of molecules are essential for accelerating material and drug discovery. These self-learning models can be trained on large collections of unlabelled molecules, enabling applications such as property… See the full description on the dataset page: https://huggingface.co/datasets/SaeedLab/MolDeBERTa.text100M<n<1B0 likes150 downloads4mo agoHugging Face26saeedzou /e-daic-ai-controlledgatedaudio1K<n<10K0 likes149 downloads2mo agoHugging Face27saeidseyfi /khattat Khattat Synthetic multilingual handwritten OCR dataset by saeidseyfi. Handwritten samples in Persian (fa), Arabic (ar) and English (en) rendered with a dynamic pen (variable stroke width and pressure along the writing path), with three configs: Configs 1. default — line crops (train 21,726 / validation 2,062 / test 924) Handwritten line images with ground-truth transcriptions. Column Type Description image image RGB handwritten line crop… See the full description on the dataset page: https://huggingface.co/datasets/saeidseyfi/khattat.imageimage-to-text10K<n<100K0 likes148 downloads9d agoHugging Face28saeedzou /iemocap-original-wavlm-layer-6-temporalaudio1K<n<10K0 likes143 downloads2mo agoHugging Face29saeedzou /iemocap-vc-wavlm-layer-6-temporalaudio1K<n<10K0 likes129 downloads2mo agoHugging Face30adamkarvonen /chess_sae_individual_games_filteredtext100K<n<1M0 likes128 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.