CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CNX-PathLLM /Llama-slideQA-Sample-Featurestextn<1K0 likes746 downloads4mo agoHugging Face02SHENJJ1017 /morph_features UniMorph + UniSegments Morph Data This dataset pairs UniMorph inflectional features with UniSegments segmentations. For languages without UniSegments coverage, segmentation defaults to the unsegmented word form itself. This resource is a necessary component for evaluating Tokenizer Morphological Plausibility, as introduced in Tokenizer Morphological Plausibility (https://arxiv.org/abs/2601.18536). The data generation process follows the implementation provided in the official… See the full description on the dataset page: https://huggingface.co/datasets/SHENJJ1017/morph_features.texttoken-classification10M<n<100M1 likes633 downloads7mo agoHugging Face03ClSu /ember-features EMBER precomputed features Concept features for EMBedding ERasure (EMBER), a plug-and-play module that uses Sparse Matrix Factorization to precisely erase concept-related features from token embeddings, making existing erasure methods more robust to relearning. For each concept, two factorizations are provided: Embedding features (EMBER): a sparse factorization of the token-embedding matrix. MLP features (SNMF): Semi-NMF over MLP activations. Models: google/gemma-2-2b-it (rank… See the full description on the dataset page: https://huggingface.co/datasets/ClSu/ember-features.tabular1K<n<10K8 likes329 downloads3mo agoHugging Face04ML-Owl /faang-engineered-time-series-features-2013-2025 FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025) Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!) DOCUMENT NAVIGATION GUIDE (ToC) 1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset 3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.imagetabular-classification10K<n<100K2 likes295 downloads6mo agoHugging Face05hhieupt /mer2026-features MER2026 Track 1 — Quickstart Guide Hướng dẫn từng bước để chạy training và tạo file submission cho MER-Cross (Track 1) sử dụng pre-extracted features tại HuggingFace: hhieupt/mer2026-features. Mục lục Mô tả bài toán và dữ liệu Yêu cầu hệ thống Clone repo ban tổ chức Cài đặt môi trường Tải dữ liệu từ HuggingFace Giải nén và tổ chức thư mục Tạo file config.py Training Tạo file submission Lưu ý và mẹo 1. Mô tả bài toán và dữ liệu Bài… See the full description on the dataset page: https://huggingface.co/datasets/hhieupt/mer2026-features.image1K<n<10K0 likes179 downloads2mo agoHugging Face06shirasko /qwen-snmf-features Qwen3.5 SNMF features for unlearning MLP Semi-NMF factorizations and (when present) LLM interpretations for SNMF concept unlearning on Qwen/Qwen3.5-2B (rank 100, seed 42). These files are the SNMF track only: per-layer MLP directions used to project concept features out of up_proj / down_proj. There is no embedding-matrix factorization in this dataset. Qwen/Qwen2.5-3B-Instruct features previously living in this repo were moved to shirasko/qwen2.5-snmf-features. Layout (same… See the full description on the dataset page: https://huggingface.co/datasets/shirasko/qwen-snmf-features.tabularn<1K0 likes126 downloads15d agoHugging Face07luckeciano /mistral8x22b-features-reddittabular100K<n<1M0 likes91 downloads2y agoHugging Face08zer0int /CLIP-ViT-L-14-336-L20-features OpenAI/CLIP-ViT-L/14@336 Layer 20 features, CLIP+BLIP labels Feature activation max visualization of the 4096 Features @ L20 CLIP+BLIP labels (may or may not describe what a neuron truly encodes!) ⚠️ May contain sensitive images, albeit abstract. Use responsibly! Examples: image1K<n<10K1 likes68 downloads2y agoHugging Face09luckeciano /pku-llama3.1-8b-answers-features-traintabular1M<n<10M0 likes62 downloads2y agoHugging Face10shirasko /qwen2.5-snmf-features Qwen2.5 SNMF features for unlearning MLP Semi-NMF factorizations and (when present) LLM interpretations for SNMF concept unlearning on Qwen/Qwen2.5-3B-Instruct (rank 100, seed 42). These files are the SNMF track only: per-layer MLP directions used to project concept features out of up_proj / down_proj. There is no embedding-matrix factorization in this dataset. Qwen/Qwen3.5-2B features live in shirasko/qwen-snmf-features. Layout (same directory scheme used by the… See the full description on the dataset page: https://huggingface.co/datasets/shirasko/qwen2.5-snmf-features.tabularn<1K0 likes61 downloads16d agoHugging Face11ALIkjjnskdjc /umerkot-aqi-featurestabular1K<n<10K0 likes56 downloads17d agoHugging Face12vancenceho /youtube-spotify-audio-features Spotify–YouTube Audio Features Tabular librosa audio features for tracks aligned with the Spotify / YouTube pipeline in the viral-content-predictor project. Each row is one Spotify track_id matched to a downloaded YouTube audio clip; features are aggregated statistics (mean / std) computed on the decoded waveform. Files File Description audio_features.csv One row per track: track_id, 89 derived feature dimensions (means/stds), extraction_success, error_message.… See the full description on the dataset page: https://huggingface.co/datasets/vancenceho/youtube-spotify-audio-features.tabular10K<n<100K0 likes48 downloads5mo agoHugging Face13SuodhanJ6 /elliptic_txs_featurestabular100K<n<1M0 likes46 downloads3y agoHugging Face14abdlh /Dataset_Automatic_Essay_Scoring_Essay-EssayScore_and_24_textual_featurestabular10K<n<100K2 likes46 downloads2y agoHugging Face15Jackdsada /nba_game_featurestabular100K<n<1M0 likes46 downloads10mo agoHugging Face16luckeciano /mistral8x22b-reddit-post-featurestabular10K<n<100K0 likes45 downloads2y agoHugging Face17kasunUdayanga /Tea_yield_6_features Tea Yield Prediction Dataset (6 Features) 📋 Quick Info Samples: 53,264 Features: 6 Task: Regression (predict tea yield) Type: Synthetic (realistic simulation) 🎯 Purpose Simple dataset for machine learning beginners to practice: Data preprocessing (missing values, outliers) Feature engineering Regression modeling Model evaluation 📊 Features # Feature Description Range 1 rainfall_mm Annual rainfall in mm 10-350 2 temperature_avg… See the full description on the dataset page: https://huggingface.co/datasets/kasunUdayanga/Tea_yield_6_features.tabulartabular-regression10K<n<100K0 likes45 downloads8mo agoHugging Face18julia-lukasiewicz-pater /small-GPT-wiki-intro-features Small-GPT-wiki-intro-features dataset This dataset is based on aadityaubhat/GPT-wiki-intro. It contains 100k randomly selected texts (50k from Wikipedia and 50k generated by ChatGPT). For each text, various complexity measures were calculated, including e.g. readibility, lexical richness etc. It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts. Dataset structure Features were calculated using… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/small-GPT-wiki-intro-features.tabulartext-classification100K<n<1M0 likes39 downloads3y agoHugging Face19luckeciano /hermes-reddit-post-featurestabular10K<n<100K0 likes39 downloads2y agoHugging Face20generality-labs /glm_features Per-transcript GLM features Deterministic features of every transcript in LASR-G5/benchmarks, for the transcript-regression protocols (glm_length, glm_regex) in small-judge. Produced by scripts/glm_features/extract.py from the benchmark logs at commit c199498b; no generative model is involved. Layout Mirrors the benchmarks repo, one CSV per successful full-run log: <benchmark>/<task_args_hash>/<model>/<repeat>/<eval id>.csv Crashed or partial logs kept beside a… See the full description on the dataset page: https://huggingface.co/datasets/generality-labs/glm_features.tabular1K<n<10K0 likes39 downloads6d agoHugging Face21harryxi /PKU-SafeRLHF-Prompts-Shift-answer-train-featurestabular100K<n<1M0 likes35 downloads1y agoHugging Face22krishan-CSE /HatEval_Relabled_with_Author_Featurestabular10K<n<100K0 likes28 downloads3y agoHugging Face23jang1563 /cbrn-physics-features CBRN Physics Features Pre-computed physics-informed distributional features for pathogen-agnostic biological threat detection in gene expression data. Overview This dataset contains per-sample and per-group features computed from the shape of gene expression distributions rather than the identity of individual genes. The four core features — Gini coefficient, Shannon entropy, normalized entropy, and Zipf exponent — are platform-agnostic: they require no gene… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/cbrn-physics-features.tabulartabular-classificationn<1K0 likes28 downloads2mo agoHugging Face24luckeciano /pku-llama3.1-8b-answers-features-testtabular1M<n<10M0 likes26 downloads2y agoHugging Face25harryxi /PKU-SafeRLHF-Prompts-Shift-alpaca-3-8b-answers-features-traintabular1M<n<10M0 likes24 downloads1y agoHugging Face26RossNeyman /fma-merged-metadata-and-featurestabular100K<n<1M0 likes23 downloads10mo agoHugging Face27luckeciano /llama370b-reddit-post-featurestabular10K<n<100K0 likes22 downloads2y agoHugging Face28julia-lukasiewicz-pater /GPT-wiki-intro-features Small-GPT-wiki-intro-features dataset This dataset is based on aadityaubhat/GPT-wiki-intro. It contains 150k short texts from Wikipedia (label 0) and corresponding texts generated by ChatGPT (label 1) (together 300k texts). For each text, various complexity measures were calculated, including e.g. readability, lexical diversity etc. It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts. For a smaller version… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/GPT-wiki-intro-features.tabulartext-classification100K<n<1M1 likes12 downloads3y agoHugging Face29luckeciano /hermes-features-ultrafeedbacktabular10K<n<100K0 likes12 downloads3y agoHugging Face30jason1966 /aadigupta1601_ai-vs-human-art-interpretable-numerical-features AI vs Human Art — Interpretable Numerical Features An Interpretable Numerical Representation of Visual Artworks for Comparative Ana Dataset Info Source: Kaggle Original Size: 0.08 MB Kaggle Downloads: 37 Files: 1 Files ai_vs_human_art_interpretable_numeric.csv Mirrored from Kaggle tabularn<1K0 likes12 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.