datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama-slideQA-Sample-Featuresmorph_features
UniMorph + UniSegments Morph Data
This dataset pairs UniMorph inflectional features with UniSegments segmentations. For languages without UniSegments coverage, segmentation defaults to the unsegmented word form itself.
This resource is a necessary component for evaluating Tokenizer Morphological Plausibility, as introduced in Tokenizer Morphological Plausibility (https://arxiv.org/abs/2601.18536). The data generation process follows the implementation provided in the official… See the full description on the dataset page: https://huggingface.co/datasets/SHENJJ1017/morph_features.ember-features
EMBER precomputed features
Concept features for EMBedding ERasure (EMBER), a plug-and-play module that uses
Sparse Matrix Factorization to precisely erase concept-related features from token
embeddings, making existing erasure methods more robust to relearning.
For each concept, two factorizations are provided:
Embedding features (EMBER): a sparse factorization of the token-embedding matrix.
MLP features (SNMF): Semi-NMF over MLP activations.
Models: google/gemma-2-2b-it (rank… See the full description on the dataset page: https://huggingface.co/datasets/ClSu/ember-features.faang-engineered-time-series-features-2013-2025
FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025)
Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!)
DOCUMENT NAVIGATION GUIDE (ToC)
1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset
3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.mer2026-features
MER2026 Track 1 — Quickstart Guide
Hướng dẫn từng bước để chạy training và tạo file submission cho MER-Cross (Track 1) sử dụng pre-extracted features tại HuggingFace: hhieupt/mer2026-features.
Mục lục
Mô tả bài toán và dữ liệu
Yêu cầu hệ thống
Clone repo ban tổ chức
Cài đặt môi trường
Tải dữ liệu từ HuggingFace
Giải nén và tổ chức thư mục
Tạo file config.py
Training
Tạo file submission
Lưu ý và mẹo
1. Mô tả bài toán và dữ liệu
Bài… See the full description on the dataset page: https://huggingface.co/datasets/hhieupt/mer2026-features.qwen-snmf-features
Qwen3.5 SNMF features for unlearning
MLP Semi-NMF factorizations and (when present) LLM interpretations for
SNMF concept unlearning on Qwen/Qwen3.5-2B (rank 100, seed 42).
These files are the SNMF track only: per-layer MLP directions used to project
concept features out of up_proj / down_proj. There is no embedding-matrix
factorization in this dataset.
Qwen/Qwen2.5-3B-Instruct features previously living in this repo were moved to
shirasko/qwen2.5-snmf-features.
Layout (same… See the full description on the dataset page: https://huggingface.co/datasets/shirasko/qwen-snmf-features.CLIP-ViT-L-14-336-L20-features
OpenAI/CLIP-ViT-L/14@336 Layer 20 features, CLIP+BLIP labels
Feature activation max visualization of the 4096 Features @ L20
CLIP+BLIP labels (may or may not describe what a neuron truly encodes!)
⚠️ May contain sensitive images, albeit abstract. Use responsibly!
Examples:
qwen2.5-snmf-features
Qwen2.5 SNMF features for unlearning
MLP Semi-NMF factorizations and (when present) LLM interpretations for
SNMF concept unlearning on Qwen/Qwen2.5-3B-Instruct (rank 100, seed 42).
These files are the SNMF track only: per-layer MLP directions used to project
concept features out of up_proj / down_proj. There is no embedding-matrix
factorization in this dataset.
Qwen/Qwen3.5-2B features live in
shirasko/qwen-snmf-features.
Layout (same directory scheme used by the… See the full description on the dataset page: https://huggingface.co/datasets/shirasko/qwen2.5-snmf-features.umerkot-aqi-featuresyoutube-spotify-audio-features
Spotify–YouTube Audio Features
Tabular librosa audio features for tracks aligned with the Spotify / YouTube pipeline in the viral-content-predictor project. Each row is one Spotify track_id matched to a downloaded YouTube audio clip; features are aggregated statistics (mean / std) computed on the decoded waveform.
Files
File
Description
audio_features.csv
One row per track: track_id, 89 derived feature dimensions (means/stds), extraction_success, error_message.… See the full description on the dataset page: https://huggingface.co/datasets/vancenceho/youtube-spotify-audio-features.nba_game_featuresDataset_Automatic_Essay_Scoring_Essay-EssayScore_and_24_textual_featuresglm_features
Per-transcript GLM features
Deterministic features of every transcript in LASR-G5/benchmarks,
for the transcript-regression protocols (glm_length, glm_regex) in
small-judge. Produced by scripts/glm_features/extract.py
from the benchmark logs at commit c199498b; no generative model is involved.
Layout
Mirrors the benchmarks repo, one CSV per successful full-run log:
<benchmark>/<task_args_hash>/<model>/<repeat>/<eval id>.csv
Crashed or partial logs kept beside a… See the full description on the dataset page: https://huggingface.co/datasets/generality-labs/glm_features.small-GPT-wiki-intro-features
Small-GPT-wiki-intro-features dataset
This dataset is based on aadityaubhat/GPT-wiki-intro.
It contains 100k randomly selected texts (50k from Wikipedia and 50k generated by ChatGPT).
For each text, various complexity measures were calculated, including e.g. readibility, lexical richness etc.
It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts.
Dataset structure
Features were calculated using… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/small-GPT-wiki-intro-features.HatEval_Relabled_with_Author_Featurescbrn-physics-features
CBRN Physics Features
Pre-computed physics-informed distributional features for pathogen-agnostic biological threat detection in gene expression data.
Overview
This dataset contains per-sample and per-group features computed from the shape of gene expression distributions rather than the identity of individual genes. The four core features — Gini coefficient, Shannon entropy, normalized entropy, and Zipf exponent — are platform-agnostic: they require no gene… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/cbrn-physics-features.fma-merged-metadata-and-featuresGPT-wiki-intro-features
Small-GPT-wiki-intro-features dataset
This dataset is based on aadityaubhat/GPT-wiki-intro.
It contains 150k short texts from Wikipedia (label 0) and corresponding texts generated by ChatGPT (label 1) (together 300k texts).
For each text, various complexity measures were calculated, including e.g. readability, lexical diversity etc.
It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts.
For a smaller version… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/GPT-wiki-intro-features.top5_leagues_features_fullOutput-features_10kSemeval_Train_Added_Featurespku-llama3.1-8b-dataset-featuresvcp-combined-features
Spotify–YouTube Combined Ensemble Features
A single-table, modeling-ready CSV that joins Spotify track metadata, librosa audio features (from matched YouTube audio), and YouTube engagement fields on a common key (track_id). Built for ensemble / viral prediction experiments in the viral-content-predictor project (e.g. 03_combined_model_training.ipynb).
File
File
Role
combined_features_cleaned.csv
One row per track (after pipeline joins); mixed numeric… See the full description on the dataset page: https://huggingface.co/datasets/vancenceho/vcp-combined-features.managed-care-features-by-qa-and-performance-incent
Managed Care Features by QA and Performance Incentive
Description
Number of Managed Care Program Types, by Quality Assurance Requirements, Performance Incentives, and Provider Value-Based Purchasing Status, at any point in 2022
Dataset Details
Publisher: Centers for Medicare & Medicaid Services
Last Modified: 2024-10-16
Contact: Medicaid.gov (Medicaid.gov@cms.hhs.gov)
Source
Original data can be found at: https://healthdata.gov/d/5wmq-82nv… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/managed-care-features-by-qa-and-performance-incent.citizen-round-featurespathorchestra-image-features
PathOrchestra Feature Representations
🔒 Access Policy
Access to this dataset is restricted and requires approval.Please request access using your official/institutional email address by contacting the dataset maintainers.
Note: Commercial use is prohibited without explicit permission.
🔄 Dataset Updates
This dataset is under continuous development as part of the broader PathOrchestra project.The current release includes the pancancer_1 subset. Additional… See the full description on the dataset page: https://huggingface.co/datasets/AI4Pathology/pathorchestra-image-features.missense-variant-featuresCHOP_inhibitors_molecular_featuresThe CHOP_inhibitors_molecular_features is a part of the study "Predicting Nanoparticle Effects on Small Biomolecule Functionalities Using the Capability of Scikit-learn and PyTorch: A Case Study on Inhibitors of the DNA Damage-Inducible Transcript 3 (CHOP)"
https://doi.org/10.48550/arXiv.2504.09537
This dataset contains 19,504 rows, each representing a unique small molecule sample. Of these, 7,909 are CHOP inhibitors and 11,592 are not CHOP inhibitors. It includes 8 columns, which are… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/CHOP_inhibitors_molecular_features.youtube-features-clean
YouTube Audio Features Cleaned
Tabular YouTube-side features for tracks that have been matched to Spotify records in the viral-content-predictor pipeline. Rows are keyed for alignment with Spotify / audio / lyrics tables; columns combine metadata and numeric signals used for engagement or popularity modeling (exact schema depends on the exporting notebook revision).
File
File
Role
youtube_features_cleaned.csv
One row per matched track; cleaned dtypes and… See the full description on the dataset page: https://huggingface.co/datasets/vancenceho/youtube-features-clean.emolia_filtered_v1_bb_features
