CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01W8Yi /tcga-wsi-uni2h-features TCGA WSI UNI2H Features Dataset Summary This dataset provides tile-level UNI2-h embeddings extracted from TCGA whole-slide images (WSIs) using a reproducible, auditable pipeline designed for computational pathology research. Data is organized by project (for example TCGA-HNSC) and currently exposes: features/ containing H5 feature files with tile-level embeddings vis/ containing overlay images for quality inspection and pipeline verification [!IMPORTANT] Unlike the… See the full description on the dataset page: https://huggingface.co/datasets/W8Yi/tcga-wsi-uni2h-features.imageimage-feature-extractionn<1K12 likes49k downloads6mo agoHugging Face02davidscripka /openwakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library. Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models. The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google. openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/openwakeword_features.2 likes16k downloads3y agoHugging Face03erl-hub /behaviour1k-Qwen3-features BEHAVIOR-1K Qwen3 skill features Per-frame conditioned features e_t = Phi(f_t, L_sub^(j), L), mean-pooled primitive skill latents S_j, aligned proprioception q_t, actions a_t, and subtask progress p_t. These are the inputs and targets for a Primitive Skill Composer VLA Skill Predictor. Ground-truth primitives come from BEHAVIOR-1K's hand-authored primitive_annotation, so the segmentation is human-labelled rather than predicted, and nothing here depends on a keyframe detector.… See the full description on the dataset page: https://huggingface.co/datasets/erl-hub/behaviour1k-Qwen3-features.tabularrobotics1K<n<10K0 likes15k downloads2mo agoHugging Face04AbstractPhil /bulk-cc12m-features bulk-cc12m-features — ten teacher towers over CC12M, plus their consensus Precomputed image-tower features for 10,968,539 CC12M images (all 2,176 shards of pixparse/cc12m-wds) from ten independent teacher extractions — eight CLIP variants across three pretraining corpora and two model scales, SigLIP, and DINOv3 — plus one derived consensus target. About 110 million feature vectors, roughly 130 GPU-hours of extraction, so that a student can be distilled against any of these… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-cc12m-features.text100M<n<1B0 likes3.5k downloads2mo agoHugging Face05vinhthuanly /tinygiant-omni-features0 likes3.4k downloads26d agoHugging Face06binhpham /livekit_wakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library. Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models. The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google. openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/livekit_wakeword_features.0 likes2.9k downloads6mo agoHugging Face07SoccerNet /SN-Features SoccerNet Features Pre-extracted per-game features for the SoccerNet benchmark, structured as <league>/<season>/<game>/<file>, one file per game half (1_.../2_...). This main branch holds no data — each feature type lives on its own branch so you only download what you need: Branch Files Description baidu-soccer-embeddings {1,2}_baidu_soccer_embeddings.npy Frame embeddings from baidu-research/vidpress-sports, used by the Action Spotting and Dense Video Captioning 2023… See the full description on the dataset page: https://huggingface.co/datasets/SoccerNet/SN-Features.other0 likes2.7k downloads15d agoHugging Face08MahmoodLab /UNI2-h-featuresgated Dataset Card for UNI2-h Pretrained Features This dataset card provides the UNI2-h features for TCGA, CPTAC, and PANDA datasets with patch size 256 x 256 pixels at 20x magnification. Requesting Access As mentioned in the gated prompt, you must agree to the outlined terms of use, with the primary email for your HuggingFace account matching your institutional email. If your primary email is a personal email (@gmail/@hotmail/@qq) your request will be denied. To fix this, you… See the full description on the dataset page: https://huggingface.co/datasets/MahmoodLab/UNI2-h-features.33 likes2.4k downloads2y agoHugging Face09snad-space /ztf-dr3-m31-featurestabular10K<n<100K0 likes2.1k downloads2y agoHugging Face10s-nlp /Mintaka_Graph_Features_T5-xl-ssm Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm" More Information needed tabular100K<n<1M0 likes1.9k downloads2y agoHugging Face11Embodied-CoT /embodied_features_and_demos_liberoDataset for Embodied Chain-of-Thought Reasoning for LIBERO-90, as used by ECoT-Lite. TFDS Demonstration Data The TFDS dataset contains successful demonstration trajectories for LIBERO-90 (50 trajectories for each of 90 tasks). It was created by rolling out the actions provided in the original LIBERO release and filtering out all unsuccessful ones, leaving 3917 successful demo trajectories. This is done via a modified version of a script from the MiniVLA codebase. In addition to… See the full description on the dataset page: https://huggingface.co/datasets/Embodied-CoT/embodied_features_and_demos_libero.robotics4 likes1.8k downloads6mo agoHugging Face12sjmathy /vitra-dinotxt-features0 likes1.8k downloads2mo agoHugging Face13DavidErikMollberg /precomputed_audio_features0 likes1.8k downloads1y agoHugging Face14Emanresu /features-dinov3-vith16plus-224-imagenet-22k-wdstext1M<n<10M0 likes1.6k downloads11mo agoHugging Face15acroitoru /features_mavos_complete0 likes1.4k downloads11mo agoHugging Face16myzhao1999 /ucf-crime-clip-features0 likes1.3k downloads2y agoHugging Face17quchenyuan /360x_dataset_features0 likes1.1k downloads2y agoHugging Face18zhaoshiwen /vtg-featuresvideon<1K0 likes922 downloads4mo agoHugging Face19Kavindu1124 /ucf-crime-processed-features0 likes909 downloads21d agoHugging Face20Koa-Chang /TissueMNIST-224-full-gpt5nano-with-vlm-features TissueMNIST 224 Full Train Val with GPT-5-nano VLM Features The full TissueMNIST train and validation splits with categorical morphology features generated by GPT-5-nano. Test is included as the full TissueMNIST passthrough split with null vlm_model_name and placeholder vlm_feature values for schema consistency. This dataset is derived from the official MedMNIST TissueMNIST 224px data. The VLM feature labels are categorical privileged-information annotations for CS231N VLM-LUPI… See the full description on the dataset page: https://huggingface.co/datasets/Koa-Chang/TissueMNIST-224-full-gpt5nano-with-vlm-features.text100K<n<1M0 likes883 downloads4mo agoHugging Face21hungphongtrn /tallyqa_extracted_featurestext10K<n<100K0 likes760 downloads1y agoHugging Face22CNX-PathLLM /Llama-slideQA-Sample-Featurestextn<1K0 likes746 downloads4mo agoHugging Face23medarc /AlgonautsDS-features Saved Features for Algonauts '25 Dataset This repository contains pre-extracted features for the Algonauts Challenge dataset using baseline models. Features Overview The developer_kit directory contains features extracted for the entire dataset using the following models: Video Features Model: SlowFast R50 Extracts spatiotemporal features from video frames Captures motion and appearance information Audio Features Model: MFCC (Mel-frequency… See the full description on the dataset page: https://huggingface.co/datasets/medarc/AlgonautsDS-features.text10K<n<100K2 likes734 downloads1y agoHugging Face24mad-bot /mind-games-features Mind Games — Features Pre-computed visual features for the mind-games project. Lunar Lander (MobileNetV3-Small) 200 episodes of MobileNetV3-Small embeddings extracted from LunarLander-v3 gameplay frames. Backbone: MobileNetV3-Small (ImageNet pretrained, frozen) Embedding dim: 576 (global average pooled) Precision: float16 Size: ~108MB Structure: lunar_lander/episode_NNNN/ with embeddings.npy (N×576) and actions.npy (N,) Index: lunar_lander/index.json with per-episode… See the full description on the dataset page: https://huggingface.co/datasets/mad-bot/mind-games-features.reinforcement-learning0 likes650 downloads6mo agoHugging Face25SHENJJ1017 /morph_features UniMorph + UniSegments Morph Data This dataset pairs UniMorph inflectional features with UniSegments segmentations. For languages without UniSegments coverage, segmentation defaults to the unsegmented word form itself. This resource is a necessary component for evaluating Tokenizer Morphological Plausibility, as introduced in Tokenizer Morphological Plausibility (https://arxiv.org/abs/2601.18536). The data generation process follows the implementation provided in the official… See the full description on the dataset page: https://huggingface.co/datasets/SHENJJ1017/morph_features.texttoken-classification10M<n<100M1 likes633 downloads7mo agoHugging Face26thienhtt20 /test-extracted-features-with-idlabel0 likes632 downloads2y agoHugging Face27Kshitijbhatt1998 /ieee-fraud-detection-pipeline-features IEEE-CIS Fraud Detection — Model-Ready Feature Dataset Dataset Summary This dataset contains 590,540 financial transactions from the IEEE-CIS Fraud Detection competition, processed through a production-grade data pipeline into a clean, labeled, model-ready feature table. The pipeline adds 16 engineered features on top of the original IEEE-CIS columns — including card-level velocity, email domain risk rates, and amount anomaly ratios — ready for direct use in fraud… See the full description on the dataset page: https://huggingface.co/datasets/Kshitijbhatt1998/ieee-fraud-detection-pipeline-features.tabulartabular-classification100K<n<1M1 likes616 downloads6mo agoHugging Face28Adam1010 /gemma-4-31b-sae-features Gemma-4-31B Sparse Autoencoder Features 3,000 interpreted and verified SAE features across all 60 layers of Google's Gemma-4-31B-IT model. What's in this dataset? For each of the 60 transformer layers in Gemma-4-31B, we trained a TopK-64 Sparse Autoencoder with 43,008 features (8x expansion from d_model=5376). We then selected the 50 most interesting features per layer using SIPIT (Sparse Input-Token Invertibility Probe) scores, interpreted them with two independent LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Adam1010/gemma-4-31b-sae-features.feature-extraction1K<n<10K2 likes579 downloads5mo agoHugging Face29Peacockery /librispeech-phoneme-featurestabular100K<n<1M0 likes555 downloads7mo agoHugging Face30ozefe /spotify_audio_features Spotify Tracks & Audio Features Dataset Overview This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research. Data Source The raw data for this dataset was originally gathered and hosted by Anna's Archive. Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/ozefe/spotify_audio_features.tabulartabular-regression100M<n<1B11 likes535 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.