features
Datasets
All datasets matching “features”tcga-wsi-uni2h-features
TCGA WSI UNI2H Features
Dataset Summary
This dataset provides tile-level UNI2-h embeddings extracted from TCGA whole-slide images (WSIs) using a reproducible, auditable pipeline designed for computational pathology research.
Data is organized by project (for example TCGA-HNSC) and currently exposes:
features/ containing H5 feature files with tile-level embeddings
vis/ containing overlay images for quality inspection and pipeline verification
[!IMPORTANT]
Unlike the… See the full description on the dataset page: https://huggingface.co/datasets/W8Yi/tcga-wsi-uni2h-features.openwakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library.
Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models.
The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google.
openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/openwakeword_features.behaviour1k-Qwen3-features
BEHAVIOR-1K Qwen3 skill features
Per-frame conditioned features e_t = Phi(f_t, L_sub^(j), L), mean-pooled primitive skill latents S_j, aligned proprioception q_t, actions a_t, and subtask progress p_t.
These are the inputs and targets for a Primitive Skill Composer VLA Skill Predictor.
Ground-truth primitives come from BEHAVIOR-1K's hand-authored
primitive_annotation, so the segmentation is human-labelled rather than
predicted, and nothing here depends on a keyframe detector.… See the full description on the dataset page: https://huggingface.co/datasets/erl-hub/behaviour1k-Qwen3-features.bulk-cc12m-features
bulk-cc12m-features — ten teacher towers over CC12M, plus their consensus
Precomputed image-tower features for 10,968,539 CC12M images (all 2,176
shards of
pixparse/cc12m-wds)
from ten independent teacher extractions — eight CLIP variants across
three pretraining corpora and two model scales, SigLIP, and DINOv3 — plus
one derived consensus target.
About 110 million feature vectors, roughly 130 GPU-hours of extraction,
so that a student can be distilled against any of these… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-cc12m-features.tinygiant-omni-featureslivekit_wakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library.
Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models.
The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google.
openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/livekit_wakeword_features.
