CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LiberCoders /FeatureBench FeatureBench: Agent Coding Evaluation Benchmark Dataset Description FeatureBench is a comprehensive benchmark designed to evaluate AI agents' capabilities in end-to-end feature-level code generation. Unlike traditional benchmarks that focus on function-level or algorithm-specific tasks, FeatureBench challenges agents to implement complete features within real-world software projects. Key Characteristics Feature-Level Tasks: Each task requires… See the full description on the dataset page: https://huggingface.co/datasets/LiberCoders/FeatureBench.texttext-generationn<1K6 likes8.1k downloads1mo agoHugging Face02AbstractPhil /bulk-cc12m-features bulk-cc12m-features — ten teacher towers over CC12M, plus their consensus Precomputed image-tower features for 10,968,539 CC12M images (all 2,176 shards of pixparse/cc12m-wds) from ten independent teacher extractions — eight CLIP variants across three pretraining corpora and two model scales, SigLIP, and DINOv3 — plus one derived consensus target. About 110 million feature vectors, roughly 130 GPU-hours of extraction, so that a student can be distilled against any of these… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-cc12m-features.text100M<n<1B0 likes4k downloads2mo agoHugging Face03Emanresu /features-dinov3-vith16plus-224-imagenet-22k-wdstext1M<n<10M0 likes2.3k downloads11mo agoHugging Face04Project-AgML /Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset Agri-LLaVA Agri-LLaVA is a large multimodal instruction dataset for agriculture, pairing crop/leaf images with multi-turn diagnostic conversations about plant diseases, pests, and nutrient deficiencies. It is compiled from 16 public source datasets (see the license table below). This dataset has been converted to Parquet format with image bytes embedded directly, standardized to the HF image_text_to_text format with a single conversational messages schema. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset.imageimage-text-to-text100K<n<1M0 likes2.3k downloads2mo agoHugging Face05snad-space /ztf-dr3-m31-featurestabular10K<n<100K0 likes2.2k downloads2y agoHugging Face06s-nlp /Mintaka_Graph_Features_T5-xl-ssm Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm" More Information needed tabular100K<n<1M0 likes2k downloads2y agoHugging Face07Koa-Chang /TissueMNIST-224-full-gpt5nano-with-vlm-features TissueMNIST 224 Full Train Val with GPT-5-nano VLM Features The full TissueMNIST train and validation splits with categorical morphology features generated by GPT-5-nano. Test is included as the full TissueMNIST passthrough split with null vlm_model_name and placeholder vlm_feature values for schema consistency. This dataset is derived from the official MedMNIST TissueMNIST 224px data. The VLM feature labels are categorical privileged-information annotations for CS231N VLM-LUPI… See the full description on the dataset page: https://huggingface.co/datasets/Koa-Chang/TissueMNIST-224-full-gpt5nano-with-vlm-features.text100K<n<1M0 likes995 downloads4mo agoHugging Face08hungphongtrn /tallyqa_extracted_featurestext10K<n<100K0 likes876 downloads1y agoHugging Face09Natt1e /FeatureBench_Litetextn<1K0 likes867 downloads3mo agoHugging Face10vinhthuan /featuretabular100K<n<1M0 likes856 downloads8mo agoHugging Face11Peacockery /librispeech-phoneme-featurestabular100K<n<1M0 likes593 downloads7mo agoHugging Face12jablonkagroup /rdkit_featurestabular10M<n<100M1 likes575 downloads1y agoHugging Face13AbstractPhil /imagenet-clip-features-orderlytimeseries1M<n<10M0 likes546 downloads1y agoHugging Face14mksethi /eli5_sae_features Dataset Card for gpt2_eli5_sae_features This dataset aims to create a corpus of data to help guide research into monosemantic features using SAE's. It has been generated using this raw template Dataset Details Dataset Description This dataset takes the eli5 subset from facebook/kilt_tasks and processes it to be used in SAE research. More specifically, we take the inputs, and tokenize them using the gpt2-small tokenizer. The outputs are tokenized, embedded , and… See the full description on the dataset page: https://huggingface.co/datasets/mksethi/eli5_sae_features.timeseries100K<n<1M0 likes537 downloads1y agoHugging Face15jiachengzhg /FeatureBench-Fast100textn<1K0 likes530 downloads6mo agoHugging Face16cat-claws /face-verification-with-features10K<n<100K2 likes464 downloads5y agoHugging Face17NuBerea /featuresgated NuBerea Features Pre-analytical feature extraction outputs for the NuBerea corpus: Stanza lemma/lexicon/syntax tables for patristic Greek and Latin, SPhilBerta/PhiloBERTa embeddings for patristic segments and biblical verses, and PhiloBERTa lexical alignment primitives. Consolidates configs formerly in NuBerea/patristics-features, NuBerea/concept-primitives, and NuBerea/greek-embeddings-primitives. Configs Patristic Greek / Latin — Syntax… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/features.tabularfeature-extraction1M<n<10M0 likes461 downloads23d agoHugging Face18NuBerea /pericope-featuresgated NuBerea/pericope-features Per-pericope feature tables for the Old and New Testaments. A pericope — the traditional unit of a self-contained biblical passage — is the unit of observation: each row is one pericope, annotated with aggregated signals covering source-critical attribution, tradition-chain and quotation linkage between the Testaments, doublet (repeated-narrative) membership, Hebrew poetic parallelism, and syntactic profile. One config covers New Testament pericopes… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/pericope-features.tabularfeature-extraction1K<n<10K0 likes411 downloads2mo agoHugging Face19mksethi /eli5-gemma-featurestext100K<n<1M0 likes406 downloads1y agoHugging Face20podongchip /kospi-daily-stock-features-2021-2026 KOSPI Stocks Daily Data with Technical & Macro Features 🇰🇷 한국어 / 🇺🇸 English 2021년부터 2026년까지 KOSPI 948개 종목의 일별 시세에 기술적 지표·거시지표·투자자 수급을 결합한 머신러닝 학습용 한국 주식 데이터셋입니다. A machine-learning–ready Korean stock dataset combining daily prices of 948 KOSPI stocks with technical indicators, macro variables, and investor flows, covering 2021 to 2026. 📊 Contents 1,214,339 rows / 121만 행 (948 stocks / 948개 종목) Period / 기간: 2021-01-04 ~ 2026-05-29 (약 5.4년) 60 columns / 60개 컬럼… See the full description on the dataset page: https://huggingface.co/datasets/podongchip/kospi-daily-stock-features-2021-2026.tabulartime-series-forecasting1M<n<10M0 likes404 downloads2mo agoHugging Face21AbstractPhil /bulk-coco-featuresHere exists the bulk prepared sets for coco 2017. With this I will begin testing the first WIDE ViT-Beatrix, ViT-Zana, ViT-Beatrix-DualStream, Clip-Vit-Beatrix, GeoVit-Beans and more. These wide vits will be using new forms of formula meant to fuse structural behaviors together which exist on multiple different manifolds simultaneously. These upcoming experiments will be with established SOTA-based processes adopted and modulated for geofractal behavior from multiple transfer learning… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-coco-features.timeseriesfeature-extraction1M<n<10M0 likes394 downloads9mo agoHugging Face22thedarkknight7 /SAE_monosemanticity_features_4x_0.01_samplingtabular100M<n<1B0 likes373 downloads6mo agoHugging Face23owkin /camelyon16-features Dataset Card for Camelyon16-features Dataset Summary The Camelyon16 dataset is a very popular benchmark dataset used in the field of cancer classification. The dataset we've uploaded here is the result of features extracted from the Camelyon16 dataset using the Phikon model, which is also openly available on Hugging Face. Dataset Creation Initial Data Collection and Normalization The initial collection of the Camelyon16 Whole Slide Images… See the full description on the dataset page: https://huggingface.co/datasets/owkin/camelyon16-features.feature-extractionn<1K1 likes326 downloads3y agoHugging Face24Ishaank18 /screenplay-features Screenplay Scene Salience Features Pre-extracted linguistic and narrative features for screenplay scene salience detection from the MENSA dataset. Dataset Description This dataset contains 913 linguistic features extracted from movie screenplays in the MENSA dataset. Features are organized into 24 feature groups covering various aspects of linguistic, narrative, and discourse analysis. Dataset Statistics Split Samples Size Train 117,503 172.9 MB… See the full description on the dataset page: https://huggingface.co/datasets/Ishaank18/screenplay-features.tabulartext-classification1M<n<10M1 likes305 downloads9mo agoHugging Face25biohub /ESMC-SAE-Features ESMC Sparse Autoencoder Features Table This dataset contains a Parquet table of the 16,384 features from the ESMC-6B-sae-layer60-k64-codebook16384, that was used for analysis in the ESMC paper and to construct the ESM Atlas. This table provides descriptions of the precomputed features that can be activated through the spotlight SAE model, assisting users for downstream interpretation of the insights revealed by ESMC. Download the table here. The features descriptions are in the… See the full description on the dataset page: https://huggingface.co/datasets/biohub/ESMC-SAE-Features.tabular10K<n<100K5 likes274 downloads4mo agoHugging Face26garima-mahato /m5-feature-storetabular10M<n<100M0 likes270 downloads5mo agoHugging Face27AntonKorznikov /feature_stories Feature Stories A large contrastive story dataset for mechanistic interpretability and alignment research. Each row contains two short matched narratives about the same situation: concept_text — written to express one behavioral / affective / epistemic pole antagonist_text — the contrast pole for the same shared setup Feature labels come from a curated concept ontology (148 classes, 1036 concept↔antagonist pairs), with narrative_guidance explaining what each dichotomy means.… See the full description on the dataset page: https://huggingface.co/datasets/AntonKorznikov/feature_stories.texttext-generation1M<n<10M2 likes254 downloads2mo agoHugging Face28thaint /540k-from-phoaudiobook-feature-whisperaudio100K<n<1M1 likes249 downloads1y agoHugging Face29mmwanje /waxal-features-v1 waxal-features-v1 Precomputed Whisper-large-v3 log-mel input_features + tokenized labels for Google WaxalNLP (Lingala, Shona, Luganda). Use this to skip FLAC download + feature extraction when fine-tuning openai/whisper-large-v3 (or any model that consumes the same Whisper-v3 mel / tokenizer layout). Contents Field Type Notes id string Clip id language string lin / sna / lug split string Source split tag input_features list[list[float16]]… See the full description on the dataset page: https://huggingface.co/datasets/mmwanje/waxal-features-v1.text10K<n<100K0 likes248 downloads2mo agoHugging Face30AbstractPhil /imagenet-clip-features Update: 10/2/2025 Claude said that I'm not being careful enough with my database curation after grilling me for 20 minutes, so I included the preparer script as well. Claude Sonnet 4.5 is kind of a chad. Update; 9/26/2025 Having to download this whole repo is annoying, so I'm making sure the splits are named train/val/test (if they exist) and the named subset is the clip name. Older non-dated updates Everything extracted with torch configured as deterministic;… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/imagenet-clip-features.tabularfeature-extraction1M<n<10M1 likes245 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.