CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01japanese-asr /whisper_transcriptions.reazon_speech_all.wer_10.0.vectorized1M<n<10M0 likes85k downloads2y agoHugging Face02japanese-asr /whisper_transcriptions.mls.wer_10.0.vectorized1M<n<10M1 likes36k downloads2y agoHugging Face03philippesaade /Wikidata_Vectors_0.2 Wikidata Entity Embeddings 0.2 Dataset Summary Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.textfeature-extraction10M<n<100M3 likes5.7k downloads28d agoHugging Face04abotresol /emotion-vectors-gemma-4-31b-it-postfix Emotion vectors, google/gemma-4-31b-it (corrected extraction) Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.feature-extraction0 likes5.6k downloads2mo agoHugging Face05vector-index-bench /vibeThis repository contains the datasets presented in VIBE: Vector Index Benchmark for Embeddings: https://github.com/vector-index-bench/vibe The datasets can be downloaded manually from this repository, but the benchmark framework also downloads them automatically. Datasets In-distribution datasets Name Type n d Distance agnews-mxbai-1024-euclidean Text 769,382 1024 euclidean arxiv-nomic-768-normalized Text 1,344,643 768 any dpr-jina-768-normalized… See the full description on the dataset page: https://huggingface.co/datasets/vector-index-bench/vibe.sentence-similarity2 likes3.8k downloads1mo agoHugging Face06lu-christina /assistant-axis-vectors The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models This repository contains pre-computed axes and persona vectors for Gemma 2 27B, Qwen 3 32B, and Llama 3.3 70B, as described in the paper The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models. Paper | Code | Demo The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's current persona is. It can be used to: Monitor persona… See the full description on the dataset page: https://huggingface.co/datasets/lu-christina/assistant-axis-vectors.other12 likes2.4k downloads8mo agoHugging Face07abotresol /emotion-vectors-gemma-4-31b-it Emotion vectors — gemma-4-31b-it (instruct) probed on the external gemma-4-4B story corpus Data provenance (what made these activations) Probed model: google/gemma-4-31b-it (instruct) Input corpus: snae/emotion_stories_gemma_4_4B — stories written by gemma-4-4B, a smaller EXTERNAL model (generator is NOT the probed model) Per-story pooled residual-stream activations and per-emotion mean vectors, extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it.feature-extraction0 likes2k downloads2mo agoHugging Face08abotresol /emotion-vectors-gemma-4-31b Emotion vectors — gemma-4-31b (base) probed on the external gemma-4-4B story corpus Data provenance (what made these activations) Probed model (whose activations these are): google/gemma-4-31b (base) Input corpus: snae/emotion_stories_gemma_4_4B — third-person emotion stories written by gemma-4-4B, a smaller EXTERNAL model (the open replication's published corpus; generator is NOT the probed model) Per-story pooled residual-stream activations and per-emotion… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b.feature-extraction0 likes2k downloads2mo agoHugging Face09Ayzengin /v32-vectorstext1M<n<10M0 likes1.9k downloads16d agoHugging Face10abotresol /emotion-vectors-gemma-4-31b-postfix Emotion vectors, google/gemma-4-31b (corrected extraction) Residual-stream activations for google/gemma-4-31b, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-vectors-gemma-4-31b. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This set re-extracts… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-postfix.feature-extraction0 likes1.8k downloads2mo agoHugging Face11vector-institute /open-pmc-18m OPEN-PMC Arxiv: Arxiv     |     Code: Open-PMC Github     |     Model Checkpoint: Hugging Face Dataset Summary This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes: Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.image10M<n<100M6 likes1.7k downloads4mo agoHugging Face12japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0.vectorized1M<n<10M0 likes1.6k downloads2y agoHugging Face13VectorSQL /VectorSQLBench0 likes1.5k downloads11mo agoHugging Face14Maktabati /openiti-vectors OpenITI Vector Database — maktabati.ai 🇬🇧 English This dataset contains the fully vectorized OpenITI RELEASE 2025-1-9 collection of classical Islamic texts, prepared for semantic search (RAG). Each entry represents a text chunk from one of 8,943 works in the OpenITI corpus, together with its embedding vector and complete metadata. Statistics: 4,696,703 chunks 8,943 works (primary editions only, status=pri from OpenITI TSV) approx. 470 Parquet files (approx.… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/openiti-vectors.tabular1M<n<10M0 likes1.2k downloads4mo agoHugging Face15abotresol /emotion-selfstory-vectors-gemma-4-31b-it-postfix Emotion vectors, google/gemma-4-31b-it (corrected extraction) Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-selfstory-vectors-gemma-4-31b-it. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it-postfix.feature-extraction0 likes1.2k downloads2mo agoHugging Face16abotresol /emotion-selfstory-vectors-gemma-4-31b-it Emotion vectors — gemma-4-31b-it probed on its OWN self-generated stories Data provenance (what made these activations) Probed model: google/gemma-4-31b-it (instruct) Input corpus: abotresol/emotion-stories-gemma-4-31b-it — stories written by the probed model itself (generator = probed model, the reference's convention; 12 the twelve emotions, up to 256 stories each — the E6 scale corpus) Per-story pooled residual-stream activations and per-emotion mean vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it.feature-extraction0 likes804 downloads2mo agoHugging Face17abotresol /emotion-dialogue-vectors-gemma-4-31b Emotion vectors — gemma-4-31b (base) probed on base-generated dialogues Data provenance (what made these activations) Probed model: google/gemma-4-31b (base) Input corpus: abotresol/emotion-dialogues-gemma-4-31b — two-person dialogues written by the base model (generator = probed model; 44% emotion-word leakage, documented) Per-story pooled residual-stream activations and per-emotion mean vectors, extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-dialogue-vectors-gemma-4-31b.feature-extraction0 likes768 downloads2mo agoHugging Face18DB-Edinburgh /VectorBenchmarkgatedtabularn<1K3 likes679 downloads2d agoHugging Face19vector-institute /sonic-o1 &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding 🎯 What is SONIC-O1? The first open-source benchmark for evaluating omnimodal video understanding with systematic fairness analysis. SONIC-O1 requires models to jointly process audio, video, and social context from real-world interactions—not just transcripts. Key… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/sonic-o1.audiovisual-question-answering1K<n<10K5 likes657 downloads4mo agoHugging Face20abotresol /emotion-deepseek-vectors-gemma-4-31b-it Emotion vectors from fixed-prompt DeepSeek stories Per-emotion vectors for google/gemma-4-31b-it, built from stories written by deepseek-v4-pro under one fixed instruction. These were the strongest detection vectors in the project's comparison of story sources: 9 of 20 layers cleared a bar fixed before scoring, against 5 for the model's own writing. 12 emotions, 20 layers, 5,376 dimensions per layer, from 3,070 stories. Contents Path Shape Contents… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-deepseek-vectors-gemma-4-31b-it.feature-extraction0 likes545 downloads2mo agoHugging Face21Maktabati /shamela-vectors 🇬🇧 English The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim). Statistics: 11,482,164 chunks from 8,589 classical Islamic books 6,236 Quran verses (one verse = one chunk, included in total) 40 categories covering the full breadth of Islamic scholarship Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/shamela-vectors.tabularfeature-extraction10M<n<100M2 likes535 downloads4mo agoHugging Face22aandreychuk /MAPF-GPT-vectors100M<n<1B0 likes533 downloads1y agoHugging Face23vector-institute /open-pmc OPEN-PMC Arxiv: Arxiv     |     Code: Open-PMC Github     |     Model Checkpoint: Hugging Face Dataset Summary This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes: Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc.image1M<n<10M9 likes503 downloads1y agoHugging Face24abotresol /neutral-vectors-gemma-4-31b-it-postfix Emotion vectors, google/gemma-4-31b-it (corrected extraction) Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/neutral-vectors-gemma-4-31b-it. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/neutral-vectors-gemma-4-31b-it-postfix.feature-extraction0 likes500 downloads2mo agoHugging Face25Yettiesoft /voice_medical_cut_medium_vector10K<n<100K0 likes484 downloads2y agoHugging Face26DataDrivenConstruction /cwicr-vector-db-bgem3-v3 CWICR Vector Database — BGE-M3 V3 Snapshots Production Qdrant snapshots for CWICR (Construction Works Items, Costs & Resources) — a multilingual catalogue of construction rate databases covering 30 countries / language locales. Each snapshot encodes one country's rate book using the BAAI/bge-m3 embedder and is ready to restore directly into a Qdrant server for hybrid semantic search. These snapshots are the V3 production artifacts produced by the OpenConstructionEstimate / CWICR… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-vector-db-bgem3-v3.imagen<1K0 likes470 downloads4mo agoHugging Face27Kandil7 /shamela-vectors 🇬🇧 English The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim). Statistics: 11,482,164 chunks from 8,589 classical Islamic books 6,236 Quran verses (one verse = one chunk, included in total) 40 categories covering the full breadth of Islamic scholarship Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/shamela-vectors.tabularfeature-extraction10M<n<100M2 likes411 downloads3mo agoHugging Face28vector-institute /s2ef-15m Dataset Description This dataset contains a collection of 3D atomistic datasets with force and energy labels gathered from a series of sources: Open Catalyst Project OC20, OC22, ODAC23 Materials Project Trajectory Dataset (MPtrj) SPICE 1.1.4 Dataset Structure Data Instances For each instance, there is set of atomic numbers (input_ids), 3-D coordinates (coords), a set of forces per atom (forces), the total and formation energy per system… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/s2ef-15m.tabular10M<n<100M0 likes410 downloads2y agoHugging Face29goea /arc_whisper_transcriptions.reazonspeech.small.wer_10.0.vectorized Dataset Card for "arc_whisper_transcriptions.reazonspeech.small.wer_10.0.vectorized" More Information needed 10K<n<100K0 likes400 downloads1y agoHugging Face30moss-vector-714 /GeoFidelity-Bench GeoFidelity-Bench GeoFidelity-Bench evaluates whether generated street-view images match a requested location at the level of named street blocks. The release contains 109 named street blocks from 25 cities, 7,117 curated Mapillary reference images, generated images from six open-weight text-to-image models, prompt control metadata, and 109-target benchmark result summaries. The generated-image index covers 15,696 released JPEG files across six models, six prompt or control… See the full description on the dataset page: https://huggingface.co/datasets/moss-vector-714/GeoFidelity-Bench.imagetext-to-image1K<n<10K0 likes397 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.