CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01philippesaade /Wikidata_Vectors_0.2 Wikidata Entity Embeddings 0.2 Dataset Summary Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.textfeature-extraction10M<n<100M3 likes5.5k downloads1mo agoHugging Face02Ayzengin /v32-vectorstext1M<n<10M0 likes1.9k downloads19d agoHugging Face03Maktabati /openiti-vectors OpenITI Vector Database — maktabati.ai 🇬🇧 English This dataset contains the fully vectorized OpenITI RELEASE 2025-1-9 collection of classical Islamic texts, prepared for semantic search (RAG). Each entry represents a text chunk from one of 8,943 works in the OpenITI corpus, together with its embedding vector and complete metadata. Statistics: 4,696,703 chunks 8,943 works (primary editions only, status=pri from OpenITI TSV) approx. 470 Parquet files (approx.… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/openiti-vectors.tabular1M<n<10M0 likes1.1k downloads4mo agoHugging Face04debajyotidasgupta /vecforge-paper-corpus Note (rebuild in progress): figures are being re-extracted with a fixed extractor (cleaner crops). The image-preview config (Parquet with an inline column + difficulty/type/score labels) returns after re-classification. The config (paper metadata + links) is live now. VecForge Paper Corpus A pristine, deduplicated collection of 59,732 top-venue AI/ML/CV/NLP/Robotics papers (2020-2024) with every captioned figure and full paper text, for research on figure understanding… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/vecforge-paper-corpus.imageimage-to-text10K<n<100K0 likes598 downloads4mo agoHugging Face05Maktabati /shamela-vectors 🇬🇧 English The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim). Statistics: 11,482,164 chunks from 8,589 classical Islamic books 6,236 Quran verses (one verse = one chunk, included in total) 40 categories covering the full breadth of Islamic scholarship Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/shamela-vectors.tabularfeature-extraction10M<n<100M2 likes585 downloads4mo agoHugging Face06EMBO /soda-vec-data-full_pmc_title_abstract SODA-VEC Clean Dataset This is a cleaned and filtered version of the SODA-VEC dataset, containing high-quality biomedical title-abstract pairs from PubMed Central (PMC) articles. Dataset Overview Total examples: 26,573,900 Training set: 26,473,900 examples (99.6%) Validation set: 50,000 examples (0.2%) Test set: 50,000 examples (0.2%) Quality Filtering Applied This dataset has been processed with the following quality filters: Abstract Length… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract.texttext-classification10M<n<100M0 likes418 downloads1y agoHugging Face07labofsahil /hackernews-vector-search-datasetThe Hacker News dataset contains 28.74 million postings and their vector embeddings. The embeddings were generated using SentenceTransformers model all-MiniLM-L6-v2. The dimension of each embedding vector is 384. Created by clickhouse more info: https://clickhouse.com/docs/getting-started/example-datasets/hackernews-vector-search-dataset tabular10M<n<100M2 likes247 downloads10mo agoHugging Face08ajaysri /route_red_yellow_vector_subtasks_pi05 Route Red-Yellow Vector Subtasks for pi0.5 This is a LeRobot v2.1 transformation of DistantSky/route at commit aced1e96f6b8bf98ffaa0754407636e82442b084. The dataset contains 212 real-robot episodes, 135699 frames, three cameras, and 14-dimensional actions at 100 Hz. Conditioning Only observation.images.video_overhead is modified. The left and right videos are byte-identical to the source dataset. The overhead image receives one fixed selected-connector pose glyph… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/route_red_yellow_vector_subtasks_pi05.tabularrobotics100K<n<1M0 likes228 downloads2mo agoHugging Face09oitnews /rss_vectorstext100K<n<1M0 likes221 downloads1y agoHugging Face10vectroid-ai /amazon-reviewsimage100K<n<1M0 likes213 downloads1y agoHugging Face11Kandil7 /shamela-vectors 🇬🇧 English The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim). Statistics: 11,482,164 chunks from 8,589 classical Islamic books 6,236 Quran verses (one verse = one chunk, included in total) 40 categories covering the full breadth of Islamic scholarship Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/shamela-vectors.tabularfeature-extraction10M<n<100M2 likes206 downloads4mo agoHugging Face12ajaysri /pi07_cable_three_vector_v1 Three-holder cable routing with vector goals Real-robot demonstrations for a vector-conditioned low-level policy: place three holders and route a cable through each holder. This is the validated LeRobot v3 dataset prepared for the first Pi0.7 four-camera-goal cable policy. Property Value Source recordings 140 Subtask episodes 840 Frames 335,897 Sampling rate 100 Hz Robot ARX bimanual Recorded state / action 14 / 14 dimensions Video resolution 448 × 448… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/pi07_cable_three_vector_v1.tabular100K<n<1M0 likes206 downloads15d agoHugging Face13OALL /details_Delta-Vector__Odin-9B Dataset Card for Evaluation run of Delta-Vector/Odin-9B Dataset automatically created during the evaluation run of model Delta-Vector/Odin-9B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Delta-Vector__Odin-9B.tabular100K<n<1M0 likes198 downloads2y agoHugging Face14vec-ai /MVRBScreenshotRetrieval MVRB Screenshot Retrieval MTEB v2 multimodal retrieval dataset layout for the seven Screenshot Retrieval subsets from MVRB. image10K<n<100K0 likes174 downloads4mo agoHugging Face15Ma-Vector /MobileGym-ConAct-Trajectories MobileGym-ConAct-Trajectories Dataset Viewer · MobileGym · MemGUI-Agent · Paper Abstract MobileGym-ConAct-Trajectories is a release of successful mobile GUI-agent rollouts collected in the MobileGym simulator. Each trajectory is selected from judge-verified rollouts using a deterministic per-task rule: retain the shortest structurally valid success, then break ties by source run and episode ID. The release preserves screenshots, the rendered prompt supplied to the… See the full description on the dataset page: https://huggingface.co/datasets/Ma-Vector/MobileGym-ConAct-Trajectories.imageimage-text-to-text1K<n<10K0 likes172 downloads2mo agoHugging Face16oitnews /newsdataio_vectorstext10K<n<100K0 likes169 downloads1y agoHugging Face17BinPrey /x86-instruction-test-vectors x86-64 Instruction Test Vectors Ground truth behavior of individual x86-64 instructions, captured by executing every encoding on real hardware and recording the resulting register and flag state. This is measured silicon behavior, not a model and not an emulator, so it also reflects implementation specific results such as the values an instruction leaves in flags that the architecture documents as undefined. How it was generated Each test case is produced by the… See the full description on the dataset page: https://huggingface.co/datasets/BinPrey/x86-instruction-test-vectors.text100M<n<1B1 likes169 downloads2mo agoHugging Face18oitnews /googlenews_vectorstext1K<n<10K0 likes167 downloads1y agoHugging Face19vectroid-ai /amazon-reviews-10Mimage10M<n<100M0 likes162 downloads1y agoHugging Face20vector-institute /HumaniBench HumaniBench: A Human-Centric Benchmark for Large Multimodal Models Evaluation **HumaniBench** is a benchmark for evaluating large multimodal models (LMMs) using real-world, human-centric criteria. It consists of 32,000+ image–question pairs across 7 tasks: ✅ Open/closed VQA 🌍 Multilingual QA 📌 Visual grounding 💬 Empathetic captioning 🧠 Robustness, reasoning, and ethics Each example is annotated with GPT-4o drafts, then verified by experts to ensure quality and… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/HumaniBench.imagevisual-question-answering10K<n<100K4 likes158 downloads18d agoHugging Face21vector-institute /hotpotqatext100K<n<1M0 likes155 downloads1y agoHugging Face22mikronai /VectorEdits VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics NOTE: Currently only test set has generated labels, other sets will have them soon Find the details in our paper: VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics Github repository: JosefKuchar/vector-edits We introduce a large-scale dataset for instruction-guided vector image editing, consisting of over 270,000 pairs of SVG images paired with natural language… See the full description on the dataset page: https://huggingface.co/datasets/mikronai/VectorEdits.text100K<n<1M7 likes153 downloads7mo agoHugging Face23vec-ai /LoCoMo LoCoMo MTEB v2 text retrieval dataset layout for LoCoMo. The candidates configs store the per-query retrieval pool for each memory retrieval subset. text10K<n<100K0 likes132 downloads4mo agoHugging Face24omrisap /nvidia-math-vectorizedtabular10K<n<100K0 likes131 downloads7mo agoHugging Face25derenrich /wiki-image-vectorsVector embeddings of images on Wikipedia using google/siglip2-base-patch16-384 Here's an example of how to use this # /// script # requires-python = ">=3.11" # dependencies = [ # "gradio", # "torch", # "transformers", # "pyarrow", # "numpy", # "huggingface_hub", # ] # /// """ Semantic search over Wikipedia/Commons image embeddings (SigLIP 2). Run locally: VECTORS=/path/to/vectors.parquet uv run app.py Or let it pull from the Hub: uv run app.py """ import… See the full description on the dataset page: https://huggingface.co/datasets/derenrich/wiki-image-vectors.tabular1M<n<10M0 likes129 downloads10d agoHugging Face26EMBO /soda-vec-data-full_pmc_title_abstract_paired SODA-VEC Paired Dataset for Negative Sampling This is a paired version of the SODA-VEC dataset, specifically formatted for negative sampling training with MultipleNegativesRankingLoss. Dataset Overview Total examples: 26,573,900 Format: Paired (anchor-positive) for contrastive learning Source: EMBO/soda-vec-data-full_pmc_title_abstract Purpose: Training sentence transformers with negative sampling Data Format Each example contains: anchor (string): The title… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract_paired.textsentence-similarity10M<n<100M1 likes127 downloads1y agoHugging Face27vectorzhou /AIME_2024_DeepSeek_R1_0528_Temp_1.0_L_16384Responses of deepseek-ai/DeepSeek-R1-0528 for AIME 2024 (original dataset: Maxwell-Jia/AIME_2024). Generation temperature is set to 1.0 and maximum token is set to 16384. textn<1K0 likes124 downloads1y agoHugging Face28kardosdrur /scandi-wiki-vector-store Dataset Card for kardosdrur/scandi-wiki-vector-store This dataset was created using the vicinity library, a lightweight nearest neighbors library with flexible backends. It contains a vector space with 3655450 items. Usage You can load this dataset using the following code: from vicinity import Vicinity vicinity = Vicinity.load_from_hub("kardosdrur/scandi-wiki-vector-store") After loading the dataset, you can use the vicinity.query method to find the nearest neighbors to… See the full description on the dataset page: https://huggingface.co/datasets/kardosdrur/scandi-wiki-vector-store.text1M<n<10M0 likes119 downloads1y agoHugging Face29TonicAI /vectrix-art-e Vectrix ART-E: Synthetic Email Agent Benchmark A fully synthetic email corpus and task dataset for training and evaluating email search agents, built as a drop-in replacement for the Enron corpus used in OpenPipe's ART-E benchmark. Key Result A Qwen3.5-35B-A3B fine-tuned via GRPO on this synthetic dataset beats o3 on real Enron emails (86% vs 85%) — despite never seeing a single real email during training. Dataset Contents The dataset is available in two… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/vectrix-art-e.tabularquestion-answering1K<n<10K1 likes99 downloads6mo agoHugging Face30vec-ai /DUDE DUDE MTEB v2 multimodal retrieval dataset layout for DUDE. image10K<n<100K0 likes89 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.