CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01epfml /FineWeb2-embedded FineWeb2-embedded Dataset summary FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research. Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer texts… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.tabulartext-generation1B<n<10B6 likes26k downloads2y agoHugging Face02permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes5.5k downloads2y agoHugging Face03embedded-language-flows /openwebtext-t51M<n<10M2 likes2.4k downloads4mo agoHugging Face04Marcus2112 /refinedweb-embedded_prototypeContaining 768-dimensional embedding vectors derived from the "content" column of the RefinedWeb dataset. The embeddings were generated using the E5-Base-4k model with a context length of 1024, employing scaled dot-product attention. Original dataset: This dataset bases as derivative work on RefinedWeb, an English web dataset created by the TII (Technology Innovation Institute). Attribution is given to the TII as original authors of the RefinedWeb dataset as per Section 4 of RefinedWeb's Open… See the full description on the dataset page: https://huggingface.co/datasets/Marcus2112/refinedweb-embedded_prototype.text10M<n<100M0 likes1.3k downloads10mo agoHugging Face05embedded-language-flows /xsum_validation_t50 likes923 downloads4mo agoHugging Face06SalihHub /Wikipedia-TR-2023-Embedded-Dump Wikipedia-TR-2023-Embedded-Dump Türkçe Vikipedi (tr.wikipedia.org) makaleleri, retrieval / RAG kullanım senaryoları için parçalanmış (chunk) ve embedding'lenmiş hâliyle. Her makale bir parent (ana) chunk'a (tüm makale metni, bağlam genişletmek için) ve birden fazla child (alt) chunk'a (her biri kendi embedding'ine sahip küçük pasajlar) bölünmüştür. İçerik Makale 348.751 Embedding'li child chunk 1.308.623 Parent chunk (embeddingsiz) 348.751… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Wikipedia-TR-2023-Embedded-Dump.tabularfeature-extraction1M<n<10M0 likes666 downloads1mo agoHugging Face07heights1976 /embedded-topology Exp 3B: Embedded topology in trained recurrent operators A 14,553-configuration computational sweep of recurrent network architectures trained on dynamical systems. Each configuration's hidden-state activations were measured for topological fidelity to the driving system using persistent homology and Gauss linking integrals. Six post-hoc analyses on saved checkpoints probe the operator properties the embedding theorems describe abstractly. This dataset is the empirical companion… See the full description on the dataset page: https://huggingface.co/datasets/heights1976/embedded-topology.10K<n<100K0 likes508 downloads4mo agoHugging Face08PwwSeniorProj /CT-RATE_RAPTOR_DINOV3_Embedded_Validtext0 likes459 downloads9mo agoHugging Face09embedded-language-flows /wmt14_de-en_validation_t50 likes453 downloads4mo agoHugging Face10neurovlm /embedded_texttextn<1K0 likes393 downloads4mo agoHugging Face11embedded-language-flows /xsum_train_t5tabular100K<n<1M0 likes386 downloads4mo agoHugging Face12vibhuiitj /UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6btabular1M<n<10M0 likes366 downloads6mo agoHugging Face13open-llm-leaderboard-old /details_EmbeddedLLM__Mistral-7B-Merge-14-v0.2 Dataset Card for Evaluation run of EmbeddedLLM/Mistral-7B-Merge-14-v0.2 Dataset automatically created during the evaluation run of model EmbeddedLLM/Mistral-7B-Merge-14-v0.2 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EmbeddedLLM__Mistral-7B-Merge-14-v0.2.0 likes332 downloads3y agoHugging Face14scaleinvariant /paired-open-images-embedded-pe-core-g14-448 Paired Open Images with PE-Core-G14-448 Embeddings This dataset contains pairs of images from Open Images along with their embeddings computed using Meta's Perception Encoder (PE-Core-G14-448). Each row contains two images (as JPEG bytes), their metadata, and their corresponding 1280-dimensional embeddings. Data Layout Column Description image1_jpeg JPEG bytes for the first image image1_metadata Metadata for the first image image2_jpeg JPEG bytes for… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-open-images-embedded-pe-core-g14-448.image-feature-extraction100K<n<1M0 likes310 downloads7mo agoHugging Face15MongoDB /embedded_movies sample_mflix.embedded_movies This data set contains details on movies with genres of Western, Action, or Fantasy. Each document contains a single movie, and information such as its title, release year, and cast. In addition, documents in this collection include a plot_embedding field that contains embeddings created using OpenAI's text-embedding-ada-002 embedding model that you can use with the Atlas Search vector search feature. Overview This dataset offers a… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/embedded_movies.image1K<n<10K18 likes263 downloads2y agoHugging Face16vibhuiitj /UltraData-Math-L2-preview-embedded-with-idstabular10M<n<100M0 likes263 downloads6mo agoHugging Face17vibhuiitj /UltraData-Math-L3-Textbook-Exercise-Synthetic-split-qwen3-0.6b-embeddedtext1M<n<10M1 likes259 downloads6mo agoHugging Face18Technoculture /chatdoctor-embedded Chat Doctor with Embeddings This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped: Add embeddings for input and output columns using BAAI/bge-small-en-v1.5 Details Sample Count 414k Token Count 1.7b Origin https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view Source of raw data ? Processing details paper Embedding Model BAAI/bge-small-en-v1.5 Data Diversity index Example Output GPT-4 Rationale GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.text100K<n<1M3 likes236 downloads3y agoHugging Face19Aquiles-ai /LLaVA-CC3M-Pretrain-595K-Embedded Dataset derived from liuhaotian/LLaVA-CC3M-Pretrain-595K Dataset details Dataset type: LLaVA Visual Instruct CC3M Pretrain 595K is a subset of CC-3M dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning. We aim to build large multimodal towards GPT-4 vision/language capability. imageimage-to-text2 likes232 downloads1mo agoHugging Face20lmcinnes /20newsgroups_embedded Dataset Card for 20-Newsgroups Embedded This provides a subset of 20-Newsgroup posts, along with sentence embeddings, and a dimension reduced 2D data map. This provides a basic setup for experimentation with various neural topic modelling approaches. Dataset Details Dataset Description This is a dataset containing posts from the classic 20-Newsgroups dataset, along with sentence embeddings, and a dimension reduced 2D data map. Per the source: The… See the full description on the dataset page: https://huggingface.co/datasets/lmcinnes/20newsgroups_embedded.text10K<n<100K0 likes227 downloads2y agoHugging Face21sproos /SlimPajama-6B-embedded Dataset Card for SlimPajama-6B-embedded This is a copy of DKYoon/SlimPajama-6B, together with embeddings generated by thenlper/gte-large. There are 5.49 million examples of text, a representative random sample of SlimPajama-627B. Each text is associated with a 1024-dimensional embedding vector that is meant to represent the semantic content. The vectors were generated by average-pooling (max-pooling dataset to come in the future). This dataset is intended to help with downstream… See the full description on the dataset page: https://huggingface.co/datasets/sproos/SlimPajama-6B-embedded.text1M<n<10M3 likes226 downloads3y agoHugging Face22embedded-language-flows /wmt14_de-en_train_t5tabular1M<n<10M0 likes211 downloads4mo agoHugging Face23much1na /miriad-embeddedtext10K<n<100K0 likes201 downloads20d agoHugging Face24forlinx-embedded /forlinx-downloads Forlinx Embedded: AI Models & Multimedia Resource Hub Welcome to the official Forlinx Embedded resource repository. This hub serves as a centralized ecosystem providing high-performance Edge AI models, Hardware Development Kits, and Creative Marketing Assets specifically tailored for our SoM (System on Module) and SBC (Single Board Computer) platforms. Resource Categories To accelerate your deployment and content creation, our resources are organized into three… See the full description on the dataset page: https://huggingface.co/datasets/forlinx-embedded/forlinx-downloads.0 likes174 downloads1mo agoHugging Face25vibhuiitj /UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6b_newtabular1M<n<10M0 likes152 downloads6mo agoHugging Face26HydraLM /embedded_1 Dataset Card for "embedded_1" More Information needed text1M<n<10M3 likes148 downloads3y agoHugging Face27open-llm-leaderboard-old /details_EmbeddedLLM__Mistral-7B-Merge-14-v0 Dataset Card for Evaluation run of EmbeddedLLM/Mistral-7B-Merge-14-v0 Dataset automatically created during the evaluation run of model EmbeddedLLM/Mistral-7B-Merge-14-v0 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EmbeddedLLM__Mistral-7B-Merge-14-v0.0 likes139 downloads3y agoHugging Face28phionyx /airep-embedded-evaluation-profile AIREP Embedded Evaluation Profile v0.1 This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical specification history lives in the AIREP GitHub repository. Byte identity between this mirror and its source commit is a distribution-integrity property, not independent scientific verification. Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.textn<1K0 likes123 downloads4d agoHugging Face29TransferGraph /embedded_dataset_domain_similarity_EleutherAI_gpt-neo-125m0 likes114 downloads3y agoHugging Face30victorych22 /lamini-embedded-instructions-only Dataset Card for "lamini-embedded-instructions-only" More Information needed text1M<n<10M2 likes108 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.