CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01epfml /FineWeb2-embedded FineWeb2-embedded Dataset summary FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research. Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer texts… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.tabulartext-generation1B<n<10B6 likes23k downloads2y agoHugging Face02permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes5.4k downloads2y agoHugging Face03Marcus2112 /refinedweb-embedded_prototypeContaining 768-dimensional embedding vectors derived from the "content" column of the RefinedWeb dataset. The embeddings were generated using the E5-Base-4k model with a context length of 1024, employing scaled dot-product attention. Original dataset: This dataset bases as derivative work on RefinedWeb, an English web dataset created by the TII (Technology Innovation Institute). Attribution is given to the TII as original authors of the RefinedWeb dataset as per Section 4 of RefinedWeb's Open… See the full description on the dataset page: https://huggingface.co/datasets/Marcus2112/refinedweb-embedded_prototype.text10M<n<100M0 likes1.3k downloads10mo agoHugging Face04SalihHub /Wikipedia-TR-2023-Embedded-Dump Wikipedia-TR-2023-Embedded-Dump Türkçe Vikipedi (tr.wikipedia.org) makaleleri, retrieval / RAG kullanım senaryoları için parçalanmış (chunk) ve embedding'lenmiş hâliyle. Her makale bir parent (ana) chunk'a (tüm makale metni, bağlam genişletmek için) ve birden fazla child (alt) chunk'a (her biri kendi embedding'ine sahip küçük pasajlar) bölünmüştür. İçerik Makale 348.751 Embedding'li child chunk 1.308.623 Parent chunk (embeddingsiz) 348.751… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Wikipedia-TR-2023-Embedded-Dump.tabularfeature-extraction1M<n<10M0 likes590 downloads1mo agoHugging Face05PwwSeniorProj /CT-RATE_RAPTOR_DINOV3_Embedded_Validtext0 likes458 downloads9mo agoHugging Face06embedded-language-flows /xsum_train_t5tabular100K<n<1M0 likes400 downloads4mo agoHugging Face07neurovlm /embedded_texttextn<1K0 likes379 downloads4mo agoHugging Face08vibhuiitj /UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6btabular1M<n<10M0 likes366 downloads6mo agoHugging Face09vibhuiitj /UltraData-Math-L2-preview-embedded-with-idstabular10M<n<100M0 likes263 downloads6mo agoHugging Face10vibhuiitj /UltraData-Math-L3-Textbook-Exercise-Synthetic-split-qwen3-0.6b-embeddedtext1M<n<10M1 likes259 downloads6mo agoHugging Face11Technoculture /chatdoctor-embedded Chat Doctor with Embeddings This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped: Add embeddings for input and output columns using BAAI/bge-small-en-v1.5 Details Sample Count 414k Token Count 1.7b Origin https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view Source of raw data ? Processing details paper Embedding Model BAAI/bge-small-en-v1.5 Data Diversity index Example Output GPT-4 Rationale GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.text100K<n<1M3 likes248 downloads3y agoHugging Face12MongoDB /embedded_movies sample_mflix.embedded_movies This data set contains details on movies with genres of Western, Action, or Fantasy. Each document contains a single movie, and information such as its title, release year, and cast. In addition, documents in this collection include a plot_embedding field that contains embeddings created using OpenAI's text-embedding-ada-002 embedding model that you can use with the Atlas Search vector search feature. Overview This dataset offers a… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/embedded_movies.image1K<n<10K18 likes245 downloads2y agoHugging Face13sproos /SlimPajama-6B-embedded Dataset Card for SlimPajama-6B-embedded This is a copy of DKYoon/SlimPajama-6B, together with embeddings generated by thenlper/gte-large. There are 5.49 million examples of text, a representative random sample of SlimPajama-627B. Each text is associated with a 1024-dimensional embedding vector that is meant to represent the semantic content. The vectors were generated by average-pooling (max-pooling dataset to come in the future). This dataset is intended to help with downstream… See the full description on the dataset page: https://huggingface.co/datasets/sproos/SlimPajama-6B-embedded.text1M<n<10M3 likes228 downloads3y agoHugging Face14lmcinnes /20newsgroups_embedded Dataset Card for 20-Newsgroups Embedded This provides a subset of 20-Newsgroup posts, along with sentence embeddings, and a dimension reduced 2D data map. This provides a basic setup for experimentation with various neural topic modelling approaches. Dataset Details Dataset Description This is a dataset containing posts from the classic 20-Newsgroups dataset, along with sentence embeddings, and a dimension reduced 2D data map. Per the source: The… See the full description on the dataset page: https://huggingface.co/datasets/lmcinnes/20newsgroups_embedded.text10K<n<100K0 likes224 downloads2y agoHugging Face15much1na /miriad-embeddedtext10K<n<100K0 likes201 downloads21d agoHugging Face16vibhuiitj /UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6b_newtabular1M<n<10M0 likes152 downloads6mo agoHugging Face17HydraLM /embedded_1 Dataset Card for "embedded_1" More Information needed text1M<n<10M3 likes148 downloads3y agoHugging Face18phionyx /airep-embedded-evaluation-profile AIREP Embedded Evaluation Profile v0.1 This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical specification history lives in the AIREP GitHub repository. Byte identity between this mirror and its source commit is a distribution-integrity property, not independent scientific verification. Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.textn<1K0 likes129 downloads4d agoHugging Face19victorych22 /lamini-embedded-instructions-only Dataset Card for "lamini-embedded-instructions-only" More Information needed text1M<n<10M2 likes112 downloads3y agoHugging Face20HydraLM /corpus_1_embedded_deduplicated Dataset Card for "corpus_1_embedded_deduplicated" More Information needed text1M<n<10M0 likes101 downloads3y agoHugging Face21victorych22 /lamini-embedded Dataset Card for "lamini-embedded" More Information needed text1M<n<10M0 likes100 downloads3y agoHugging Face22Technoculture /synthetic-clinical-notes-embedded Synthetic Clinical Notes This dataset is post-processed version of starmpcc/Asclepius-Synthetic-Clinical-Notes: Turn into Alpaca format (instruction, input, and output) Add embeddings for input and output columns using BAAI/bge-small-en-v1.5 Details Sample Count 158k Token Count 648m Origin https://figshare.com/authors/Zhengyun_Zhao/16480335 Source of raw data PubMed Central (PMC) and MIMIC 3 Processing details original, paper Embedding Model… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/synthetic-clinical-notes-embedded.textquestion-answering100K<n<1M10 likes88 downloads3y agoHugging Face23HydraLM /embedded_datasets_0822 Dataset Card for "combined_embedded_v2" More Information needed text1M<n<10M0 likes86 downloads3y agoHugging Face24tennisb /github-embedded github-embedded A bunch of python code, gotten from Github, along with a json dump of their ast's and a vector embedding of the code using openai's text-embedding-3-small. tabulartext-classification10M<n<100M1 likes86 downloads1y agoHugging Face25podcusdace /donaroma3421.42_embeddedaudio10K<n<100K0 likes82 downloads10mo agoHugging Face26Ailiance-fr /kill-life-embedded-qa Ailiance — Kill-LIFE Embedded Knowledge Base 🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/kill-life-embedded-qa. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025). Knowledge-base Q&A spécifique au projet Kill_LIFE (compagnon vocal embarqué ESP32-S3 + Mascarade) : composants matériels, schémas KiCad du board minimal, simulations SPICE de l'alimentation/I2C/I2S/audio, et architecture du firmware C++ (pipeline… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/kill-life-embedded-qa.texttext-generationn<1K0 likes74 downloads4mo agoHugging Face27Alignment-Lab-AI /embeddedflantext100K<n<1M0 likes71 downloads2y agoHugging Face28DopeorNope /orca_embeddedtext100K<n<1M0 likes67 downloads2y agoHugging Face29eniomecaj /embedded-systems-qa Embedded Systems Engineering Q&A — Instruction Dataset A hand-authored instruction-tuning dataset of technical question/answer pairs for embedded systems engineering, formatted for supervised fine-tuning of Mistral 7B (Alpaca-style instruction / input / output schema). At a glance Entries 302 Format JSONL, one JSON object per line Schema {"instruction": <question>, "input": "", "output": <answer>} Language English Avg. answer length ~590… See the full description on the dataset page: https://huggingface.co/datasets/eniomecaj/embedded-systems-qa.textquestion-answeringn<1K0 likes66 downloads2mo agoHugging Face30Sam04 /BlueLionEcho_tran_embedded_filteredaudio10K<n<100K0 likes65 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.