CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KShivendu /dbpedia-entities-openai-1M1M OpenAI Embeddings -- 1536 dimensions Created: June 2023. Text used for Embedding: title (string) + text (string) Embedding Model: text-embedding-ada-002 First used for the pgvector vs VectorDB (Qdrant) benchmark: https://nirantk.com/writing/pgvector-vs-qdrant/ Citation @dataset{dbpedia-entities-openai-1M, doi = {10.57967/hf/6768}, url = {https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M}, author = {{Kumar Shivendu} and {Nirant Kasliwal}}, title =… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M.textfeature-extraction1M<n<10M26 likes3.9k downloads11mo agoHugging Face02Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1536-1M1M OpenAI Embeddings: text-embedding-3-large 1536 dimensions Created: February 2024. Text used for Embedding: title (string) + text (string) Embedding Model: OpenAI text-embedding-3-large This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here textfeature-extraction1M<n<10M14 likes1.3k downloads3y agoHugging Face03Qdrant /dbpedia-entities-openai3-text-embedding-3-large-3072-1M1M OpenAI Embeddings: text-embedding-3-large 3072 dimensions + ada-002 1536 dimensions — parallel dataset Created: February 2024. Text used for Embedding: title (string) + text (string) Embedding Model: text-embedding-3-large This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here textfeature-extraction1M<n<10M27 likes1k downloads3y agoHugging Face04ghanaopenai /ghana-named-entities-tts-twi This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana Named Entities TTS — Twi A Twi-language speech dataset built from descriptions of Ghana named entities (people, places, organisations, and concepts). Each audio clip is a synthesised reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.audiotext-to-speech1K<n<10K0 likes607 downloads3mo agoHugging Face05gptilt /lol-esports-entities GPTilt: League of Legends Esports Directory This dataset is part of the GPTilt open-source initiative, aimed at democratizing access to high-quality LoL data for research and analysis, fostering public exploration, and advancing the community's understanding of League of Legends through data science and AI. It provides a clean, canonical reference for the people and organizations of competitive League of Legends. By using this dataset, users accept full responsibility for any… See the full description on the dataset page: https://huggingface.co/datasets/gptilt/lol-esports-entities.text100K<n<1M1 likes267 downloads10d agoHugging Face06Qdrant /dbpedia-entities-openai3-text-embedding-3-small-1536-100Ktext100K<n<1M7 likes206 downloads3y agoHugging Face07Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1536-100Ktext100K<n<1M2 likes205 downloads3y agoHugging Face08nirantk /dbpedia-entities-efficient-splade-100K DBPedia SPLADE + OpenAI: 100,000 SPLADE Sparse Vectors + OpenAI Embedding This dataset has both OpenAI and SPLADE vectors for 100,000 DBPedia entries. This adds SPLADE Vectors to KShivendu/dbpedia-entities-openai-1M/ Model id used to make these vectors: model_id = "naver/efficient-splade-VI-BT-large-doc" For processing the query, use this: model_id = "naver/efficient-splade-VI-BT-large-query" If you'd like to extract the indices and weights/values from the vectors, you can do so… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/dbpedia-entities-efficient-splade-100K.textfeature-extraction100K<n<1M3 likes184 downloads3y agoHugging Face09garutyunov /litbank-entities Dataset Card for "litbank-entities" More Information needed textn<1K1 likes167 downloads4y agoHugging Face10Qdrant /dbpedia-entities-openai3-text-embedding-3-large-3072-100Ktext100K<n<1M2 likes152 downloads3y agoHugging Face11DidulaThavishaPro /blab-all-entities-30s-chunksaudio1K<n<10K0 likes102 downloads1y agoHugging Face12ACCORD-NLP /CODE-ACCORD-Entities CODE-ACCORD: A Corpus of Building Regulatory Data for Rule Generation towards Automatic Compliance Checking The CODE-ACCORD corpus contains annotated sentences from the building regulations of England and Finland and has been developed as part of the Horizon European project for Automated Compliance Checks for Construction, Renovation or Demolition Works (ACCORD). The corpus is in English, and it consists of both the English Building Regulations and the English translation of the… See the full description on the dataset page: https://huggingface.co/datasets/ACCORD-NLP/CODE-ACCORD-Entities.textn<1K0 likes97 downloads2y agoHugging Face13RodrigoLimaRFL /named-entities-tts-pt-br-audio-qwen3tts Qwen3-TTS (checkpoint base) — áudio sintetizado de entidades nomeadas Dataset de áudio gerado sinteticamente a partir do glossário de entidades nomeadas do projeto de IC "Aprimoramento de modelos de reconhecimento automático de fala em relação ao reconhecimento de nomes próprios", usando o modelo Qwen3-TTS (Qwen/Qwen3-TTS-12Hz-1.7B-Base) em seu checkpoint base, sem fine-tuning — etapa de baseline do projeto. Splits Split Nº de exemplos train 17309… See the full description on the dataset page: https://huggingface.co/datasets/RodrigoLimaRFL/named-entities-tts-pt-br-audio-qwen3tts.audio10K<n<100K0 likes87 downloads12d agoHugging Face14agoulah /ontario-hansard-entities Ontario Hansard — entity mentions (rule-extracted), v0.2 A mention-level entity layer over ontario-hansard-official (the speaker-resolved Ontario Hansard corpus, 4,316,254 paragraphs, parliaments 29–44). 658,556 entity mentions of four types — BILL, ACT, RIDING, MEMBER_REF_RIDING — extracted deterministically by rule (no model). A companion dataset: it carries no text of its own beyond verbatim entity surfaces and joins back to the corpus on id. v0.2 supersedes v0.1 (replaces it… See the full description on the dataset page: https://huggingface.co/datasets/agoulah/ontario-hansard-entities.tabulartoken-classification100K<n<1M0 likes85 downloads3mo agoHugging Face15RodrigoLimaRFL /named-entities-tts-pt-br-audio YourTTS (checkpoint base) — áudio sintetizado de entidades nomeadas Dataset de áudio gerado sinteticamente a partir do glossário de entidades nomeadas do projeto de IC "Aprimoramento de modelos de reconhecimento automático de fala em relação ao reconhecimento de nomes próprios", usando o modelo YourTTS (tts_models/multilingual/multi-dataset/your_tts) em seu checkpoint base, sem fine-tuning — etapa de baseline do projeto. Splits Split Nº de exemplos… See the full description on the dataset page: https://huggingface.co/datasets/RodrigoLimaRFL/named-entities-tts-pt-br-audio.audio10K<n<100K0 likes75 downloads13d agoHugging Face16Qdrant /dbpedia-entities-openai3-text-embedding-3-small-512-100Ktext100K<n<1M4 likes70 downloads3y agoHugging Face17proj-airi /games-balatro-2024-entities-detection Project AIRI's Games Datasets - Balatro (2024, game) - Entities detection This project is part of (and also associate to) the Project AIRI, we aim to build a LLM-driven VTuber like Neuro-sama (subscribe if you didn't!) if you are interested in, please do give it a try on live demo. Who are we? We are a group of currently non-funded talented people made up with computer scientists, experts in multi-modal fields, designers, product managers, and popular open source contributors… See the full description on the dataset page: https://huggingface.co/datasets/proj-airi/games-balatro-2024-entities-detection.imageobject-detectionn<1K0 likes64 downloads1y agoHugging Face18Qdrant /dbpedia-entities-openai3-text-embedding-3-small-1024-100Ktext100K<n<1M1 likes61 downloads3y agoHugging Face19nirantk /dbpedia-entities-google-palm-gemini-embedding-001-100K Dataset Card for DBPedia 100K: Gemini Google Embedding Model 001 100K vectors from DBPedia! Embedding Model: Google's latest Embedding Model 001 -- the successor to the Gecko Models! Dataset Details Dataset Description 100K Google Embeddings -- 768 dimensions Created: December 2023 Text used for Embedding: title (string) + text (string) Embedding Model: Google's models/embedding-001 Curated by: Nirant Kasliwal Funded by: Qdrant Gmbh Language(s) (NLP): English… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/dbpedia-entities-google-palm-gemini-embedding-001-100K.textfeature-extraction100K<n<1M3 likes55 downloads3y agoHugging Face20Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1024-100Ktext100K<n<1M2 likes55 downloads3y agoHugging Face21Vijaysr4 /en_wikidata_5M_entities en_wikidata_5M_entities Hugging Face dataset card for a large, English-only Wikidata slice with optional Wikipedia links and Wikimedia Commons image URLs. One file, five million entities. Filename: en_wikidata_5M_entities.jsonl.gz TL;DR Format: JSON Lines, gzip-compressed (.jsonl.gz) Rows: 5,000,000 entities (one JSON object per line) Language: English labels/descriptions Fields: qid, label, description, enwiki_title, wikipedia_url, images (list of URLs), has_image… See the full description on the dataset page: https://huggingface.co/datasets/Vijaysr4/en_wikidata_5M_entities.textother1M<n<10M2 likes51 downloads1y agoHugging Face22Nooxus-AI /NOO-Verified-Global-Entities 🌐 NOO-Verified-Global-Entities (NOO AI互联网商业实体验证归属) Official Data Infrastructure by Nooxus.com [ 🚀 ANGEL ROUND INVESTOR NOTICE / 天使轮国际融资公告 ] EN: Nooxus-AI is raising its Angel Round to scale nooxus.com — the world's first dedicated B2B Trading & Clearing Network for AI Agents. If your fund recognizes the trillion-dollar potential of building the "Visa / SWIFT network for the Agentic Web," powered by our production-ready 0.8ms RST signaling and Zero-Inbound stealth architecture… See the full description on the dataset page: https://huggingface.co/datasets/Nooxus-AI/NOO-Verified-Global-Entities.texttext-generation10K<n<100K1 likes51 downloads4mo agoHugging Face23BoB14TeamSentinel /sentinel-kr-sensitive-entities-synthetic-v3 Sentinel KR Sensitive Entities (Synthetic) v3 Overview Sentinel KR Sensitive Entities (Synthetic) v3 is a Korean synthetic (AI-generated) dataset for whitelist-only sensitive-entity detection in DLP / LLM guardrail scenarios. All sensitive values in this dataset (e.g., phone numbers, emails, IDs, tokens, keys) are artificially generated by AI and do not come from real individuals, real incidents, or collected private datasets. Any resemblance to real persons or real… See the full description on the dataset page: https://huggingface.co/datasets/BoB14TeamSentinel/sentinel-kr-sensitive-entities-synthetic-v3.texttoken-classification10K<n<100K0 likes44 downloads9mo agoHugging Face24AdityaMayukhSom /MixSub-LLaMA-3.2-Entities-Overlap-GPU-Scoretabular1K<n<10K0 likes43 downloads1y agoHugging Face25BioMedBigDataCenter /ben-entities BEN Entities Full BEN entity extraction results exported from MongoDB as Hub-native jsonl.gz shards. Each row contains only document_id and entities. Scores are filtered with threshold 0.6 and rounded to two decimals. Configs pubmed from Mongo collection pubmed_ncbi pmc from Mongo collection pmc_xml uspto from Mongo collection patent_uspto clinical_trial from Mongo collection clinical_trial_gov Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/BioMedBigDataCenter/ben-entities.text10M<n<100M0 likes41 downloads5mo agoHugging Face26boapps /kmdb_bert_entitiestext10K<n<100K0 likes34 downloads2y agoHugging Face27thanhduycao /data_for_synthesis_with_entities_align_v3 Dataset Card for "data_for_synthesis_with_entities_align_v3" More Information needed text1K<n<10K0 likes32 downloads3y agoHugging Face28HKAI-Sci /qkg-primekg-entities-with-cui Data Card: qkg-primekg-entities-with-cui Summary qkg-primekg-entities-with-cui.jsonl is the QKG entity table derived from PrimeKG and enriched with UMLS CUI annotations. It provides the entity inventory used by the QKG runtime for entity lookup and UMLS- backed synonym matching. This artifact is intended to be loaded into MongoDB collection: primeKG.entities Paper This artifact is released with the paper: Yao Wang, Zixu Geng, and Jun Yan. Quantum… See the full description on the dataset page: https://huggingface.co/datasets/HKAI-Sci/qkg-primekg-entities-with-cui.tabular100K<n<1M0 likes32 downloads5mo agoHugging Face29thanhduycao /data_for_synthesis_with_entities_align_v5 Dataset Card for "data_for_synthesis_with_entities_align_v5" More Information needed text1K<n<10K0 likes29 downloads3y agoHugging Face30nirantk /dbpedia-entities-mistral-embeddings-100Ktext100K<n<1M0 likes28 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.