datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dbpedia-entities-openai-1M1M OpenAI Embeddings -- 1536 dimensions
Created: June 2023.
Text used for Embedding: title (string) + text (string)
Embedding Model: text-embedding-ada-002
First used for the pgvector vs VectorDB (Qdrant) benchmark: https://nirantk.com/writing/pgvector-vs-qdrant/
Citation
@dataset{dbpedia-entities-openai-1M,
doi = {10.57967/hf/6768},
url = {https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M},
author = {{Kumar Shivendu} and {Nirant Kasliwal}},
title =… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M.dbpedia-entities-openai3-text-embedding-3-large-1536-1M1M OpenAI Embeddings: text-embedding-3-large 1536 dimensions
Created: February 2024.
Text used for Embedding: title (string) + text (string)
Embedding Model: OpenAI text-embedding-3-large
This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here
dbpedia-entities-openai3-text-embedding-3-large-3072-1M1M OpenAI Embeddings: text-embedding-3-large 3072 dimensions + ada-002 1536 dimensions — parallel dataset
Created: February 2024.
Text used for Embedding: title (string) + text (string)
Embedding Model: text-embedding-3-large
This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here
ghana-named-entities-tts-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Named Entities TTS — Twi
A Twi-language speech dataset built from descriptions of Ghana named entities
(people, places, organisations, and concepts). Each audio clip is a synthesised
reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.lol-esports-entities
GPTilt: League of Legends Esports Directory
This dataset is part of the GPTilt open-source initiative, aimed at democratizing access to high-quality LoL data for research and analysis, fostering public exploration, and advancing the community's understanding of League of Legends through data science and AI. It provides a clean, canonical reference for the people and organizations of competitive League of Legends.
By using this dataset, users accept full responsibility for any… See the full description on the dataset page: https://huggingface.co/datasets/gptilt/lol-esports-entities.dbpedia-entities-openai3-text-embedding-3-small-1536-100Kdbpedia-entities-openai3-text-embedding-3-large-1536-100Kdbpedia-entities-efficient-splade-100K
DBPedia SPLADE + OpenAI: 100,000 SPLADE Sparse Vectors + OpenAI Embedding
This dataset has both OpenAI and SPLADE vectors for 100,000 DBPedia entries. This adds SPLADE Vectors to KShivendu/dbpedia-entities-openai-1M/
Model id used to make these vectors:
model_id = "naver/efficient-splade-VI-BT-large-doc"
For processing the query, use this:
model_id = "naver/efficient-splade-VI-BT-large-query"
If you'd like to extract the indices and weights/values from the vectors, you can do so… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/dbpedia-entities-efficient-splade-100K.litbank-entities
Dataset Card for "litbank-entities"
More Information needed
dbpedia-entities-openai3-text-embedding-3-large-3072-100Kblab-all-entities-30s-chunksCODE-ACCORD-Entities
CODE-ACCORD: A Corpus of Building Regulatory Data for Rule Generation towards Automatic Compliance Checking
The CODE-ACCORD corpus contains annotated sentences from the building regulations of England and Finland and has been developed as part of the Horizon European project for Automated Compliance Checks for Construction, Renovation or Demolition Works (ACCORD). The corpus is in English, and it consists of both the English Building Regulations and the English translation of the… See the full description on the dataset page: https://huggingface.co/datasets/ACCORD-NLP/CODE-ACCORD-Entities.named-entities-tts-pt-br-audio-qwen3tts
Qwen3-TTS (checkpoint base) — áudio sintetizado de entidades nomeadas
Dataset de áudio gerado sinteticamente a partir do glossário de entidades nomeadas do
projeto de IC "Aprimoramento de modelos de reconhecimento automático de fala em relação
ao reconhecimento de nomes próprios", usando o modelo Qwen3-TTS (Qwen/Qwen3-TTS-12Hz-1.7B-Base) em seu
checkpoint base, sem fine-tuning — etapa de baseline do projeto.
Splits
Split
Nº de exemplos
train
17309… See the full description on the dataset page: https://huggingface.co/datasets/RodrigoLimaRFL/named-entities-tts-pt-br-audio-qwen3tts.ontario-hansard-entities
Ontario Hansard — entity mentions (rule-extracted), v0.2
A mention-level entity layer over
ontario-hansard-official (the speaker-resolved Ontario Hansard corpus,
4,316,254 paragraphs, parliaments 29–44). 658,556 entity mentions
of four types — BILL, ACT, RIDING, MEMBER_REF_RIDING — extracted
deterministically by rule (no model). A companion dataset: it carries no text
of its own beyond verbatim entity surfaces and joins back to the corpus on id.
v0.2 supersedes v0.1 (replaces it… See the full description on the dataset page: https://huggingface.co/datasets/agoulah/ontario-hansard-entities.named-entities-tts-pt-br-audio
YourTTS (checkpoint base) — áudio sintetizado de entidades nomeadas
Dataset de áudio gerado sinteticamente a partir do glossário de entidades nomeadas do
projeto de IC "Aprimoramento de modelos de reconhecimento automático de fala em relação
ao reconhecimento de nomes próprios", usando o modelo YourTTS (tts_models/multilingual/multi-dataset/your_tts) em seu
checkpoint base, sem fine-tuning — etapa de baseline do projeto.
Splits
Split
Nº de exemplos… See the full description on the dataset page: https://huggingface.co/datasets/RodrigoLimaRFL/named-entities-tts-pt-br-audio.dbpedia-entities-openai3-text-embedding-3-small-512-100Kgames-balatro-2024-entities-detection
Project AIRI's Games Datasets - Balatro (2024, game) - Entities detection
This project is part of (and also associate to) the Project AIRI, we aim to build a LLM-driven VTuber like Neuro-sama (subscribe if you didn't!) if you are interested in, please do give it a try on live demo.
Who are we?
We are a group of currently non-funded talented people made up with computer scientists, experts in multi-modal fields, designers, product managers, and popular open source contributors… See the full description on the dataset page: https://huggingface.co/datasets/proj-airi/games-balatro-2024-entities-detection.dbpedia-entities-openai3-text-embedding-3-small-1024-100Kdbpedia-entities-google-palm-gemini-embedding-001-100K
Dataset Card for DBPedia 100K: Gemini Google Embedding Model 001
100K vectors from DBPedia!
Embedding Model: Google's latest Embedding Model 001 -- the successor to the Gecko Models!
Dataset Details
Dataset Description
100K Google Embeddings -- 768 dimensions
Created: December 2023
Text used for Embedding: title (string) + text (string)
Embedding Model: Google's models/embedding-001
Curated by: Nirant Kasliwal
Funded by: Qdrant Gmbh
Language(s) (NLP): English… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/dbpedia-entities-google-palm-gemini-embedding-001-100K.dbpedia-entities-openai3-text-embedding-3-large-1024-100Ken_wikidata_5M_entities
en_wikidata_5M_entities
Hugging Face dataset card for a large, English-only Wikidata slice with optional Wikipedia links and Wikimedia Commons image URLs.
One file, five million entities.
Filename: en_wikidata_5M_entities.jsonl.gz
TL;DR
Format: JSON Lines, gzip-compressed (.jsonl.gz)
Rows: 5,000,000 entities (one JSON object per line)
Language: English labels/descriptions
Fields: qid, label, description, enwiki_title, wikipedia_url, images (list of URLs), has_image… See the full description on the dataset page: https://huggingface.co/datasets/Vijaysr4/en_wikidata_5M_entities.NOO-Verified-Global-Entities
🌐 NOO-Verified-Global-Entities (NOO AI互联网商业实体验证归属)
Official Data Infrastructure by Nooxus.com
[ 🚀 ANGEL ROUND INVESTOR NOTICE / 天使轮国际融资公告 ]
EN: Nooxus-AI is raising its Angel Round to scale nooxus.com — the world's first dedicated B2B Trading & Clearing Network for AI Agents. If your fund recognizes the trillion-dollar potential of building the "Visa / SWIFT network for the Agentic Web," powered by our production-ready 0.8ms RST signaling and Zero-Inbound stealth architecture… See the full description on the dataset page: https://huggingface.co/datasets/Nooxus-AI/NOO-Verified-Global-Entities.sentinel-kr-sensitive-entities-synthetic-v3
Sentinel KR Sensitive Entities (Synthetic) v3
Overview
Sentinel KR Sensitive Entities (Synthetic) v3 is a Korean synthetic (AI-generated) dataset for whitelist-only sensitive-entity detection in DLP / LLM guardrail scenarios.
All sensitive values in this dataset (e.g., phone numbers, emails, IDs, tokens, keys) are artificially generated by AI and do not come from real individuals, real incidents, or collected private datasets.
Any resemblance to real persons or real… See the full description on the dataset page: https://huggingface.co/datasets/BoB14TeamSentinel/sentinel-kr-sensitive-entities-synthetic-v3.MixSub-LLaMA-3.2-Entities-Overlap-GPU-Scoreben-entities
BEN Entities
Full BEN entity extraction results exported from MongoDB as Hub-native
jsonl.gz shards.
Each row contains only document_id and entities. Scores are filtered with
threshold 0.6 and rounded to two decimals.
Configs
pubmed from Mongo collection pubmed_ncbi
pmc from Mongo collection pmc_xml
uspto from Mongo collection patent_uspto
clinical_trial from Mongo collection clinical_trial_gov
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/BioMedBigDataCenter/ben-entities.kmdb_bert_entitiesdata_for_synthesis_with_entities_align_v3
Dataset Card for "data_for_synthesis_with_entities_align_v3"
More Information needed
qkg-primekg-entities-with-cui
Data Card: qkg-primekg-entities-with-cui
Summary
qkg-primekg-entities-with-cui.jsonl is the QKG entity table derived from PrimeKG and enriched with UMLS CUI annotations. It provides the entity inventory used by the QKG runtime for entity lookup and UMLS-
backed synonym matching.
This artifact is intended to be loaded into MongoDB collection:
primeKG.entities
Paper
This artifact is released with the paper:
Yao Wang, Zixu Geng, and Jun Yan. Quantum… See the full description on the dataset page: https://huggingface.co/datasets/HKAI-Sci/qkg-primekg-entities-with-cui.data_for_synthesis_with_entities_align_v5
Dataset Card for "data_for_synthesis_with_entities_align_v5"
More Information needed
dbpedia-entities-mistral-embeddings-100K
