datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vc-deal-flow-signal-glossary
VC Deal Flow Signal Glossary
The controlled vocabulary used across the VC Deal Flow Signal site —
84 definitions covering code-side sourcing, engineering acceleration
metrics, discoverability surfaces (programmatic SEO, AEO, GEO, AIO),
agent infrastructure (MCP, A2A, x402), academic citation infrastructure,
and venture vocabulary including the SaaS efficiency quintet (burn
multiple, magic number, CAC payback, LTV, quick ratio).
Maintained as a single source of truth and refreshed… See the full description on the dataset page: https://huggingface.co/datasets/the-data-nerd/vc-deal-flow-signal-glossary.ner_court_decisions
Basic Information
This dataset is converted from fewshot-goes-multilingual/cs_czech-court-decisions-ner using script convert_ner_court_decisions.py.
For longer texts (>200 ws tokens), the script samples text around the selected entity. It always follows form "<initial 20 ws tokens>, ..., <sampled window>".
Then it extracts category name for the entity, all occurences of such entity in the text, and creates simple json representation. For example:
{
"label": "Reference na rozhodnutí… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/ner_court_decisions.my_uni_nerFunder-NER
Dataset Card for Dataset Named Entity Recognition of funders of scientific research
Dataset Summary
Training/test set for automatically identifying funder entities mentioned in scientific papers. This data set is generated from Open Access documents hosted at https://econstor.eu and manually curated/labeled.
Supported Tasks and Leaderboards
The dataset is for training and testing the automatic recognition of funders as they are acknowledged in scientific… See the full description on the dataset page: https://huggingface.co/datasets/ZBWatHF/Funder-NER.Zarma_NER
ZarmaNER-600 Dataset
Dataset Description
ZarmaNER-600 is a gold-standard dataset for Named Entity Recognition (NER) in Zarma. This dataset contains 600 manually annotated sentences, making it the first publicly available NER corpus for Zarma. It was created to support research in low-resource NLP, particularly for sequence tagging tasks, as part of the Rule-to-Tag (R2T) framework introduced in our paper, "R2T: A Case Study in Principled Learning for Low-Resource POS… See the full description on the dataset page: https://huggingface.co/datasets/27Group/Zarma_NER.
