datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
complex_ner
Elephant Labs Complex PII Dataset for Long Contexts and Advanced Anonymization (with Business and Software-related Entities)
Developed by: Elephant Labs
LinkedIn: Elephant Labs
Dataset Size: 20,0000 synthetic documents
Number of tokens in text: 14,140,795 (Tokenized with tiktoken.encoding_for_model("gpt-3.5-turbo"))
Dataset Summary
Purpose: A synthetically generated dataset for advanced NER tasks, supporting both token classification and LLM fine-tuning (enabling… See the full description on the dataset page: https://huggingface.co/datasets/MorryShah/complex_ner.medical-ner-sft
Medical Named Entity Recognition (NER)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Clinical text → structured JSON with conditions, drugs, dosages, procedures
Why download this
Train clinical NER models to extract structured data from unstructured clinical notes. Output is JSON-formatted for downstream pipeline… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/medical-ner-sft.pharmacy-ner-sft
Pharmacy NER — Drug Entity Extraction
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Clinical/biomedical text → drug name, dosage, frequency, route, indication
Why download this
Automate medication extraction from clinical notes, discharge summaries, or biomedical literature. Powers medication reconciliation and… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/pharmacy-ner-sft.esic-nerDataset sintético para treinamento em tarefa de extração de entidades (NER) para uso em classificação de dados pessoais (PII) em formulários e-SIC.
Estatísticas do train split
Summary
samples: 4473
samples_with_any_entity: 3571 (79.83%)
samples_with_any_pii (excludes ORG_JURIDICA, DOC_EMPRESA): 2244 (50.17%)
entity_records_total: 14510
literal_occurrences_total: 14686
Note: ORG_JURIDICA and DOC_EMPRESA are labels but are treated as non-PII (excluded from PII-only… See the full description on the dataset page: https://huggingface.co/datasets/EliMC/esic-ner.Neru-66K-Bilingual-SFT
Neru-66K-Bilingual-SFT
This dataset is a high-quality, professional 66,000 (66K) row bilingual Supervised Fine-Tuning (SFT) instruction set optimized for training Large Language Models (LLMs) in both Turkish-to-English and English-to-Turkish translation tasks.
Non-Synthetic
Dataset Details
Curated by: ezfiez dev
Language(s) (NLP): Turkish, English
License: CC-BY-4.0 (Permissive license. Free to use for both commercial and personal projects, provided appropriate… See the full description on the dataset page: https://huggingface.co/datasets/ezfiez/Neru-66K-Bilingual-SFT.Mental_Health_Support_ChatBOT_Conversation
Mental Health Support Dataset
Instruction–response pairs for training supportive, non-diagnostic,
safety-aware mental health chatbots.
Fields
instruction: user message
response: Bot reposne
category: intent label
Safety
This dataset includes crisis escalation examples and refusal patterns.
Not a replacement for professional care.
autotrain-data-chinese-nerrmj_covid_ner
