openchs
Datasets
All datasets matching “openchs”synthetic-helpline-ner-v1
Dataset Card for Child Helpline NER Dataset
Dataset Details
Dataset Description
This dataset contains synthetic English conversations from Tanzanian child helpline services for training Named Entity Recognition (NER) models. The dataset extracts critical information from helpline conversations including caller names, ages, locations, incident types, and other key entities essential for case management and child protection services.
Key features:
Character-span… See the full description on the dataset page: https://huggingface.co/datasets/openchs/synthetic-helpline-ner-v1.sum-synthetic-data-v1
Child Protection Helpline Case Summarization Dataset
Overview
This dataset provides 1,000 synthetic child protection helpline call transcripts paired with concise, structured summaries for training sequence-to-sequence summarization models. It simulates real-world calls reporting various forms of abuse and exploitation in East African contexts while ensuring no real cases or personal data are included
The data was created to fine-tune FLAN-T5 and related models for rapid… See the full description on the dataset page: https://huggingface.co/datasets/openchs/sum-synthetic-data-v1.synthetic_helpine_classification_v1
Synthetic Helpline Call Classification Dataset
Synthetic conversation data for training and evaluating child helpline case classification and quality assessment models in East African contexts.
Dataset Details
Dataset Description
This is a fully synthetic dataset of helpline conversations generated for training multitask classification models for the OpenCHS AI Pipeline. The dataset simulates realistic child helpline scenarios across various case categories… See the full description on the dataset page: https://huggingface.co/datasets/openchs/synthetic_helpine_classification_v1.synthetic_helpline_qa_scoring_v1
Dataset Card for OpenCHS Child Helpline QA Scoring Dataset
Dataset Details
Dataset Description
This dataset contains simulated, annotated child helpline conversation transcripts for automated quality assurance (QA) scoring. Each conversation is evaluated across six key performance dimensions using binary labels to assess counselor effectiveness and call quality. The dataset is designed for training multi-head DistilBERT classification models that can… See the full description on the dataset page: https://huggingface.co/datasets/openchs/synthetic_helpline_qa_scoring_v1.synthetic-helpline-sw-en-translation-v1
Dataset Card for the Synthetic Swahili-English Helpline Translation Dataset
Dataset Details
Dataset Description
This dataset contains synthetic parallel Swahili-English translations from Tanzanian child helpline conversations, designed for training and evaluating neural machine translation (NMT) models. The dataset addresses the critical need for high-quality translation systems in child protection services across East Africa, where multilingual support is… See the full description on the dataset page: https://huggingface.co/datasets/openchs/synthetic-helpline-sw-en-translation-v1.
