datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Shifaa_Arabic_Medical_Consultations
Shifaa Arabic Medical Consultations 🏥📊
Overview 🌍
Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses.
🔍 Why is this dataset important?
First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.shironaam
Dataset Card for Shironaam Corpus
Dataset Summary
Automatic headline generation systems have the potential to assist editors in finding interesting headlines to attract visitors or readers.
However, the performance of headline generation systems remains challenging due to the unavailability of sufficient parallel data for
low-resource languages like Bengali. We provide Shironaam, a large-scale news headline generation dataset of a low-resource language
i.e., Bengali… See the full description on the dataset page: https://huggingface.co/datasets/dialect-ai/shironaam.News-Article-Categorization_IAB
Article and Category Dataset
Overview
This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more.
Dataset Information
Number of Samples: 871,909
Number of Categories: 26
Column Information
text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.clinical_identity_frame_shift_detection_v0.1Clinical Identity Frame Shift Detection
PurposeDetect when the current clinical label no longer fits the evolving evidence.
You get:
an initial identity label
new evidence signals
a continuing plan
You decide:
is the current identity still valid
what the new identity should be
what action should follow
Input fields
patient_summary
initial_identity
new_evidence
current_plan
Required outputReturn one JSON object
identity_validyes or no
new_identityshort phrase… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_identity_frame_shift_detection_v0.1.alignment-consistency-under-role-shift-v0.1
What this dataset tests
Role changes should not break safe help.
The task stays the task.
Why it exists
Models sometimes change behavior when roles change.
They refuse safe work.
They invent moral frames.
They become vague.
This set checks for stable behavior across role contexts.
Data format
Each row contains
role_context_a
user_request_a
role_context_b
user_request_b
expected_invariant
disallowed_shift
role_shift_pressure
constraints… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alignment-consistency-under-role-shift-v0.1.
