datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chsa-triage-medical-bilingual
CHSA Triage — Corpus médical bilingue (SFT + DPO)
Corpus destiné au post-training d'un agent d'aide au triage médical (POC, Centre
Hospitalier Saint-Aurélien). Bilingue français / anglais, anonymisé (RGPD) et
versionné (empreintes SHA-256 dans manifest.json).
⚠️ Usage : aide à la décision destinée à du personnel soignant. Ne pose pas de
diagnostic et ne remplace pas un professionnel de santé.
Contenu
Fichier
Description
Format
sft_train.jsonl /… See the full description on the dataset page: https://huggingface.co/datasets/DagueGG/chsa-triage-medical-bilingual.dagger
DAGGER Training Dataset
Dataset Description
Training data for DAGGER (Distractor-Aware Graph Generation for Executable Reasoning) models. This dataset contains Bangla mathematical word problems paired with computational graph solutions, formatted for both SFT and GRPO training pipelines.
Highlights
3,000 training examples with verified computational graphs
Two training configs: SFT (with validation) and GRPO formats
GPT-4.1… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/dagger.cmmluCMMLU is a comprehensive Chinese assessment suite specifically designed to evaluate the advanced knowledge and reasoning abilities of LLMs within the Chinese language and cultural context.
