datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3.6-27b-h100-bf16-benchmark
Two GPUs do not mean twice the users
This study answers a serving decision, not a hardware trivia question: when a
27B model already fits on one H100, should a second GPU shard the model or run a
second independent replica?
The answer
Concurrency is a load-generator setting, not a user count and not a promise.
Capacity is the highest real arrival rate that satisfies a declared service
objective. That is why this study measures both a saturated concurrency curve… See the full description on the dataset page: https://huggingface.co/datasets/cogeanu-marius/qwen3.6-27b-h100-bf16-benchmark.H100-MonEspaceSante-CPT-corpus
🔧 Code & reproduction complète (scripts, RUNBOOK, reproduce.sh, conversations) : https://github.com/AlexandreFenyo/MonEspaceSante-H100-reproduction
Mon espace santé — Corpus synthétique pour CPT (EntiGraph, FR)
Corpus synthétique en français pour le continued pre-training (CPT) d'un LLM, destiné à
injecter dans les poids les connaissances de la FAQ du service public « Mon espace santé ».
Provenance (1-hop, anti model-collapse)
Généré uniquement à partir des 88… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/H100-MonEspaceSante-CPT-corpus.H100-MonEspaceSante-SFT
🔧 Code & reproduction complète (scripts, RUNBOOK, reproduce.sh, conversations) : https://github.com/AlexandreFenyo/MonEspaceSante-H100-reproduction
Mon espace santé — Données SFT (format Q/R, FR)
Données de supervised fine-tuning servant à restaurer le format question/réponse (et le refus
hors-périmètre) après le CPT — sans injecter de faits nouveaux (Gekhman et al., arXiv:2405.05904 :
la connaissance s'injecte en CPT, le SFT ne sert qu'au format/instruction-following).… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/H100-MonEspaceSante-SFT.
