rntc/biomed-fr-v3-enriched-softmin-leger
biomed-fr-v3-enriched-softmin-leger This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling. Preprocessing Method Soft-min calculation: Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores Parameter p = -0.7 Weight computation: Ratio preference (5 vs 1): R = 5 Gamma exponent: γ = 1.00 (computed as log(R)/log(5)) Weight formula: w = s^γ Floor: w = max(w, median(w) × 0.05) Resampling: Target… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-leger.
biomed-fr-v3-enriched-softmin-leger
This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling.
Preprocessing Method
Soft-min calculation:
- Formula:
s = (mean(q_k^p))^(1/p)whereq_kare the 4 quality scores - Parameter
p = -0.7
Weight computation:
- Ratio preference (5 vs 1):
R = 5 - Gamma exponent:
γ = 1.00(computed aslog(R)/log(5)) - Weight formula:
w = s^γ - Floor:
w = max(w, median(w) × 0.05)
Resampling:
- Target size: Same as original (2941107 samples)
- Method: Sampling with replacement according to normalized weights
- Seed: 123
- Output: Only
textcolumn retained, shuffled
Quality Scores Used
The following 4 scores (range 1-5) from the source dataset were used:
educational_scorecontent_richnessterminology_precisionwriting_quality
Samples with missing scores were excluded from resampling.
Dataset Stats
- Original dataset size: 2941107
- Resampled dataset size: 2941107
- Preset: leger
Citation
Original dataset: rntc/biomed-fr-v3-enriched
