CoolFace
Datasetpublic

rntc/biomed-fr-v3-enriched-softmin-leger

biomed-fr-v3-enriched-softmin-leger This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling. Preprocessing Method Soft-min calculation: Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores Parameter p = -0.7 Weight computation: Ratio preference (5 vs 1): R = 5 Gamma exponent: γ = 1.00 (computed as log(R)/log(5)) Weight formula: w = s^γ Floor: w = max(w, median(w) × 0.05) Resampling: Target… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-leger.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes40downloads
Dataset Card

biomed-fr-v3-enriched-softmin-leger

This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling.

Preprocessing Method

Soft-min calculation:

  • —Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores
  • —Parameter p = -0.7

Weight computation:

  • —Ratio preference (5 vs 1): R = 5
  • —Gamma exponent: γ = 1.00 (computed as log(R)/log(5))
  • —Weight formula: w = s^γ
  • —Floor: w = max(w, median(w) × 0.05)

Resampling:

  • —Target size: Same as original (2941107 samples)
  • —Method: Sampling with replacement according to normalized weights
  • —Seed: 123
  • —Output: Only text column retained, shuffled

Quality Scores Used

The following 4 scores (range 1-5) from the source dataset were used:

  • —educational_score
  • —content_richness
  • —terminology_precision
  • —writing_quality

Samples with missing scores were excluded from resampling.

Dataset Stats

  • —Original dataset size: 2941107
  • —Resampled dataset size: 2941107
  • —Preset: leger

Citation

Original dataset: rntc/biomed-fr-v3-enriched