MadlabOSS/LFM2-2.6B-SDG-GGUF
Madlab Synthetic Data Generator
π§ Overview
The Madlab SDG 2.6B is part of the MadlabOSS Synthetic Data Generator family β a suite of small, efficient synthetic data generators designed for ruleβconsistent, semantically coherent variation. This model was trained on a closed-source dataset created through a multi-stage synthetic data generation process using a modified Madlab training pipeline.
π Intended Use
This model is optimized for:
- Madlab synthetic data generation
It is not intended as a general-purpose chatbot.
π§© Model Details
Base Model: LFM2-2.6B Parameter Count: 2.6 Billion Training Type: Supervised fine-tuning Sequence Length: 1024 Precision: FP16 Framework: PyTorch / Transformers
π¦ Training Data
The model was trained on:
- 1444 compressed and encoded dataset pairs
- High variation in output
- Preservation of semantic meaning
- Data entirely generated with Madlab
ποΈ Training Procedure
Hyperparameters
- Epochs: 6
- Batch size: 48
- Learning rate: cosine schedule, peak ~4e-5
- Optimizer: AdamW
- Gradient clipping: 1.0
- Gradient accumulation: 1
Hardware
Training was performed on:
- RTX 6000 Blackwell (96GB)
π Evaluation
Synthetic Data Expansion Benchmark
A curated set of 30 input/target pairs was programmatically expanded using a Python script. Metrics include seed pairs covered, total variation count, and semantic quality. The task is to generate 5 variations of each incoming pair.
Qualitative Behavior
- Overperforms in variation count
- Maintains strict semantic correctness
π Safety
This model is a synthetic data generator. It is not designed for conversational use and is not suitable for anything other than generating synthetic datasets.
It is not designed for:
- Political advice
- Medical advice
- Legal advice
- General-purpose conversation
β οΈ Limitations
- Not a general assistant
- Not trained for coding, math, or open-domain reasoning
- May refuse tasks outside the Madlab SDG scope
