datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lfm-2.5-1.2b-instruct-308xTrace of LFM2.5 1.2B Instruct LLM.
Data count (Total: 308):
English - 198
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
otb-augmented-lfm1.2b
davanstrien/otb-augmented-lfm1.2b
LLM-annotated dataset produced by classify-and-augment.
Configuration
Model: LiquidAI/LFM2.5-1.2B-Instruct
Labels: jim_crow, no_jim_crow
Input rows: 200
Output rows: 200
Label distribution
Label
Real
Synthetic
Total
jim_crow
22
0
22
no_jim_crow
178
0
178
Synthesis audit
Class
Needed
Generated
Validated
Kept
Acceptance
jim_crow
28
160
1
0
0.6%
Acceptance = synthetic candidates… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/otb-augmented-lfm1.2b.otb-baseline-lfm1.2b
davanstrien/otb-baseline-lfm1.2b
LLM-annotated dataset produced by classify-and-augment.
Configuration
Model: LiquidAI/LFM2.5-1.2B-Instruct
Labels: jim_crow, no_jim_crow
Input rows: 200
Output rows: 200
Label distribution
Label
Real
Synthetic
Total
jim_crow
22
0
22
no_jim_crow
178
0
178
