CoolFace
Modelpublic

ManhHoDinh/lfm25-titlegen-dpo

sourceHugging Faceotherupdated 18d agoView on Hugging Face
0likes1kdownloads
Model Card

lfm25-titlegen-dpo

Abstract

This model release is evaluated as the Direct preference optimization variant in a reproducible, 17-language conversation-title benchmark. The study contains 3,400 inputs per model and distinguishes 2,550 evaluation inputs from 850 development inputs. Each generated title receives one automated Luna assessment on seven explicit ordinal criteria where applicable. This card reports observed scores and denominators; it does not establish calibrated acceptance rates, causal training effects, or production readiness.

Dataset and detailed results · Interactive benchmark · Pinned aggregate JSON

1. Model and task

The intended task is to generate a concise, informative title from a user conversation in its target language. Titles should preserve central intent, essential entities and literals, and meaningful uncertainty. Training details and the original model documentation are retained verbatim in Appendix A. Variant names identify checkpoints; this comparison does not isolate the causal effect of a training method.

  • —Checkpoint evaluated: c77e8b5fa2fc4f7d562bf599fd5433d04d12c467
  • —Repository: ManhHoDinh/lfm25-titlegen-dpo
  • —Weight manifest fingerprint: 86b7ee563e1ef3a4c6147a9ed94e69f325100a22bd52f88e7f14ee30a1882f47
  • —This documentation update changes no model weights.

2. Experimental design

The mixed-v2 set contains 200 inputs in each of 17 languages. The primary evaluation partition has 150 inputs per language; 50 development inputs per language are reported separately. All four local model variants use the same input set. Generation uses greedy decoding, float32, and at most 32 new tokens (Torch 2.14.0; Transformers 5.13.1). Local generation timing is not a deployment load test.

Original benchmark prompts contain the user message only, passed through each local tokenizer's chat template with an assistant generation prompt. The separately generated Luna reference adds an explicit title-task system instruction; it is not a matched-prompt baseline.

The pointwise judge is cx/gpt-5.6-luna, temperature 0, low reasoning effort, with structured JSON output. The rubric covers relevance, intent coverage, faithfulness, entity fidelity, concision, fluency and safety. Each applicable criterion is scored from 0 to 4. Criterion applicability is input-specific: unassessed criteria are excluded rather than assigned zero. Pending agent-authored expectations guide evaluation; they are not native-speaker gold labels. Requests, raw replies and parsed checkpoints are validated before aggregate publication.

3. Primary evaluation results

Cells show mean / 4 (applicable n) over the 2,550 evaluation inputs. Means summarize ordinal scores; they are not percentages or pass probabilities.

Variantrelevanceintent_coveragefaithfulnessentity_fidelityconcisionfluencysafety
sft2.547 (n=2,550)1.455 (n=2,505)2.976 (n=2,550)2.064 (n=2,229)3.551 (n=2,550)2.694 (n=2,550)3.868 (n=318)
dpo2.591 (n=2,550)1.478 (n=2,505)3.018 (n=2,550)2.105 (n=2,229)3.573 (n=2,550)2.676 (n=2,550)3.896 (n=318)
curriculum2.524 (n=2,550)1.404 (n=2,505)3.021 (n=2,550)2.074 (n=2,229)3.511 (n=2,550)2.694 (n=2,550)3.928 (n=318)
multilingual2.467 (n=2,550)1.249 (n=2,505)2.611 (n=2,550)2.056 (n=2,229)3.245 (n=2,550)2.573 (n=2,550)3.912 (n=318)

For this checkpoint, the lowest observed mean criterion is intent_coverage (1.478/4; n=2,505). This identifies cases to inspect; it does not by itself diagnose a data or training cause.

Results by language

Languagerelevanceintent_coveragefaithfulnessentity_fidelityconcisionfluencysafety
de3.027 (n=150)1.987 (n=150)3.273 (n=150)2.644 (n=149)3.933 (n=150)3.273 (n=150)3.900 (n=20)
en3.453 (n=150)2.333 (n=150)3.673 (n=150)2.828 (n=145)4.000 (n=150)3.893 (n=150)4.000 (n=14)
es3.233 (n=150)2.307 (n=150)3.533 (n=150)2.535 (n=127)3.980 (n=150)3.720 (n=150)3.917 (n=12)
fil2.153 (n=150)1.060 (n=150)2.733 (n=150)1.904 (n=135)3.293 (n=150)1.833 (n=150)3.824 (n=17)
fr3.167 (n=150)2.120 (n=150)3.360 (n=150)2.627 (n=142)3.980 (n=150)3.453 (n=150)3.929 (n=14)
id2.067 (n=150)0.893 (n=150)2.600 (n=150)1.619 (n=105)3.260 (n=150)2.247 (n=150)3.850 (n=20)
ja2.967 (n=150)1.807 (n=150)3.320 (n=150)2.338 (n=145)3.973 (n=150)3.533 (n=150)4.000 (n=20)
ko3.180 (n=150)2.060 (n=150)3.460 (n=150)2.484 (n=124)3.967 (n=150)3.660 (n=150)3.792 (n=24)
lo1.780 (n=150)0.460 (n=150)2.520 (n=150)1.286 (n=105)2.507 (n=150)0.753 (n=150)4.000 (n=20)
ms2.333 (n=150)1.207 (n=150)2.827 (n=150)1.914 (n=116)3.560 (n=150)2.493 (n=150)3.733 (n=15)
my1.373 (n=150)0.520 (n=150)1.680 (n=150)1.035 (n=114)3.020 (n=150)0.640 (n=150)4.000 (n=27)
pt3.120 (n=150)2.113 (n=150)3.473 (n=150)2.420 (n=131)3.987 (n=150)3.627 (n=150)3.882 (n=17)
ru2.733 (n=150)1.352 (n=105)2.987 (n=150)2.128 (n=133)3.933 (n=150)2.820 (n=150)3.722 (n=18)
ta1.787 (n=150)0.453 (n=150)2.460 (n=150)1.385 (n=135)2.633 (n=150)1.213 (n=150)3.857 (n=14)
th2.053 (n=150)0.927 (n=150)2.827 (n=150)1.297 (n=128)2.947 (n=150)1.967 (n=150)3.875 (n=16)
vi2.447 (n=150)1.333 (n=150)3.187 (n=150)1.931 (n=145)3.787 (n=150)2.920 (n=150)4.000 (n=24)
zh3.173 (n=150)2.153 (n=150)3.393 (n=150)2.733 (n=150)3.973 (n=150)3.447 (n=150)3.885 (n=26)

Development and combined partitions

PartitionInputsrelevanceintent_coveragefaithfulnessentity_fidelityconcisionfluencysafety
development8502.548 (n=850)1.462 (n=833)2.932 (n=850)2.109 (n=740)3.555 (n=850)2.613 (n=850)3.923 (n=104)
combined3,4002.580 (n=3,400)1.474 (n=3,338)2.996 (n=3,400)2.106 (n=2,969)3.568 (n=3,400)2.660 (n=3,400)3.903 (n=422)

4. Detailed diagnostics and Luna reference

The linked dataset and interactive report expose per-input outputs, original judge reasons, and supplemental feedback. The supplemental protocol generates a Luna reference title and assesses five anonymously ordered candidates. It uses a different prompting protocol and output budget. Luna also judges its own generated reference, so self-judge bias is a material limitation. The verified release at dataset revision 779c3bedc27a2eaa6a173a69eebb4deb1215bf9d contains 3,400 Luna reference titles and 3,400 complete feedback rows (17,000 candidate feedback records). Feedback provenance is 3,227 original responses, 172 candidate-remediation-v2 responses and one issue-schema-remediation-v3 response. Paired comparisons are separated by feedback protocol; remediation cases are a selected failure subset. Supplemental scores are kept separate from the pointwise results in Section 3. Suggested titles are proposals, not gold labels.

5. Limitations and production readiness

Production readiness: `NOT_ESTABLISHED`. The study has 200 inputs per language versus the configured 500-item requirement, one automated judge versus three, and no verified native audit versus the required 100 items per language. Deterministic compliance rates, calibrated semantic acceptance, native agreement, real deployment load tests, fairness acceptance gaps and regression gates remain unestablished. Ordinal rubric means cannot be directly substituted for similarly named production-policy thresholds.

The published evaluation inputs are now exposed. Tuning on these inputs or their feedback invalidates future held-out claims for this set; use a newly collected, independently audited evaluation set for subsequent optimization claims.

Additional limitations include pending annotation review, unmeasured independent judge agreement, potential language-dependent judge bias, and incomplete evidence of exclusion from every training corpus. A reached-limit flag is a proxy, not proof of truncation. Low intent coverage motivates a human label and intent-diversity audit; low entity fidelity motivates entity-retention examples and a controlled output-cap experiment. These are hypotheses requiring held-out validation, not demonstrated optimization gains.

6. Reproducibility and release artifacts

The aggregate JSON preserves score distributions and applicable denominators. Generation and judge checkpoints bind inputs, output manifests, prompts and response evidence to immutable identities. The dataset repository distributes versioned diagnostic snapshots and source artifacts; the Space provides an inspection interface. Use the checkpoint revision above and the dataset snapshot revision when citing or comparing results. Inspect the license of each distributed artifact before reuse; the original model license metadata remains unchanged.

Inference quickstart

The following matches the benchmark's greedy, user-only title-task interface. It is an illustrative invocation; the recorded benchmark runtime remains authoritative.

bash
pip install torch==2.14.0 transformers==5.13.1
python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = 'ManhHoDinh/lfm25-titlegen-dpo'
revision = 'c77e8b5fa2fc4f7d562bf599fd5433d04d12c467'
tokenizer = AutoTokenizer.from_pretrained(repo_id, revision=revision, trust_remote_code=False)
model = AutoModelForCausalLM.from_pretrained(
    repo_id, revision=revision, torch_dtype=torch.float32,
    trust_remote_code=False, use_safetensors=True,
).eval()
messages = [{'role': 'user', 'content': 'How can I plan a three-day trip to Hanoi with children?'}]
inputs = tokenizer.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_tensors='pt', return_dict=True,
)
with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=32, do_sample=False, num_beams=1)
title_tokens = output[0, inputs['input_ids'].shape[-1]:]
print(tokenizer.decode(title_tokens, skip_special_tokens=True).strip())

Citation

This is a versioned software and benchmark release, not a claim of peer-reviewed publication.

bibtex
@misc{titlegen_mixed_v2_luna_2026,
  title = {TitleGen Mixed-v2: Multilingual Title Generation with Luna Evaluation},
  author = {ManhHoDinh},
  year = {2026},
  howpublished = {Hugging Face model and benchmark release},
  url = {https://huggingface.co/datasets/ManhHoDinh/titlegen-benchmark-luna},
  note = {Model dpo, evaluated revision c77e8b5fa2fc4f7d562bf599fd5433d04d12c467; cite the dataset snapshot revision used}
}

Appendix A. Historical model card

The original card is preserved below as historical documentation; the current benchmark qualifications and scope above govern interpretation of this new evaluation.

<details><summary>Original documentation</summary>

<!-- titlegen-research-history:start -->

LFM2.5 TitleGen DPO

This model is the direct preference optimization stage of the LFM2.5 TitleGen experiment. Model repository: ManhHoDinh/lfm25-titlegen-dpo. Benchmark publication timestamp: 2026-08-16T00:00:00+07:00.

Preliminary benchmark

The current model passed 293 of 300 evaluated English and Vietnamese examples (97.7%).

ModelPassedOverallEnglishVietnameseEvidence
DPO293 / 30097.7%99.5%94.1%Preliminary automated EN/VI
SFT v2291 / 30097.0%99.5%92.2%Preliminary automated EN/VI
Curriculum289 / 30096.3%98.5%92.2%Preliminary automated EN/VI

Language coverage

LanguageSamplesPassedRateStatus
German000.0%NOT_EVALUATED
English19819799.5%EVALUATED
Spanish000.0%NOT_EVALUATED
Filipino000.0%NOT_EVALUATED
French000.0%NOT_EVALUATED
Indonesian000.0%NOT_EVALUATED
Japanese000.0%NOT_EVALUATED
Korean000.0%NOT_EVALUATED
Lao000.0%NOT_EVALUATED
Malay000.0%NOT_EVALUATED
Burmese000.0%NOT_EVALUATED
Portuguese000.0%NOT_EVALUATED
Russian000.0%NOT_EVALUATED
Tamil000.0%NOT_EVALUATED
Thai000.0%NOT_EVALUATED
Vietnamese1029694.1%EVALUATED
Chinese000.0%NOT_EVALUATED

0 means no evaluated examples when the status is NOT_EVALUATED; it is not a measured zero score.

Methodology

This preliminary benchmark evaluates aggregate English and Vietnamese results with deterministic decoding (do_sample: false, max_new_tokens: 32). Automated rubric identifiers: 3-8_tu, khong_cham_cuoi, mot_dong, dung_ngon_ngu, khong_chep. The remaining contracted languages are shown explicitly as not evaluated.

Limitations

  • —Only English and Vietnamese have evaluated examples.
  • —Language correctness uses an automated heuristic.
  • —No native review or blind preference evidence is included.
  • —These aggregate automated results do not establish production readiness, causal improvement, or statistical significance.

Machine-readable results

See benchmark-report.json for the validated aggregate report.

<!-- titlegen-mixed-v2-luna:start -->

Mixed-v2 Luna benchmark: dpo

3,400 / 3,400 judgments for this model. Single Luna automated assessment on a 0..4 ordinal scale. Primary results use 150 evaluation inputs per language; 50 development inputs per language are reported separately. Agent-authored annotations are not native-speaker gold. Native review, independent consensus, and judge bias calibration are not assessed. This release covers only the named model. Missing applicability is excluded from criterion counts. Provider billing and upstream retry totals are unknown.

  • —Report (HTML)
  • —Aggregate JSON
  • —Methodology and per-language scores

<!-- titlegen-mixed-v2-luna:end --> <!-- titlegen-research-history:end -->

</details>