CoolFace
Datasetpublic

puttatidam/liputan6-ind-bitextmining

liputan6-ind-bitextmining Deduplicated copy of kornwtp/liputan6-ind-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/liputan6-ind-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: test, validation What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/liputan6-ind-bitextmining.

sourceHugging Faceupdated 6d agoView on Hugging Face
0likes84downloads
Dataset Card

liputan6-ind-bitextmining

Deduplicated copy of `kornwtp/liputan6-ind-bitextmining`, part of the SEA-BED data-quality work.

  • Source dataset: kornwtp/liputan6-ind-bitextmining
  • Deduplicated on: 2026-09-04
  • Task type: bitext_mining
  • Splits: test, validation

What changed

Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no row was removed for being low quality.

Row counts and the exact policy used are recorded in bitext_mining_dedup_summary.csv in the analysis repo. This dataset is not identical to the published SEA-BED data -- scores computed on it are not directly comparable to published SEA-BED numbers.