puttatidam/liputan6-ind-bitextmining
liputan6-ind-bitextmining Deduplicated copy of kornwtp/liputan6-ind-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/liputan6-ind-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: test, validation What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/liputan6-ind-bitextmining.
liputan6-ind-bitextmining
Deduplicated copy of `kornwtp/liputan6-ind-bitextmining`, part of the SEA-BED data-quality work.
- Source dataset:
kornwtp/liputan6-ind-bitextmining - Deduplicated on: 2026-09-04
- Task type: bitext_mining
- Splits: test, validation
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no row was removed for being low quality.
Row counts and the exact policy used are recorded in bitext_mining_dedup_summary.csv in the analysis repo. This dataset is not identical to the published SEA-BED data -- scores computed on it are not directly comparable to published SEA-BED numbers.
