maayan890/miluim-career-bridge
Miluim Career Bridge - Dataset Synthetic Hebrew-context job-ad pairs translating Israeli reserve-service (miluim) experience into civilian job language, across 8 categories. Generated with Qwen/Qwen2.5-1.5B-Instruct (see notebooks/01_generation.ipynb for the full pipeline). ⚠️ Note: the Hub split shown above (train) is a technical wrapper around the single parquet file. The dataset's actual train/test division is the split column inside the data (9536 train / 1000 test) - filter… See the full description on the dataset page: https://huggingface.co/datasets/maayan890/miluim-career-bridge.
notebook 04 with outputs
04: fine-tune outcome (adopted or not)
notebook 04 with outputs
04: fine-tune outcome (adopted or not)
notebook 04 with outputs
04: fine-tune outcome (adopted or not)
notebook 03 with outputs
03: model comparison + winner
notebook 03 with outputs
03: model comparison + winner
notebook 02 with outputs
notebook 02 with outputs
02: dataset card + Hub viewer config fix
notebook 01 with outputs
data: generation stats
notebook 01 with outputs
remove checkpoint
data: generation stats
data: dataset (10536 rows)
rejected batch 288
checkpoint batch 288: 10536 rows
rejected batch 256
checkpoint batch 256: 9394 rows
rejected batch 224
checkpoint batch 224: 8228 rows
rejected batch 192
checkpoint batch 192: 7056 rows
rejected batch 159
checkpoint batch 159: 5835 rows
rejected batch 127
checkpoint batch 127: 4663 rows
rejected batch 96
checkpoint batch 96: 3510 rows
rejected batch 64
checkpoint batch 64: 2352 rows
rejected batch 32
checkpoint batch 32: 1176 rows
notebook 04 with outputs
04: fine-tune outcome (adopted or not)
notebook 03 with outputs
03: FAISS index
03: embedding row-id order
03: bge-small corpus embeddings
03: comparison chart
03: model comparison + winner
03: FAISS index
03: embedding row-id order
03: MiniLM-L6 corpus embeddings
03: comparison chart
03: model comparison + winner
