vudang449/AIO2025-Final-Exam-Dataset
version https://git-lfs.github.com/spec/v1 oid sha256:072d081dec61aad02275a7ed608c136a28f1de6672c7ff3a520bdbf7b0952e53 size 2225
v11.3: Fix toàn bộ audit
v11.2.1: Add cleandataset.py + 20-test suite
v11.2: Fix VI+English glue (31/31 PASS) + add dict-aware splitter
v11.1: Add cleandataset.py + 14 tests
v11 full upload: data + engine (vi_respace.py) + freq tables (vi_syllables.json 5052 entries, vi_syllable_freq.json 494 entries) + verify.py
v11: Apply VI DP respacing (frequency-weighted Viterbi). 384 splits across 97/145 records. dữliệu 62→0 occurrences. Skip-guards: digits/URL/code/English vocab/consonant clusters. Whitespace-strip validation: 0 mismatches (no chars added/removed). Verify: PASS_10_10.
v10: Fix 2 wrong-option records (qz_1_2, qz_2_4) from raw PDF. Full cross-check against raw PDF done. VI spacing + PUA + newline + Unicode noise = 0 issues across ALL fields. 145 records PASS_10_10.
v9: fix truncated/broken options (qz_10_8, ve_44, ve_20, ve_17, ve_16). Clean PUA chars (ex_42, ex_61). Fix VI spacing in ALL fields. 145 records PASS_10_10.
v8: ex_48 + qz_9_2 options fixed from raw PDF. All VI spacing fixed in ALL fields (41 fields). 145 records PASS_10_10.
v7: sửdụng added to whitelist — all VI spacing artifacts resolved. qz_9_3: 'Theo đoạn mã ở trên, lớp mô hình được sử dụng để xây dựng bài toán ATSC là gì?'
v6: position-based fix — all VI spacing artifacts resolved (sửdụng, thứtựlà, ởtrên, etc.)
v5: corpus-based VI spacing fix — ởtrên, quảcủa, thứtự auto-detected and fixed
Fix: join mid-sentence newlines + VI spacing (ở trên, sử dụng, hệ thống, mô hình) — PASS_10_10, verify.py updated
Fix: join 94 broken questions (newline mid-sentence split), e.g. qz_9_3 now reads as single sentence
Delete dataset_stats_upgraded.json with huggingface_hub
Delete aio2025_upgraded_hf.json with huggingface_hub
Cleanup: delete redundant CSV/MD files, keep only train/test JSONL
Cleanup: remove redundant files, keep only train/test JSONL + README
Re-upload: 145 verified questions, clean split (train=116, test=29)
Fix dataset format: proper train/test JSONL splits
Upload AIO2025 Final Exam Dataset - 145 verified questions
Upload README.md with huggingface_hub
Upload aio2025_final_canonical.csv with huggingface_hub
Upload aio2025_final_canonical.jsonl with huggingface_hub
Upload README.md with huggingface_hub
Delete export_final.py with huggingface_hub
Delete aio2025_final_canonical.md with huggingface_hub
Delete aio2025_final_canonical.csv with huggingface_hub
Delete dataset_info.json with huggingface_hub
Delete dataset_stats_canonical.json with huggingface_hub
Delete /_metadata.json with huggingface_hub
Delete /_stats_canonical.json with huggingface_hub
Upload folder using huggingface_hub
Delete aio2025_hf_canonical.json with huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Add dataset_info.json
Fix README YAML: task_categories=question-answering
Fix README metadata: task_categories=multiple-choice
Add HuggingFace-native JSON format
Add metadata files
Add CSV export
Add canonical JSONL dataset
Add README
initial commit
