iput-tk230215/steel-ocr-dataset-960
Steel OCR Dataset (Resized 960px) 鉄骨の手書き製品コード認識のための学習データセット(リサイズ版)。 学習高速化のため、画像を長辺960pxにリサイズしたバージョン。 バウンディングボックス座標も同じスケールで変換済み。 オリジナル版との違い 項目 オリジナル リサイズ版 総ピクセル数 5077.8 MP 215.1 MP 削減率 - 95.8% 用途 評価、高精度学習 高速学習 データセット構成 train / val はリーク無しで分割済み(画像レベルで互いに素、sha256 + dHash 監査で 0 leak)。 固定 14 枚 test セットとも画像・クロップが重複しない。 train_data/ ├── det/ # 検出モデル用 │ ├── train.txt # 学習データラベル (222画像) │ ├── val.txt… See the full description on the dataset page: https://huggingface.co/datasets/iput-tk230215/steel-ocr-dataset-960.
Clean DET val.txt: drop non-code boxes + code-less images (is_code filter). See data_quality_audit_20260601.
Clean DET train.txt: drop non-code boxes + code-less images (is_code filter). See data_quality_audit_20260601.
Update README counts to leak-free split (222/95/317/323/142) + train/val rationale
Remove stale files after leak-free refresh (2 test-leak images + abandoned negative_v1 labels)
Refresh 960px dataset from leak-free split (train 222 / val 95)
Add negative sample file: val_with_neg.txt
Add negative sample file: train_with_neg.txt
Add negative sample file: val_with_neg.txt
Add negative sample file: train_with_neg.txt
Upload README.md with huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
initial commit
