chitransh001/googlenet-imagenet100
GoogLeNet (Inception v1) — ImageNet100
Trained from scratch on ImageNet100. No pretrained weights. No transfer learning. Just raw training.
Final val accuracy: 92.3%
Model Details
Training Config
optimizer = Adam(lr=0.001, weight_decay=1e-4)
scheduler = CosineAnnealingLR(T_max=100, eta_min=1e-6)
criterion = CrossEntropyLoss(label_smoothing=0.1)
batch_size = 64
epochs = 100
aux_weight = 0.3 # auxiliary classifier loss weightData Augmentation
# Train
transforms.RandomResizedCrop(224)
transforms.RandomHorizontalFlip()
transforms.ColorJitter(brightness=0.2, contrast=0.2, saturation=0.2)
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
# Val
transforms.Resize(256)
transforms.CenterCrop(224)
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])Architecture Details
Standard GoogLeNet (Inception v1) as described in the original paper:
- 9 Inception blocks
- 2 auxiliary classifiers (weighted 0.3 during training, disabled at inference)
- Global average pooling before classifier
- Dropout 0.4 on both main and auxiliary heads
- BatchNorm added to all conv blocks (not in original paper — improves training stability)
Key Bug Worth Documenting
The ImageNet100 Kaggle dataset (ambityga/imagenet100) splits classes across 4 folders (train.X1 to train.X4) with ~25 unique classes per folder — not samples of shared classes.
Naively using ConcatDataset on 4 ImageFolder objects gives each shard independent 0-24 indices, making label 0 mean four different things. Val accuracy stays pinned at ~1% (exact random chance for 100 classes) regardless of training time or LR.
Fix — remap all shards to a single global class index before concatenating:
global_classes = sorted(set(cls for ds in train_datasets for cls in ds.classes))
global_class_to_idx = {cls: i for i, cls in enumerate(global_classes)}
def remap_dataset(ds, mapping):
old_idx_to_class = {v: k for k, v in ds.class_to_idx.items()}
remap = {old_idx: mapping[cls] for old_idx, cls in old_idx_to_class.items()}
ds.samples = [(path, remap[label]) for path, label in ds.samples]
ds.targets = [remap[label] for label in ds.targets]
ds.class_to_idx = mapping
ds.classes = global_classes
for ds in train_datasets:
remap_dataset(ds, global_class_to_idx)
remap_dataset(val_dataset, global_class_to_idx)Usage
import torch
from model import Inception # your model definition file
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = Inception(num_classes=100)
ckpt = torch.load('best_model.pth', map_location=device)
model.load_state_dict(ckpt['model_state_dict'])
model.eval()
model.to(device)
# inference
with torch.no_grad():
outputs, _, _ = model(images) # returns (main, aux1, aux2) — aux are None at eval
preds = outputs.argmax(dim=1)Results
What Made the Difference
Things that actually moved the needle vs things that didn't:
Helped a lot:
- Label remapping fix (was literally the difference between 1% and learning)
- CosineAnnealingLR over ReduceLROnPlateau
- Label smoothing 0.1
- Auxiliary classifier weight 0.3
- Gradient clipping max_norm=5.0
Helped somewhat:
- BatchNorm in conv blocks
- Dropout 0.4 (was 0.7 on aux heads — too aggressive)
- Separate train/val transforms with augmentation
Author
Chitransh Panwar B.Tech CSE — JIIT Noida GitHub · LinkedIn · HuggingFace
Trained entirely on free-tier GPUs across multiple overnight Kaggle sessions. Every checkpoint survived via HuggingFace Hub auto-upload.
