CoolFace
Modelpublic

chitransh001/googlenet-imagenet100

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes
Model Card

GoogLeNet (Inception v1) — ImageNet100

Trained from scratch on ImageNet100. No pretrained weights. No transfer learning. Just raw training.

Final val accuracy: 92.3%


Model Details

ArchitectureGoogLeNet (Inception v1)
DatasetImageNet100 (130,000 images, 100 classes)
Training time~45 hours across multiple sessions
HardwareKaggle T4 x2 (free tier) + Google Colab
FrameworkPyTorch
Parameters~6.8M trainable

Training Config

python
optimizer    = Adam(lr=0.001, weight_decay=1e-4)
scheduler    = CosineAnnealingLR(T_max=100, eta_min=1e-6)
criterion    = CrossEntropyLoss(label_smoothing=0.1)
batch_size   = 64
epochs       = 100
aux_weight   = 0.3  # auxiliary classifier loss weight

Data Augmentation

python
# Train
transforms.RandomResizedCrop(224)
transforms.RandomHorizontalFlip()
transforms.ColorJitter(brightness=0.2, contrast=0.2, saturation=0.2)
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])

# Val
transforms.Resize(256)
transforms.CenterCrop(224)
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])

Architecture Details

Standard GoogLeNet (Inception v1) as described in the original paper:

  • —9 Inception blocks
  • —2 auxiliary classifiers (weighted 0.3 during training, disabled at inference)
  • —Global average pooling before classifier
  • —Dropout 0.4 on both main and auxiliary heads
  • —BatchNorm added to all conv blocks (not in original paper — improves training stability)

Key Bug Worth Documenting

The ImageNet100 Kaggle dataset (ambityga/imagenet100) splits classes across 4 folders (train.X1 to train.X4) with ~25 unique classes per folder — not samples of shared classes.

Naively using ConcatDataset on 4 ImageFolder objects gives each shard independent 0-24 indices, making label 0 mean four different things. Val accuracy stays pinned at ~1% (exact random chance for 100 classes) regardless of training time or LR.

Fix — remap all shards to a single global class index before concatenating:

python
global_classes      = sorted(set(cls for ds in train_datasets for cls in ds.classes))
global_class_to_idx = {cls: i for i, cls in enumerate(global_classes)}

def remap_dataset(ds, mapping):
    old_idx_to_class = {v: k for k, v in ds.class_to_idx.items()}
    remap       = {old_idx: mapping[cls] for old_idx, cls in old_idx_to_class.items()}
    ds.samples  = [(path, remap[label]) for path, label in ds.samples]
    ds.targets  = [remap[label] for label in ds.targets]
    ds.class_to_idx = mapping
    ds.classes  = global_classes

for ds in train_datasets:
    remap_dataset(ds, global_class_to_idx)
remap_dataset(val_dataset, global_class_to_idx)

Usage

python
import torch
from model import Inception  # your model definition file

device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')

model = Inception(num_classes=100)
ckpt  = torch.load('best_model.pth', map_location=device)
model.load_state_dict(ckpt['model_state_dict'])
model.eval()
model.to(device)

# inference
with torch.no_grad():
    outputs, _, _ = model(images)  # returns (main, aux1, aux2) — aux are None at eval
    preds = outputs.argmax(dim=1)

Results

MetricValue
Best Val Accuracy92.3%
DatasetImageNet100
Training images130,000
Val images5,000

What Made the Difference

Things that actually moved the needle vs things that didn't:

Helped a lot:

  • —Label remapping fix (was literally the difference between 1% and learning)
  • —CosineAnnealingLR over ReduceLROnPlateau
  • —Label smoothing 0.1
  • —Auxiliary classifier weight 0.3
  • —Gradient clipping max_norm=5.0

Helped somewhat:

  • —BatchNorm in conv blocks
  • —Dropout 0.4 (was 0.7 on aux heads — too aggressive)
  • —Separate train/val transforms with augmentation

Author

Chitransh Panwar B.Tech CSE — JIIT Noida GitHub · LinkedIn · HuggingFace

Trained entirely on free-tier GPUs across multiple overnight Kaggle sessions. Every checkpoint survived via HuggingFace Hub auto-upload.