CoolFace
Datasetpublic

gaunernst/ms1mv3-wds

MS-Celeb-1M (v3) This dataset is introduced in the Lightweight Face Recognition Challenge at ICCV 2019. Paper. There are 5,179,510 images and 93,431 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112. This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_ (MS1M-RetinaFace). The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 100… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ms1mv3-wds.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes1.9kdownloads
Dataset Card

MS-Celeb-1M (v3)

This dataset is introduced in the Lightweight Face Recognition Challenge at ICCV 2019. Paper.

There are 5,179,510 images and 93,431 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112.

This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_ (MS1M-RetinaFace). The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 100 shards in total.

Usage

python
import webdataset as wds

url = "https://huggingface.co/datasets/gaunernst/ms1mv3-wds/resolve/main/ms1mv3-{{0000..0099}}.tar"
ds = wds.WebDataset(url).decode("pil").to_tuple("jpg", "cls")

img, label = next(iter(ds))