CoolFace
Modelpublic

empirischtech/yolox-pylon-xl

sourceHugging Faceapache-2.0updated 29d agoView on Hugging Face
0likes
Model Card

YOLOX-Pylon-XL

A YOLOX-X detector extended with one additional class, traffic cones (traffic_cone), on top of the 80 COCO classes, for 81 classes total. We train the adaptation so the original COCO capabilities are kept, not traded away. This checkpoint retains 98% of the official YOLOX-X baseline while the added cone class becomes the highest-scoring class in the model, ahead of all 80 originals.

Cones are our public demo class. The same adaptation recipe adds arbitrary custom classes such as defects, parts, or PPE to a proven detector without losing what it already knows.

Part of the YOLOX-Pylon family, S · M · L · XL. This is the accuracy ceiling of the family.

Built by Empirisch Tech GmbH (Vienna, Austria) under our Chaperone AI brand. See About Empirisch Tech below.

Results

We evaluate on COCO val2017 plus a held-out traffic-cone split, 81 classes in a single pass, 640×640 input, IoU 0.50:0.95 unless noted.

MetricValue
mAP 50:95 (81 classes)50.4
mAP 50:95, original 80 COCO classes only50.0
AP50 / AP7567.3 / 54.4
AP small / medium / large30.9 / 55.1 / 65.4
AR@10061.9
traffic_cone AP78.8
Inference (forward + NMS, batch 1, A100)6.01 ms

Two things stand out.

  • Retention held at 98%. The official YOLOX-X val2017 baseline is 51.2 mAP on COCO. After adding the cone class, this checkpoint keeps 50.0 on the same 80 classes, so we traded 1.2 points for an entire new class arriving at the top of the table.
  • The added class outscores all 80 originals. At 78.8 AP, traffic_cone is the single best- scoring class on this checkpoint, ahead of bear (76.4), cat (76.0), bus (75.9) and fire hydrant (75.8).

Highest and lowest scoring classes on this checkpoint.

Highest APLowest AP
traffic_cone78.8book17.9
bear76.4apple20.2
cat76.0handbag21.2
bus75.9broccoli22.7
fire hydrant75.8hair drier22.7

The bottom of the table is where the extra capacity shows up most. hair drier, the hardest class for every size in the family, more than doubles from 10.2 AP on the L checkpoint to 22.7 here.

Comparison with the base model

The comparison that matters is against the checkpoint we adapted from, with the same architecture, the same parameter count, the same FLOPs, and one extra class. Baseline figures are the official COCO val2017 numbers from the YOLOX model table.

ModelCOCO mAP 50:95ParamsFLOPsCustom classesLicense
yolox_pylon_xl (this model)50.0 kept + traffic_cone 78.899.1M281.9Gcone added, COCO keptApache-2.0
YOLOX-X (base)51.299.1M281.9GCOCO onlyApache-2.0

Reading that table, the adaptation costs 1.2 mAP on the original 80 classes and buys an entire new class at 78.8 AP, which lands above every one of those 80. Nothing else about the model changes. Parameters, FLOPs, and inference cost are the same as stock YOLOX-X, and Apache-2.0 carries over from the base, so the weights can be deployed commercially with no per-deployment license and no obligation to open-source derivative work.

Siblings for scale, same recipe and same eval protocol.

Family membermAP (81 cls)COCO keptCone APInference
yoloxpylons42.041.674.51.7 ms
yoloxpylonm47.246.877.52.6 ms
yoloxpylonl48.948.578.63.7 ms
yoloxpylonxl50.450.078.86.0 ms

We measure inference as forward plus NMS on an A100. Those times are not comparable to the V100 figures published in the official YOLOX table.

Usage

The checkpoint loads with the official YOLOX codebase. The only change from stock YOLOX-X is num_classes = 81, with traffic_cone as class index 80.

python
import torch
from yolox.exp import get_exp
from yolox.utils import postprocess

# stock yolox-x exp, patched to 81 classes
exp = get_exp(exp_name="yolox-x")
exp.num_classes = 81

model = exp.get_model()
ckpt = torch.load("yolox_pylon_xl.pth", map_location="cpu")
model.load_state_dict(ckpt["model"])
model.eval().cuda()

# img is a float32 tensor [1, 3, 640, 640], preprocessed YOLOX-style
with torch.no_grad():
    outputs = model(img)
outputs = postprocess(outputs, num_classes=81, conf_thre=0.25, nms_thre=0.45)

COCO_CLASSES = [...]                        # standard 80-class list
CLASSES = COCO_CLASSES + ["traffic_cone"]   # index 80

Or with the repo's demo tool.

bash
git clone https://github.com/Megvii-BaseDetection/YOLOX && cd YOLOX
python tools/demo.py image \
    -f exps/default/yolox_x.py \
    -c yolox_pylon_xl.pth \
    --path your_image.jpg --conf 0.25 --nms 0.45 --tsize 640 --device gpu
# patch exps/default/yolox_x.py with self.num_classes = 81 first

Training

  • Base. YOLOX-X (99.1M params), initialized from COCO-pretrained weights
  • Data. 147k images across 81 classes, COCO train2017 plus roughly 30k traffic-cone images, trained jointly so the original 80 classes stay in the mix during adaptation
  • Eval. COCO val2017 plus a held-out cone split, single 81-class evaluation pass
  • Input. 640×640

Intended use and limitations

We built this for roadside and infrastructure perception where traffic cones matter, such as work zones, lane closures, and autonomous driving research, and as a template for class-extension on YOLOX. The XL size is the accuracy ceiling of the released family and is the right pick for offline analysis, auto-labeling, and batch processing where 6 ms per frame is affordable.

One caveat we want to be clear about. Its advantage over yolox_pylon_l is concentrated in medium and large objects, at 55.1 and 65.4 AP against 53.2 and 63.5. On small objects it does not improve, scoring 30.9 against the L checkpoint's 31.1. Where small-object recall dominates the workload, L is the better trade at 60% of the latency.

The model detects boxes for 81 classes. It does not segment, track, or estimate distance. For fixed single-camera deployments, yolox_pylon_m offers most of the accuracy at 2.6 ms. For embedded and edge boards, yolox_pylon_s runs in 1.7 ms. As with any detector, we recommend validating on the target cameras before production use.

About Empirisch Tech

We are Empirisch Tech GmbH, a Vienna-based AI company, and we publish YOLOX-Pylon under our Chaperone AI brand. We run one recipe across three domains. We adapt a proven foundation model to a specific domain, keep what the base already knows, and ship the checkpoint together with the data it was trained on.

  • Language. Thinking-LQ-1.0 (84% MedQA, within 4 points of GPT-4o at ~20GB) and Coder-LQ-1.0
  • Physics. Chaperone-Flow-1.0 (Poseidon-B extended to new CFD regimes, 1.8% wake error) and Palace-LoRA (electromagnetics solver configs)
  • Vision. The YOLOX-Pylon family and a road-scene anomaly segmentation pipeline

These models power our production platforms, including NumericalAI (GPU physics simulation) and Simvera (industrial perception trained in simulation, deployed on real cameras). We self-host everything in our own Vienna datacenter and use no third- party model APIs. We are a member of the NVIDIA Inception and Microsoft for Startups programs, and our open checkpoints have passed 30,000 downloads on Hugging Face.

Custom builds. The cone class took one adaptation run. For other classes, cameras, or datasets, reach out via chaperoneai.com/contact.

License

Apache-2.0, matching the YOLOX base.

Citation

bibtex
@article{yolox2021,
  title={YOLOX: Exceeding YOLO Series in 2021},
  author={Ge, Zheng and Liu, Songtao and Wang, Feng and Li, Zeming and Sun, Jian},
  journal={arXiv preprint arXiv:2107.08430},
  year={2021}
}