why the loss is nan while the loss of the same data training with SSD and RefineDet is normal
why the loss is nan while the loss of the same data training with SSD and RefineDet is normal