Hasty Briefsbeta

双语

The Softmax function and its derivative

a day ago
  • Softmax transforms an N-dimensional vector of arbitrary real values into a vector of probabilities in (0, 1) that sum to 1.0, preserving relative order.
  • The softmax function is a 'soft' version of the maximum function, giving a proportionally larger share to the maximal element while distributing some probability to others.
  • Softmax has a probabilistic interpretation, making it suitable for multiclass classification tasks in machine learning, such as multiclass logistic regression.
  • The derivative of softmax is a Jacobian matrix, and it can be computed using cases with the Kronecker delta: ∂S_i/∂z_j = S_i (δ_{ij} - S_j).
  • Numerical stability when computing softmax can be improved by shifting inputs by their maximum, preventing overflow or underflow due to exponentiation.
  • In a softmax layer (fully-connected matrix multiplication followed by softmax), the derivative with respect to weights can be computed via the multivariate chain rule, resulting in a sparse Jacobian.
  • Cross-entropy loss is commonly used with softmax for training; its gradient with respect to weights simplifies to (P - Y) times the input, where P is the softmax output and Y is the one-hot true label.