Mish was obtained by experimenting with functions similar to Swish (SiLU, see above). It is non-monotonic (has a "bump") like Swish. The main new feature is that it exhibits a "self-regularizing" behavior attributed to a term in its first derivative.[34][35]
where is a hyperparameter that determines the "size" of the curved region near . (For example, letting yields ReLU, and letting yields the metallic mean function.) Squareplus shares many properties with softplus: It is monotonic, strictly positive, approaches 0 as , approaches the identity as , and is smooth. However, squareplus can be computed using only algebraic functions, making it well-suited for settings where computational resources or instruction sets are limited. Additionally, squareplus requires no special consideration to ensure numerical stability when is large.
DELU
ExtendeD Exponential Linear Unit (DELU, 2023) is an activation function which is smoother within the neighborhood of zero and sharper for bigger values, allowing better allocation of neurons in the learning process for higher performance. Thanks to its unique design, it has been shown that DELU may obtain higher classification accuracy than ReLU and ELU.[37]
In these formulas, , and are hyperparameter values which could be set as default constraints , and , as done in the original work.
↑Brownlee, Jason (8 January 2019). "A Gentle Introduction to the Rectified Linear Unit (ReLU)". Machine Learning Mastery. Retrieved 8 April 2021.
↑Liu, Danqing (30 November 2017). "A Practical Guide to ReLU". Medium. Retrieved 8 April 2021.
↑Ramachandran, Prajit; Barret, Zoph; Quoc, V. Le (October 16, 2017). "Searching for Activation Functions". arXiv:1710.05941 [cs.NE].
12345Xavier Glorot; Antoine Bordes; Yoshua Bengio (2011). Deep sparse rectifier neural networks(PDF). AISTATS. Rectifier and softplus activation functions. The second one is a smooth version of the first.
↑ László Tóth (2013). Phone Recognition with Deep Sparse Rectifier Neural Networks (PDF) . ICASSP .
1 2 Andrew L. Maas、Awni Y. Hannun、Andrew Y. Ng (2014)。整流器の非線形性によりニューラルネットワーク音響モデルが改善される。
↑ Hansel, D.; van Vreeswijk, C. (2002). "How noise contributes to contrast invariance of orientation tuning in cat visual cortex" . J. Neurosci. 22 (12): 5118– 5128. doi : 10.1523/JNEUROSCI.22-12-05118.2002 . PMC 6757721 . PMID 12077207 .
↑ Woodbury, G.; van der Zwan, R.; Gibson, WG (2002). "網膜局所マップと眼優位性列の共同発達に関する相関モデル". Vision Research . 42 (19): 2295– 2310. doi : 10.1016/s0042-6989(02)00190-6 . PMID 12220585 .
↑ Woodbury, G. (2004).網膜局所性と眼優位性の活動依存的発達に関する相関モデル(博士論文)。シドニー大学。
↑ Hahnloser, Richard; Seung, H. Sebastian (2000). "対称閾値線形ネットワークにおける許可セットと禁止セット" . Advances in Neural Information Processing Systems . 13 . MIT Press.
↑ Yann LeCun ; Leon Bottou ; Genevieve B. Orr; Klaus-Robert Müller (1998). "Efficient BackProp" (PDF) . In G. Orr; K. Müller (eds.). Neural Networks: Tricks of the Trade . Springer. 2018-08-31 のオリジナル(PDF)からアーカイブ済み。2012-12-07に取得。
1 2 He, Kaiming; Zhang, Xiangyu; Ren, Shaoqing; Sun, Jian (2015). "Delving Deep into Rectifiers: Surpassing Human-Level Performance on Image Net Classification". arXiv : 1502.01852 [ cs.CV ].
↑ Shang, Wenling; Sohn, Kihyuk; Almeida, Diogo; Lee, Honglak (2016-06-11). "Understanding and Improving Convolutional Neural Networks via Concatenated Rectified Linear Units" . Proceedings of the 33rd International Conference on Machine Learning . PMLR: 2217–2225 . arXiv : 1603.05201 .
↑ Dugas, Charles; Bengio, Yoshua; Bélisle, François; Nadeau, Claude; Garcia, René (2000-01-01). "Incorporating second-order functional knowledge for better option pricing" (PDF) . Proceedings of the 13th International Conference on Neural Information Processing Systems (NIPS'00) . MIT Press: 451– 457.シグモイドhは正の一次導関数を持つため、その原始関数 (softplus と呼ぶ) は凸関数である。
↑ 「Smooth Rectifier Linear Unit (SmoothReLU) Forward Layer」。Intel Data Analytics Acceleration Library の開発者ガイド。2017 年。2018年 12 月 4 日に取得。
↑ Barron, Jonathan T. (2021年12月22日). "Squareplus: Softplusに似た代数的整流器". arXiv : 2112.11687 [ cs.NE ].
↑ Çatalbaş, Burak; Morgül, Ömer (2023年8月16日). "Deep learning with Extended Exponential Linear Unit (DELU)" . Neural Computing and Applications . 35 (30): 22705– 22724. doi : 10.1007/s00521-023-08932-z . 2025年4月20日取得.