pouyaardehkhani/ActTensor

ActTensor: Activation Functions for TensorFlow. https://pypi.org/project/ActTensor-tf/ Authors: Pouya Ardehkhani, Pegah Ardehkhani

Python

27

91 commits

updated Jun 11, 2024

See the code

README



ActTensor: Activation Functions for TensorFlow

license releases

What is it?

ActTensor is a Python package that provides state-of-the-art activation functions which facilitate using them in Deep Learning projects in an easy and fast manner.

Why not using tf.keras.activations?

As you may know, TensorFlow only has a few defined activation functions and most importantly it does not include newly-introduced activation functions. Wrting another one requires time and energy; however, this package has most of the widely-used, and even state-of-the-art activation functions that are ready to use in your models.

Requirements

Install the required dependencies by running the following command:

  • conda env create -f environment.yml

Where to get it?

The source code is currently hosted on GitHub at: https://github.com/pouyaardehkhani/ActTensor

Binary installers for the latest released version are available at the Python Package Index (PyPI)

# PyPI
pip install ActTensor-tf

License

MIT

How to use?

import tensorflow as tf
import numpy as np
from ActTensor_tf import ReLU # name of the layer

functional api

inputs = tf.keras.layers.Input(shape=(28,28))
x = tf.keras.layers.Flatten()(inputs)
x = tf.keras.layers.Dense(128)(x)
# wanted class name
x = ReLU()(x)
output = tf.keras.layers.Dense(10,activation='softmax')(x)

model = tf.keras.models.Model(inputs = inputs,outputs=output)

sequential api

model = tf.keras.models.Sequential([tf.keras.layers.Flatten(),
                                    tf.keras.layers.Dense(128),
                                    # wanted class name
                                    ReLU(),
                                    tf.keras.layers.Dense(10, activation = tf.nn.softmax)])

NOTE:

The main functions of the activation layers are also available, but they may be defined by different names. Check this for more information.

from ActTensor_tf import relu

Activations

Classes and Functions are available in ActTensor_tf

Activation NameUse CaseProsConsExample Usage in Known Network
SoftShrinkDenoising autoencodersGood for noise reductionLimited usage scenariosUsed in image denoising autoencoders
HardShrinkDenoising autoencodersEffective noise removalLimited usage scenariosUsed in image denoising autoencoders
GLUGated networksHelps with learning complex functionsRequires additional gating mechanismGated Linear Units in NLP models like ELMo
BilinearBilinear interpolationEfficient image processingNot used for non-image dataBilinear interpolation in super-resolution networks
ReGLUTransformer modelsEnhanced gating mechanismComputationally expensiveEnhanced transformer models
GeGLUTransformer modelsEnhanced gating mechanismComputationally expensiveEnhanced transformer models
SwiGLUTransformer modelsEnhanced gating mechanismComputationally expensiveEnhanced transformer models
SeGLUTransformer modelsEnhanced gating mechanismComputationally expensiveEnhanced transformer models
ReLUGeneral purposeSimple, efficient, avoids vanishing gradientsDying ReLU problemUsed in almost all CNN architectures like VGG, ResNet
IdentityLinear networksRetains input valuesNo non-linearityIdentity mapping in residual networks
StepBinary classificationSimple thresholdingNon-differentiableUsed in simple binary classifiers
SigmoidBinary classification, output layersSmooth gradient, probabilistic interpretationVanishing gradient problemOutput layer in binary classification networks
HardSigmoidLow-power devicesSimple and efficientNon-differentiableMobile networks for power efficiency
LogSigmoidBinary classification, probabilistic outputsStabilizes trainingVanishing gradient problemBinary classification in networks
SiLUAdvanced networksCombines ReLU and Sigmoid benefitsComputationally expensiveUsed in Swish-activated networks
PLinearCustomizable linear transformationFlexibilityRequires parameter tuningCustom layers in experimental networks
Piecewise-LinearCustomizable piecewise transformationsFlexibilityRequires parameter tuningCustom layers in experimental networks
Complementary Log-LogProbabilistic outputsUseful for binary classificationLimited use in deep networksOutput layers in certain probabilistic models
BipolarBinary classificationSimple bipolar outputNon-differentiableBinary classification networks
Bipolar-SigmoidBinary classificationCombines benefits of Sigmoid and BipolarVanishing gradient problemBinary classification networks
TanhHidden layersZero-centered output, smooth gradientVanishing gradient problemRNNs and LSTMs like in original LSTM paper
TanhShrinkDenoising autoencodersCombines Tanh with shrinkageLimited usage scenariosUsed in denoising autoencoders
LeCun's TanhHidden layersScaled Tanh for better performanceVanishing gradient problemApplied in LeNet-5 network
HardTanhLow-power devicesSimple and efficientNon-differentiableEfficient models for mobile devices
TanhExpAdvanced networksCombines Tanh and exponential benefitsComputationally expensiveExperimental deep networks
AbsoluteSimple tasksEasy to implementNon-differentiableSimple experimental networks
Squared-ReLUAdvanced networksCombines ReLU and squaring benefitsComputationally expensiveExperimental networks with custom activations
P-ReLUCustomizable ReLU variantLearnable parametersRequires parameter tuningVariants of ResNet
R-ReLURegularizationReduces overfittingComputationally expensiveApplied in CNNs for added regularization
LeakyReLUGeneral purposePrevents dying ReLU problemSlightly more computationally expensive than ReLULeakyReLU in networks like YOLO
ReLU6Mobile networksBounded outputDying ReLU problemEfficientNet and MobileNet
Mod-ReLUAdvanced networksCombines ReLU and modulationComputationally expensiveCustom experimental networks
Cosine-ReLUAdvanced networksCombines ReLU and cosine benefitsComputationally expensiveCustom experimental networks
Sin-ReLUAdvanced networksCombines ReLU and sine benefitsComputationally expensiveCustom experimental networks
ProbitProbabilistic outputsUseful for binary classificationLimited use in deep networksCertain probabilistic models
CosPeriodic tasksHandles periodicity wellNon-differentiableNetworks dealing with periodic signals
GaussianRadial basis functionsSmooth gradient, radial basis functionComputationally expensiveRadial basis function networks
MultiquadraticRadial basis functionsSmooth gradient, radial basis functionComputationally expensiveRadial basis function networks
Inverse-MultiquadraticRadial basis functionsSmooth gradient, radial basis functionComputationally expensiveRadial basis function networks
SoftPlusAdvanced networksSmooth approximation to ReLUComputationally expensiveExperimental networks
MishAdvanced networksSmooth gradient, non-monotonicComputationally expensiveExperimental networks
SMishAdvanced networksSmooth gradient, non-monotonicComputationally expensiveExperimental networks
P-SMishCustomizable Mish variantLearnable parametersRequires parameter tuningExperimental networks
SwishAdvanced networksSmooth gradient, non-monotonicComputationally expensiveEfficientNet
ESwishAdvanced networksSmooth gradient, non-monotonicComputationally expensiveExperimental networks
HardSwishLow-power devicesSimple and efficientNon-differentiableMobileNetV3
GCUAdvanced networksGradient-controlled unitsComputationally expensiveExperimental networks
CoLUAdvanced networksCombines linear and unit step benefitsComputationally expensiveExperimental networks
PELUCustomizable ELU variantLearnable parametersRequires parameter tuningCustom experimental networks
SELUSelf-normalizing networksMaintains mean and varianceRequires careful initialization and architecture choicesSelf-normalizing networks like in self-normalizing neural networks paper
CELUAdvanced networksContinuously differentiable ELUComputationally expensiveExperimental networks
ArcTanPeriodic tasksHandles periodicity wellNon-differentiableNetworks dealing with periodic signals
Shifted-SoftPlusAdvanced networksSmooth gradientComputationally expensiveExperimental networks
SoftmaxOutput layer for multi-class classificationConverts logits to probabilitiesNot suitable for hidden layersOutput layer in classification networks like AlexNet
LogitProbabilistic outputsUseful for binary classificationLimited use in deep networksCertain probabilistic models
GELUAdvanced networksCombines Gaussian and ReLU benefitsComputationally expensiveTransformer networks like BERT
SoftsignGeneral purposeSmooth approximation to sign functionSlower convergenceApplied in some RNN architectures
ELiSHAdvanced networksCombines ELU and Swish benefitsComputationally expensiveExperimental networks
HardELiSHLow-power devicesSimple and efficientNon-differentiableEfficient models for mobile devices
SerfAdvanced networksCombines several benefits of other functionsComputationally expensiveExperimental networks
ELUDeep networksSmooth gradient, avoids dying ReLU problemComputationally expensiveDeep CNNs like in ELU paper
PhishAdvanced networksCombines several benefits of other functionsComputationally expensiveExperimental networks
QReLUQuantized networksEfficient in low-bit precisionLess flexible than regular ReLUEfficient quantized networks
MQReLUQuantized networksEfficient in low-bit precisionLess flexible than regular ReLUEfficient quantized networks
FReLUAdvanced networksCombines ReLU and filter benefitsComputationally expensiveExperimental networks

Which activation functions it supports?

  1. Soft Shrink:

  1. Hard Shrink:

  1. GLU:

  1. Bilinear:
  1. ReGLU:

    ReGLU is an activation function which is a variant of GLU.

  1. GeGLU:

    GeGLU is an activation function which is a variant of GLU.

  1. SwiGLU:

    SwiGLU is an activation function which is a variant of GLU.

  1. SeGLU:

    SeGLU is an activation function which is a variant of GLU.

  2. ReLU:

  1. Identity:

    $f(x) = x$

  1. Step:

  1. Sigmoid:

  1. Hard Sigmoid:

  1. Log Sigmoid:

  1. SiLU:

  1. ParametricLinear:

    $f(x) = a*x$

  2. PiecewiseLinear:

    Choose some xmin and xmax, which is our "range". Everything less than than this range will be 0, and everything greater than this range will be 1. Anything else is linearly-interpolated between.

  1. Complementary Log-Log (CLL):

  1. Bipolar:

  1. Bipolar Sigmoid:

  1. Tanh:

  1. Tanh Shrink:

  1. LeCunTanh:

  1. Hard Tanh:

  1. TanhExp:

  1. ABS:

  1. SquaredReLU:

  1. ParametricReLU (PReLU):

  1. RandomizedReLU (RReLU):

  1. LeakyReLU:

  1. ReLU6:

  1. ModReLU:

  1. CosReLU:

  1. SinReLU:

  1. Probit:

  1. Cosine:

  1. Gaussian:

  1. Multiquadratic:

    Choose some point (x,y).

  1. InvMultiquadratic:

  1. SoftPlus:

  1. Mish:

  1. Smish:

  1. ParametricSmish (PSmish):

  1. Swish:

  1. ESwish:

  1. Hard Swish:

  1. GCU:

  1. CoLU:

  1. PELU:

  1. SELU:

    where $\alpha \approx 1.6733$ & $\lambda \approx 1.0507$

  1. CELU:

  1. ArcTan:

  1. ShiftedSoftPlus:

  1. Softmax:

  1. Logit:

  1. GELU:

  1. Softsign:

  1. ELiSH:

  1. Hard ELiSH:

  1. Serf:

  1. ELU:

  1. Phish:

  1. QReLU:

  1. modified QReLU (m-QReLU):

  1. FReLU:

Cite this repository

@software{Pouya_ActTensor_2022,
author = {Pouya, Ardehkhani and Pegah, Ardehkhani},
license = {MIT},
month = {7},
title = {{ActTensor}},
url = {https://github.com/pouyaardehkhani/ActTensor},
version = {1.0.0},
year = {2022}
}
activation
activation-code
activation-function
activation-functions
activation-methods
activations
ai
artificial-intelligence
data-science
deep-learning
deeplearning
deep-neural-networks
neural-network-architectures
neural-networks
python
python3
python-library
python-package
tensorflow
tensorflow2

pouyaardehkhani/ActTensor

ActTensor: Activation Functions for TensorFlow. https://pypi.org/project/ActTensor-tf/ Authors: Pouya Ardehkhani, Pegah Ardehkhani

Python

27

91 commits

updated Jun 11, 2024

See the code

README



ActTensor: Activation Functions for TensorFlow

license releases

What is it?

ActTensor is a Python package that provides state-of-the-art activation functions which facilitate using them in Deep Learning projects in an easy and fast manner.

Why not using tf.keras.activations?

As you may know, TensorFlow only has a few defined activation functions and most importantly it does not include newly-introduced activation functions. Wrting another one requires time and energy; however, this package has most of the widely-used, and even state-of-the-art activation functions that are ready to use in your models.

Requirements

Install the required dependencies by running the following command:

  • conda env create -f environment.yml

Where to get it?

The source code is currently hosted on GitHub at: https://github.com/pouyaardehkhani/ActTensor

Binary installers for the latest released version are available at the Python Package Index (PyPI)

# PyPI
pip install ActTensor-tf

License

MIT

How to use?

import tensorflow as tf
import numpy as np
from ActTensor_tf import ReLU # name of the layer

functional api

inputs = tf.keras.layers.Input(shape=(28,28))
x = tf.keras.layers.Flatten()(inputs)
x = tf.keras.layers.Dense(128)(x)
# wanted class name
x = ReLU()(x)
output = tf.keras.layers.Dense(10,activation='softmax')(x)

model = tf.keras.models.Model(inputs = inputs,outputs=output)

sequential api

model = tf.keras.models.Sequential([tf.keras.layers.Flatten(),
                                    tf.keras.layers.Dense(128),
                                    # wanted class name
                                    ReLU(),
                                    tf.keras.layers.Dense(10, activation = tf.nn.softmax)])

NOTE:

The main functions of the activation layers are also available, but they may be defined by different names. Check this for more information.

from ActTensor_tf import relu

Activations

Classes and Functions are available in ActTensor_tf

Activation NameUse CaseProsConsExample Usage in Known Network
SoftShrinkDenoising autoencodersGood for noise reductionLimited usage scenariosUsed in image denoising autoencoders
HardShrinkDenoising autoencodersEffective noise removalLimited usage scenariosUsed in image denoising autoencoders
GLUGated networksHelps with learning complex functionsRequires additional gating mechanismGated Linear Units in NLP models like ELMo
BilinearBilinear interpolationEfficient image processingNot used for non-image dataBilinear interpolation in super-resolution networks
ReGLUTransformer modelsEnhanced gating mechanismComputationally expensiveEnhanced transformer models
GeGLUTransformer modelsEnhanced gating mechanismComputationally expensiveEnhanced transformer models
SwiGLUTransformer modelsEnhanced gating mechanismComputationally expensiveEnhanced transformer models
SeGLUTransformer modelsEnhanced gating mechanismComputationally expensiveEnhanced transformer models
ReLUGeneral purposeSimple, efficient, avoids vanishing gradientsDying ReLU problemUsed in almost all CNN architectures like VGG, ResNet
IdentityLinear networksRetains input valuesNo non-linearityIdentity mapping in residual networks
StepBinary classificationSimple thresholdingNon-differentiableUsed in simple binary classifiers
SigmoidBinary classification, output layersSmooth gradient, probabilistic interpretationVanishing gradient problemOutput layer in binary classification networks
HardSigmoidLow-power devicesSimple and efficientNon-differentiableMobile networks for power efficiency
LogSigmoidBinary classification, probabilistic outputsStabilizes trainingVanishing gradient problemBinary classification in networks
SiLUAdvanced networksCombines ReLU and Sigmoid benefitsComputationally expensiveUsed in Swish-activated networks
PLinearCustomizable linear transformationFlexibilityRequires parameter tuningCustom layers in experimental networks
Piecewise-LinearCustomizable piecewise transformationsFlexibilityRequires parameter tuningCustom layers in experimental networks
Complementary Log-LogProbabilistic outputsUseful for binary classificationLimited use in deep networksOutput layers in certain probabilistic models
BipolarBinary classificationSimple bipolar outputNon-differentiableBinary classification networks
Bipolar-SigmoidBinary classificationCombines benefits of Sigmoid and BipolarVanishing gradient problemBinary classification networks
TanhHidden layersZero-centered output, smooth gradientVanishing gradient problemRNNs and LSTMs like in original LSTM paper
TanhShrinkDenoising autoencodersCombines Tanh with shrinkageLimited usage scenariosUsed in denoising autoencoders
LeCun's TanhHidden layersScaled Tanh for better performanceVanishing gradient problemApplied in LeNet-5 network
HardTanhLow-power devicesSimple and efficientNon-differentiableEfficient models for mobile devices
TanhExpAdvanced networksCombines Tanh and exponential benefitsComputationally expensiveExperimental deep networks
AbsoluteSimple tasksEasy to implementNon-differentiableSimple experimental networks
Squared-ReLUAdvanced networksCombines ReLU and squaring benefitsComputationally expensiveExperimental networks with custom activations
P-ReLUCustomizable ReLU variantLearnable parametersRequires parameter tuningVariants of ResNet
R-ReLURegularizationReduces overfittingComputationally expensiveApplied in CNNs for added regularization
LeakyReLUGeneral purposePrevents dying ReLU problemSlightly more computationally expensive than ReLULeakyReLU in networks like YOLO
ReLU6Mobile networksBounded outputDying ReLU problemEfficientNet and MobileNet
Mod-ReLUAdvanced networksCombines ReLU and modulationComputationally expensiveCustom experimental networks
Cosine-ReLUAdvanced networksCombines ReLU and cosine benefitsComputationally expensiveCustom experimental networks
Sin-ReLUAdvanced networksCombines ReLU and sine benefitsComputationally expensiveCustom experimental networks
ProbitProbabilistic outputsUseful for binary classificationLimited use in deep networksCertain probabilistic models
CosPeriodic tasksHandles periodicity wellNon-differentiableNetworks dealing with periodic signals
GaussianRadial basis functionsSmooth gradient, radial basis functionComputationally expensiveRadial basis function networks
MultiquadraticRadial basis functionsSmooth gradient, radial basis functionComputationally expensiveRadial basis function networks
Inverse-MultiquadraticRadial basis functionsSmooth gradient, radial basis functionComputationally expensiveRadial basis function networks
SoftPlusAdvanced networksSmooth approximation to ReLUComputationally expensiveExperimental networks
MishAdvanced networksSmooth gradient, non-monotonicComputationally expensiveExperimental networks
SMishAdvanced networksSmooth gradient, non-monotonicComputationally expensiveExperimental networks
P-SMishCustomizable Mish variantLearnable parametersRequires parameter tuningExperimental networks
SwishAdvanced networksSmooth gradient, non-monotonicComputationally expensiveEfficientNet
ESwishAdvanced networksSmooth gradient, non-monotonicComputationally expensiveExperimental networks
HardSwishLow-power devicesSimple and efficientNon-differentiableMobileNetV3
GCUAdvanced networksGradient-controlled unitsComputationally expensiveExperimental networks
CoLUAdvanced networksCombines linear and unit step benefitsComputationally expensiveExperimental networks
PELUCustomizable ELU variantLearnable parametersRequires parameter tuningCustom experimental networks
SELUSelf-normalizing networksMaintains mean and varianceRequires careful initialization and architecture choicesSelf-normalizing networks like in self-normalizing neural networks paper
CELUAdvanced networksContinuously differentiable ELUComputationally expensiveExperimental networks
ArcTanPeriodic tasksHandles periodicity wellNon-differentiableNetworks dealing with periodic signals
Shifted-SoftPlusAdvanced networksSmooth gradientComputationally expensiveExperimental networks
SoftmaxOutput layer for multi-class classificationConverts logits to probabilitiesNot suitable for hidden layersOutput layer in classification networks like AlexNet
LogitProbabilistic outputsUseful for binary classificationLimited use in deep networksCertain probabilistic models
GELUAdvanced networksCombines Gaussian and ReLU benefitsComputationally expensiveTransformer networks like BERT
SoftsignGeneral purposeSmooth approximation to sign functionSlower convergenceApplied in some RNN architectures
ELiSHAdvanced networksCombines ELU and Swish benefitsComputationally expensiveExperimental networks
HardELiSHLow-power devicesSimple and efficientNon-differentiableEfficient models for mobile devices
SerfAdvanced networksCombines several benefits of other functionsComputationally expensiveExperimental networks
ELUDeep networksSmooth gradient, avoids dying ReLU problemComputationally expensiveDeep CNNs like in ELU paper
PhishAdvanced networksCombines several benefits of other functionsComputationally expensiveExperimental networks
QReLUQuantized networksEfficient in low-bit precisionLess flexible than regular ReLUEfficient quantized networks
MQReLUQuantized networksEfficient in low-bit precisionLess flexible than regular ReLUEfficient quantized networks
FReLUAdvanced networksCombines ReLU and filter benefitsComputationally expensiveExperimental networks

Which activation functions it supports?

  1. Soft Shrink:

  1. Hard Shrink:

  1. GLU:

  1. Bilinear:
  1. ReGLU:

    ReGLU is an activation function which is a variant of GLU.

  1. GeGLU:

    GeGLU is an activation function which is a variant of GLU.

  1. SwiGLU:

    SwiGLU is an activation function which is a variant of GLU.

  1. SeGLU:

    SeGLU is an activation function which is a variant of GLU.

  2. ReLU:

  1. Identity:

    $f(x) = x$

  1. Step:

  1. Sigmoid:

  1. Hard Sigmoid:

  1. Log Sigmoid:

  1. SiLU:

  1. ParametricLinear:

    $f(x) = a*x$

  2. PiecewiseLinear:

    Choose some xmin and xmax, which is our "range". Everything less than than this range will be 0, and everything greater than this range will be 1. Anything else is linearly-interpolated between.

  1. Complementary Log-Log (CLL):

  1. Bipolar:

  1. Bipolar Sigmoid:

  1. Tanh:

  1. Tanh Shrink:

  1. LeCunTanh:

  1. Hard Tanh:

  1. TanhExp:

  1. ABS:

  1. SquaredReLU:

  1. ParametricReLU (PReLU):

  1. RandomizedReLU (RReLU):

  1. LeakyReLU:

  1. ReLU6:

  1. ModReLU:

  1. CosReLU:

  1. SinReLU:

  1. Probit:

  1. Cosine:

  1. Gaussian:

  1. Multiquadratic:

    Choose some point (x,y).

  1. InvMultiquadratic:

  1. SoftPlus:

  1. Mish:

  1. Smish:

  1. ParametricSmish (PSmish):

  1. Swish:

  1. ESwish:

  1. Hard Swish:

  1. GCU:

  1. CoLU:

  1. PELU:

  1. SELU:

    where $\alpha \approx 1.6733$ & $\lambda \approx 1.0507$

  1. CELU:

  1. ArcTan:

  1. ShiftedSoftPlus:

  1. Softmax:

  1. Logit:

  1. GELU:

  1. Softsign:

  1. ELiSH:

  1. Hard ELiSH:

  1. Serf:

  1. ELU:

  1. Phish:

  1. QReLU:

  1. modified QReLU (m-QReLU):

  1. FReLU:

Cite this repository

@software{Pouya_ActTensor_2022,
author = {Pouya, Ardehkhani and Pegah, Ardehkhani},
license = {MIT},
month = {7},
title = {{ActTensor}},
url = {https://github.com/pouyaardehkhani/ActTensor},
version = {1.0.0},
year = {2022}
}
activation
activation-code
activation-function
activation-functions
activation-methods
activations
ai
artificial-intelligence
data-science
deep-learning
deeplearning
deep-neural-networks
neural-network-architectures
neural-networks
python
python3
python-library
python-package
tensorflow
tensorflow2

Significant stargazers

Martin Kubovčík

161 followers · starred May 2023