SirojFlow is a lightweight educational deep learning framework that implements neural networks from first principles. Every stage of training, from forward propagation to backpropagation, is written manually using NumPy.
- Fully Connected (
Linear) - Sequential model container (
Sequential) - Activation Functions
- ReLU
- LeakyReLU
- Sigmoid
- Tanh
- Softmax
- Swish
- HeavySide
- MSELoss
- MAELoss
- BinaryCrossentropyLoss
- SparseCategoricalCrossentropyLoss
- SGD
- Adam
DataLoader- Batch creation
- Dataset shuffling
- Optional dropping of incomplete batches
- He
- Xavier
- Zero
- Random (Default)
pip install sirojflowimport numpy as np
from SirojFlow.engine.nn import Sequential, Linear
from SirojFlow.engine.act import ReLU, Softmax
from SirojFlow.losses import SparseCategoricalCrossentropyLoss
from SirojFlow.optims import Adam
from SirojFlow.utils import DataLoader
x = np.random.randn(500, 20)
y = np.random.randint(5, size=500)
loader = DataLoader(x, y, batch_size=32, shuffle=True)
model = Sequential(
Linear(20, 64),
ReLU(),
Linear(64, 5),
Softmax()
)
criterion = SparseCategoricalCrossentropyLoss(model.layers[-1])
optimizer = Adam(model, lr=1e-3)
for epoch in range(20):
total_loss = 0
for xb, yb in loader:
pred = model(xb)
loss = criterion(pred, yb)
criterion.backward()
optimizer.step()
total_loss += loss
print(f"Epoch {epoch+1}: {total_loss:.4f}")print(model.summary())src/
└── SirojFlow/
│ ├── engine/
│ │ ├── _sirojflow.py
│ │ ├── act.py
│ │ └── nn.py
│ ├── losses.py
│ ├── optims.py
│ ├── utils.py
│ └── LICENSE
└──README.md
Forward Pass
Input
│
Model
│
Prediction
│
Loss
Backward Pass
Loss
│
criterion.backward()
│
optimizer.step()
│
*layers.backward()
|
Parameter Update
SirojFlow follows a modular object-oriented design.
- Layers perform forward propagation and gradient computation.
- Losses compute the initial gradient i.e. of gradient of loss function with respect to the output of the last layer.
- Optimizers drive the complete backpropagation process, and update parameters.
No automatic differentiation or computational graph is used.
The syntax in training process is same for all cases
prediction = model(x_batch)
loss = criterion(prediction, y_batch)
criterion.backward()
optimizer.step()-
First, import
DataLoaderobject, it is located assirojflow.utils. Then, load your data asloaded_data = DataLoader(x: numpy.ndarray, y: numpy.ndarray, batch_size=<int>, shuffle=<bool>, drop_last=<bool>)here,batch_size; if None: full-batch, if value <int> given, mini-batch shuffle; if True: shuffles batches every epoch else: doesn't shuffle batches drop_last; if True: drops incomplete batch else: doesn't drop incomplete batch -
Import the
Sequentiallayer container fromsirojflow.engine.nn -
Now, import the
Linearlayer object fromsirojflow.engine.nnand essential activations fromsirojflow.engine.nn.act -
You can define your model in two ways:
model = Sequential( <*layers> )
or
model = Sequential() model.add(layer1) model.add(layer2) model.add(layer3) ... ...
Note: The
Linearlayer requires two parameterin_featuresandout_featuresbecause shape of linear layer is (in_features x out_features). You can even specify weight initialization for each linear layer. Shape of weight: (out_features x in_feature). -
Now, import required loss from
sirojflow.lossesand required optimizer fromsirojflow.optims. -
To setup optimizer, you should pass entire model
modelinto it, and you can put enter the learning ratelr:optim = <Optimizer>(model: Sequential, lr=1e-3). You can also addmomentin SGD. Similarly you can experiment with moment parametersbetaandgammain Adam. Nevertheless, you can also use L2 regularization (penalty) by addingweight_decayin any optimizer. -
To setup the criterion function, you should pass the last layer of model
model.layers[-1]to define the criterion;criterion = <Loss>(model.layers[-1])and to find the loss, you can do:loss = criterion(pred, label) -
Now, you can iterate through epochs and
loaded_datalike:for epoch in range(epochs): for x_batch, y_batch in loaded_data: ... ...
and use the same training flow syntax mentioned above
model(x_batch)does forward propagation for each layers, caches and intermediate values inside each layer objectcriterion(prediction, y_batch)returns the loss valuecritetion.backward()calls the.backward()of the loss function, this calculates gradient of loss with respect to output of modelprediction, this gets cached asgrad_nextof the last layer of the model, which is passed into the loss object during criterion definitioncriterion = <Loss>(model.layers[-1])- Optimizer does backward pass to every layer of model from last layer, from what it caches the chained gradient upto previous layer as
grad_nextin the current layer by calling each layer's.backward(next_grad) - Then the optimizer filters updatable layer
Linearwhich has parametersdWanddB, fetches those and updates the parameters of all linear layers.
- Python 3.7+
- NumPy
MIT License.