A simple and lightweight genetic algorithm for optimization of any machine learning model

Last update: Aug 10, 2022

Overview

geneticml

This package contains a simple and lightweight genetic algorithm for optimization of any machine learning model.

Installation

Use pip to install the package from PyPI:

pip install geneticml

Usage

This package provides a easy way to create estimators and perform the optimization with genetic algorithms. The example below describe in details how to create a simulation with genetic algorithms using evolutionary approach to train a sklearn.neural_network.MLPClassifier. A full list of examples could be found here.

from geneticml.optimizers import GeneticOptimizer
from geneticml.strategy import EvolutionaryStrategy
from geneticml.algorithms import EstimatorBuilder
from metrics import metric_accuracy
from sklearn.neural_network import MLPClassifier
from sklearn.datasets import load_iris

# Creates a custom fit method
def fit(model, x, y):
    return model.fit(x, y)

# Creates a custom predict method
def predict(model, x):
    return model.predict(x)

if __name__ == "__main__":

    seed = 11412

    # Creates an estimator
    estimator = EstimatorBuilder()\
        .of(model_type=MLPClassifier)\
        .fit_with(func=fit)\
        .predict_with(func=predict)\
        .build()

    # Defines a strategy for the optimization
    strategy = EvolutionaryStrategy(
        estimator_type=estimator,
        parameters=parameters,
        retain=0.4,
        random_select=0.1,
        mutate_chance=0.2,
        max_children=2,
        random_state=seed
    )

    # Creates the optimizer
    optimizer = GeneticOptimizer(strategy=strategy)

    # Loads the data
    data = load_iris()

    # Defines the metric
    metric = metric_accuracy
    greater_is_better = True

    # Create the simulation using the optimizer and the strategy
    models = optimizer.simulate(
        data=data.data, 
        target=data.target,
        generations=generations,
        population=population,
        evaluation_function=metric,
        greater_is_better=greater_is_better,
        verbose=True
    )

The estimator is the way you define an algorithm or a class that will be used for model instantiation

estimator = EstimatorBuilder().of(model_type=MLPClassifier).fit_with(func=fit).predict_with(func=predict).build()

You need to speficy a custom fit and predict functions. These functions need to use the same signature than the below ones. This happens because the algorithm is generic and needs to know how to perform the fit and predict functions for the models.

# Creates a custom fit method
def fit(model, x, y):
    return model.fit(x, y)

# Creates a custom predict method
def predict(model, x):
    return model.predict(x)

Custom strategy

You can create custom strategies for the optimizers by extending the geneticml.strategy.BaseStrategy and implementing the execute(...) function.

class MyCustomStrategy(BaseStrategy):
    def __init__(self, estimator_type: Type[BaseEstimator]) -> None:
        super().__init__(estimator_type)

    def execute(self, population: List[Type[T]]) -> List[T]:
        return population

The custom strategies will allow you to create optimization strategies to archive your goals. We currently have the evolutionary strategy but you can define your own :)

Custom optimizer

You can create custom optimizers by extending the geneticml.optimizers.BaseOptimizer and implementing the simulate(...) function.

class MyCustomOptimizer(BaseOptimizer):
    def __init__(self, strategy: Type[BaseStrategy]) -> None:
        super().__init__(strategy)

    def simulate(self, data, target, verbose: bool = True) -> List[T]:
        """
        Generate a network with the genetic algorithm.

        Parameters:
            data (?): The data used to train the algorithm
            target (?): The targets used to train the algorithm
            verbose (bool): True if should verbose or False if not

        Returns:
            (List[BaseEstimator]): A list with the final population sorted by their loss
        """
        estimators = self._strategy.create_population()
        for x in estimators:
            x.fit(data, target)
            y_pred = x.predict(target)
        pass

Custom optimizers will let you define how you want your algorithm to optimize the selected strategy. You can also combine custom strategies and optimizers to archive your desire objective.

Testing

The following are the steps to create a virtual environment into a folder named "venv" and install the requirements.

# Create virtualenv
python3 -m venv venv
# activate virtualenv
source venv/bin/activate
# update packages
pip install --upgrade pip setuptools wheel
# install requirements
python setup.py install

Tests can be run with python setup.py test when the virtualenv is active.

Contributing

All contributions, bug reports, bug fixes, documentation improvements, enhancements, and ideas are welcome.

A detailed overview on how to contribute can be found in the contributing guide. There is also an overview on GitHub.

If you are simply looking to start working with the geneticml codebase, navigate to the GitHub "issues" tab and start looking through interesting issues. Or maybe through using geneticml you have an idea of your own or are looking for something in the documentation and thinking ‘this can be improved’...you can do something about it!

Feel free to ask questions on the mailing the contributors.

Changelog

1.0.3 - Included pytorch example

1.0.2 - Minor fixes on naming

1.0.1 - README fixes

1.0.0 - First release

Comments

feature/data_sampling

We added support to run your own data sampling (e.g., imblearn.SMOTE) and use the genetic algorithms to find the best set parameters for them. Also, you can find the best set of parameters for your machine learning model at same time that find the best minority class size that maximizes the model score

opened by albarsil 0

Releases(1.0.8)

1.0.8(Mar 2, 2022)

Full Changelog: https://github.com/albarsil/geneticml/compare/1.0.7...1.0.8
Source code(tar.gz)
Source code(zip)
1.0.7(Mar 2, 2022)

Source code(tar.gz)
Source code(zip)
1.0.6(Feb 25, 2022)

Full Changelog: https://github.com/albarsil/geneticml/compare/1.0.5...1.0.6
Source code(tar.gz)
Source code(zip)
1.0.5(Feb 25, 2022)
What's Changed

feature/data_sampling by @albarsil in https://github.com/albarsil/geneticml/pull/5

New Contributors

@albarsil made their first contribution in https://github.com/albarsil/geneticml/pull/5

Full Changelog: https://github.com/albarsil/geneticml/compare/1.0.4...1.0.5
Source code(tar.gz)
Source code(zip)
1.0.4(Feb 18, 2022)

Full Changelog: https://github.com/albarsil/geneticml/commits/1.0.4
Source code(tar.gz)
Source code(zip)
1.0.3(Dec 7, 2021)

Included pytorch example
Source code(tar.gz)
Source code(zip)
1.0.2(Dec 7, 2021)

Full Changelog: https://github.com/albarsil/geneticml/compare/1.0.1...1.0.2
Source code(tar.gz)
Source code(zip)
1.0.1(Dec 7, 2021)

Full Changelog: https://github.com/albarsil/geneticml/commits/1.0.1
Source code(tar.gz)
Source code(zip)

H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.

H2O H2O is an in-memory platform for distributed, scalable machine learning. H2O uses familiar interfaces like R, Python, Scala, Java, JSON and the Fl

6.1k Jan 5, 2023

A very simple tool for situations where optimization with onnx-simplifier would exceed the Protocol Buffers upper file size limit of 2GB, or simply to separate onnx files to any size you want.

sne4onnx A very simple tool for situations where optimization with onnx-simplifier would exceed the Protocol Buffers upper file size limit of 2GB, or

10 Aug 30, 2022

library for nonlinear optimization, wrapping many algorithms for global and local, constrained or unconstrained, optimization

NLopt is a library for nonlinear local and global optimization, for functions with and without gradient information. It is designed as a simple, unifi

1.4k Dec 25, 2022

Ever felt tired after preprocessing the dataset, and not wanting to write any code further to train your model? Ever encountered a situation where you wanted to record the hyperparameters of the trained model and able to retrieve it afterward? Models Playground is here to help you do that. Models playground allows you to train your models right from the browser.

Models Playground 🗂️ Upload a Preprocessed Dataset 🌠 Choose whether to perform Classification or Regression 🦹 Enter the Dependent Variable ?

19 Dec 10, 2022

A simple and lightweight genetic algorithm for optimization of any machine learning model

Related tags

Overview

geneticml

Installation

Usage

Custom strategy

Custom optimizer

Testing

Contributing

Changelog

You might also like...

A very simple tool for situations where optimization with onnx-simplifier would exceed the Protocol Buffers upper file size limit of 2GB, or simply to separate onnx files to any size you want.

library for nonlinear optimization, wrapping many algorithms for global and local, constrained or unconstrained, optimization

A Lightweight Hyperparameter Optimization Tool 🚀

A Genetic Programming platform for Python with TensorFlow for wicked-fast CPU and GPU support.

Simulate genealogical trees and genomic sequence data using population genetic models

MBPO (paper: When to trust your model: Model-based policy optimization) in offline RL settings

RoMA: Robust Model Adaptation for Offline Model-based Optimization

Comments

feature/data_sampling

Releases(1.0.8)

1.0.8(Mar 2, 2022)

1.0.7(Mar 2, 2022)

1.0.6(Feb 25, 2022)

1.0.5(Feb 25, 2022)

What's Changed

New Contributors

1.0.4(Feb 18, 2022)

1.0.3(Dec 7, 2021)

1.0.2(Dec 7, 2021)

1.0.1(Dec 7, 2021)

Owner

Allan Barcelos

Continuous Query Decomposition for Complex Query Answering in Incomplete Knowledge Graphs

Libraries, tools and tasks created and used at DeepMind Robotics.

A non-linear, non-parametric Machine Learning method capable of modeling complex datasets

This is the latest version of the PULP SDK

Cmsc11 arcade - Final Project for CMSC11

Implementation of "Learning Multi-Granular Hypergraphs for Video-Based Person Re-Identification"

Rank 3 : Source code for OPPO 6G Data Generation Challenge

A simple baseline for 3d human pose estimation in PyTorch.

Benchmarks for Model-Based Optimization

A torch implementation of "Pixel-Level Domain Transfer"

A library to inspect itermediate layers of PyTorch models.

《Fst Lerning of Temporl Action Proposl vi Dense Boundry Genertor》(AAAI 2020)

catch-22: CAnonical Time-series CHaracteristics

Python codes for Lite Audio-Visual Speech Enhancement.

UltraGCN: An Ultra Simplification of Graph Convolutional Networks for Recommendation

Visual Question Answering in Pytorch

Official pytorch implementation of "DSPoint: Dual-scale Point Cloud Recognition with High-frequency Fusion"

code for paper "Not All Unlabeled Data are Equal: Learning to Weight Data in Semi-supervised Learning" by Zhongzheng Ren*, Raymond A. Yeh*, Alexander G. Schwing.

A Pytorch Implementation of Source Data-free Domain Adaptation for a Faster R-CNN

ML-Decoder: Scalable and Versatile Classification Head

code for paper "Not All Unlabeled Data are Equal: Learning to Weight Data in Semi-supervised Learning" by Zhongzheng Ren, Raymond A. Yeh, Alexander G. Schwing.