High performance Python GLMs with all the features!
See the codeGeneralized linear models (GLM) are a core statistical tool that include many common methods like least-squares regression, Poisson regression, and logistic regression as special cases. At QuantCo, we have used GLMs in e-commerce pricing, insurance claims prediction, and more. We have developed glum, a fast Python-first GLM library. The development was based on a fork of scikit-learn, so it has a scikit-learn-like API. We are thankful for the starting point provided by Christian Lorentzen in that PR!
We believe that for GLM development, broad support for distributions, regularization, and statistical inference, along with fast formula-based specification, is key. glum supports
formulaic, including monotonic constraintsnarwhalsPerformance also matters, so we conducted extensive benchmarks against other modern libraries. Although performance depends on the specific problem, we find that when N >> K (there are more observations than predictors), glum is consistently much faster for a wide range of problems. This repo includes the benchmarking tools in the glum_benchmarks module. For details, see here.
For more information on glum, including tutorials and API reference, please see the documentation.
Why did we choose the name glum? We wanted a name that had the letters GLM and wasn't easily confused with any existing implementation. And we thought glum sounded like a funny name (and not glum at all!). If you need a more professional-sounding name, feel free to pronounce it as G-L-um. Or maybe it stands for "Generalized linear... ummm... modeling?"
>>> import pandas as pd
>>> from glum import GeneralizedLinearRegressor
>>>
>>> # This dataset contains house sale prices for King County, which includes
>>> # Seattle. It includes homes sold between May 2014 and May 2015.
>>> # To download, use: sklearn.datasets.fetch_openml(name="house_sales", version=3)
>>> house_data = pd.read_parquet("data/housing.parquet")
>>>
>>> # Use only select features
>>> X = house_data[
... [
... "bedrooms",
... "bathrooms",
... "sqft_living",
... "floors",
... "waterfront",
... "view",
... "condition",
... "grade",
... "yr_built",
... "yr_renovated",
... ]
... ].copy()
>>>
>>> # Model whether a house had an above or below median price via a Binomial
>>> # distribution. We'll be doing L1-regularized logistic regression.
>>> price = house_data["price"]
>>> y = (price < price.median()).values.astype(int)
>>> model = GeneralizedLinearRegressor(
... family='binomial',
... l1_ratio=1.0,
... alpha=0.001
... )
>>>
>>> _ = model.fit(X=X, y=y)
>>>
>>> # Models can also be built with formulas from formulaic.
>>> # Monotonic constraints ensure bathrooms has a non-decreasing effect.
>>> model_formula = GeneralizedLinearRegressor(
... family='binomial',
... alpha=0.001,
... formula="bedrooms + bs(bathrooms, df=5) + bs(sqft_living, 3) + C(waterfront)",
... monotonic_constraints={"bathrooms": "increasing"}
... )
>>> _ = model_formula.fit(X=house_data, y=y)
Please install the package through conda-forge:
conda install glum -c conda-forge
For optimal performance on an x86_64 architecture, we recommend using the MKL library
(conda install mkl). By default, conda usually installs the openblas version, which
is slower, but supported on all major architectures and operating systems.
(top 30 of 32)
Python
95.9%
Cython
4.0%
High performance Python GLMs with all the features!
See the codeGeneralized linear models (GLM) are a core statistical tool that include many common methods like least-squares regression, Poisson regression, and logistic regression as special cases. At QuantCo, we have used GLMs in e-commerce pricing, insurance claims prediction, and more. We have developed glum, a fast Python-first GLM library. The development was based on a fork of scikit-learn, so it has a scikit-learn-like API. We are thankful for the starting point provided by Christian Lorentzen in that PR!
We believe that for GLM development, broad support for distributions, regularization, and statistical inference, along with fast formula-based specification, is key. glum supports
formulaic, including monotonic constraintsnarwhalsPerformance also matters, so we conducted extensive benchmarks against other modern libraries. Although performance depends on the specific problem, we find that when N >> K (there are more observations than predictors), glum is consistently much faster for a wide range of problems. This repo includes the benchmarking tools in the glum_benchmarks module. For details, see here.
For more information on glum, including tutorials and API reference, please see the documentation.
Why did we choose the name glum? We wanted a name that had the letters GLM and wasn't easily confused with any existing implementation. And we thought glum sounded like a funny name (and not glum at all!). If you need a more professional-sounding name, feel free to pronounce it as G-L-um. Or maybe it stands for "Generalized linear... ummm... modeling?"
>>> import pandas as pd
>>> from glum import GeneralizedLinearRegressor
>>>
>>> # This dataset contains house sale prices for King County, which includes
>>> # Seattle. It includes homes sold between May 2014 and May 2015.
>>> # To download, use: sklearn.datasets.fetch_openml(name="house_sales", version=3)
>>> house_data = pd.read_parquet("data/housing.parquet")
>>>
>>> # Use only select features
>>> X = house_data[
... [
... "bedrooms",
... "bathrooms",
... "sqft_living",
... "floors",
... "waterfront",
... "view",
... "condition",
... "grade",
... "yr_built",
... "yr_renovated",
... ]
... ].copy()
>>>
>>> # Model whether a house had an above or below median price via a Binomial
>>> # distribution. We'll be doing L1-regularized logistic regression.
>>> price = house_data["price"]
>>> y = (price < price.median()).values.astype(int)
>>> model = GeneralizedLinearRegressor(
... family='binomial',
... l1_ratio=1.0,
... alpha=0.001
... )
>>>
>>> _ = model.fit(X=X, y=y)
>>>
>>> # Models can also be built with formulas from formulaic.
>>> # Monotonic constraints ensure bathrooms has a non-decreasing effect.
>>> model_formula = GeneralizedLinearRegressor(
... family='binomial',
... alpha=0.001,
... formula="bedrooms + bs(bathrooms, df=5) + bs(sqft_living, 3) + C(waterfront)",
... monotonic_constraints={"bathrooms": "increasing"}
... )
>>> _ = model_formula.fit(X=house_data, y=y)
Please install the package through conda-forge:
conda install glum -c conda-forge
For optimal performance on an x86_64 architecture, we recommend using the MKL library
(conda install mkl). By default, conda usually installs the openblas version, which
is slower, but supported on all major architectures and operating systems.
(top 30 of 32)
Python
95.9%
Cython
4.0%