Cheap and reliable Node.js hosting starts at $3/month, and $1/month static HTML hosting

Created with love in Canada, visit hostnodejs.com today

Feel like to post an Ad? Learn Details

All Projects → scikit-learn-contrib → Stability Selection

scikit-learn-contrib / Stability Selection

Licence: bsd-3-clause

scikit-learn compatible implementation of stability selection.

Programming Languages

139335 projects - #7 most used programming language

Labels

machine-learning statistics scikit-learn

Projects that are alternatives of or similar to Stability Selection

Interpretable ML package 🔍 for concise, transparent, and accurate predictive modeling (sklearn-compatible).

Stars: ✭ 194 (+25.16%)

Mutual labels: statistics, scikit-learn

Machine Learning With Python

Practice and tutorial-style notebooks covering wide variety of machine learning techniques

Stars: ✭ 2,197 (+1317.42%)

Mutual labels: statistics, scikit-learn

Data Science Best Resources

Carefully curated resource links for data science in one place

Stars: ✭ 1,104 (+612.26%)

Mutual labels: statistics, scikit-learn

Virgilio is developed and maintained by these awesome people. You can email us virgilio.datascience (at) gmail.com or join the Discord chat.

Stars: ✭ 13,200 (+8416.13%)

Mutual labels: statistics, scikit-learn

Awesome Python Data Science

Probably the best curated list of data science software in Python.

Stars: ✭ 812 (+423.87%)

Mutual labels: statistics, scikit-learn

50% faster, 50% less RAM Machine Learning. Numba rewritten Sklearn. SVD, NNMF, PCA, LinearReg, RidgeReg, Randomized, Truncated SVD/PCA, CSR Matrices all 50+% faster

Stars: ✭ 1,204 (+676.77%)

Mutual labels: statistics, scikit-learn

Interactive machine learning

IPython widgets, interactive plots, interactive machine learning

Stars: ✭ 140 (-9.68%)

Mutual labels: statistics, scikit-learn

Analytics for CSS

Stars: ✭ 146 (-5.81%)

Mutual labels: statistics

Spacexstats React

SpaceXStats, with React and using the SpaceX API

Stars: ✭ 148 (-4.52%)

Mutual labels: statistics

Python Machine Learning Book

The "Python Machine Learning (1st edition)" book code repository and info resource

Stars: ✭ 11,428 (+7272.9%)

Mutual labels: scikit-learn

A library for machine learning research on motion capture data

Stars: ✭ 150 (-3.23%)

Mutual labels: scikit-learn

Uncertainty Metrics

An easy-to-use interface for measuring uncertainty and robustness.

Stars: ✭ 145 (-6.45%)

Mutual labels: statistics

Uncommons Maths

Random number generators, probability distributions, combinatorics and statistics for Java.

Stars: ✭ 149 (-3.87%)

Mutual labels: statistics

Transform ML models into a native code (Java, C, Python, Go, JavaScript, Visual Basic, C#, R, PowerShell, PHP, Dart, Haskell, Ruby, F#, Rust) with zero dependencies

Stars: ✭ 1,962 (+1165.81%)

Mutual labels: scikit-learn

🛠 All-in-one web-based IDE specialized for machine learning and data science.

Stars: ✭ 2,337 (+1407.74%)

Mutual labels: scikit-learn

An implementation of Caruana et al's Ensemble Selection algorithm in Python, based on scikit-learn

Stars: ✭ 145 (-6.45%)

Mutual labels: scikit-learn

Python solutions to the problem sets of Stanford's graduate course on Machine Learning, taught by Prof. Andrew Ng

Stars: ✭ 151 (-2.58%)

Mutual labels: statistics

Pg stat monitor

PostgreSQL Statistics Collector

Stars: ✭ 145 (-6.45%)

Mutual labels: statistics

Contributions Importer For Github

This tool helps users to import contributions to GitHub from private git repositories, or from public repositories that are not hosted in GitHub.

Stars: ✭ 147 (-5.16%)

Mutual labels: statistics

Hands On Machine Learning With Scikit Learn Keras And Tensorflow

Notes & exercise solutions of Part I from the book: "Hands-On ML with Scikit-Learn, Keras & TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems" by Aurelien Geron

Stars: ✭ 151 (-2.58%)

Mutual labels: scikit-learn

View All Similar Projects ➔

stability-selection - A scikit-learn compatible implementation of stability selection

stability-selection is a Python implementation of the stability selection feature selection algorithm, first proposed by Meinshausen and Buhlmann.

The idea behind stability selection is to inject more noise into the original problem by generating bootstrap samples of the data, and to use a base feature selection algorithm (like the LASSO) to find out which features are important in every sampled version of the data. The results on each bootstrap sample are then aggregated to compute a stability score for each feature in the data. Features can then be selected by choosing an appropriate threshold for the stability scores.

Installation

To install the module, clone the repository

git clone https://github.com/scikit-learn-contrib/stability-selection.git

Before installing the module you will need numpy, matplotlib, and sklearn. Install these modules separately, or install using the requirements.txt file:

pip install -r requirements.txt

and execute the following in the project directory to install stability-selection:

python setup.py install

Documentation and algorithmic details

See the documentation for details on the module, and the accompanying blog post for details on the algorithmic details.

Example usage

stability-selection implements a class StabilitySelection, that takes any scikit-learn compatible estimator that has either a feature_importances_ or coef_ attribute after fitting. Important other parameters are

lambda_name: the name of the penalization parameter of the base estimator (for example, C in the case of LogisticRegression).
lambda_grid: an array of values of the penalization parameter to iterate over.

After instantiation, the algorithm can be run with the familiar fit and transform calls.

Basic example

See below for an example:

import numpy as np

from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.utils import check_random_state
from stability_selection import StabilitySelection


def _generate_dummy_classification_data(p=1000, n=1000, k=5, random_state=123321):

    rng = check_random_state(random_state)

    X = rng.normal(loc=0.0, scale=1.0, size=(n, p))
    betas = np.zeros(p)
    important_betas = np.sort(rng.choice(a=np.arange(p), size=k))
    betas[important_betas] = rng.uniform(size=k)

    probs = 1 / (1 + np.exp(-1 * np.matmul(X, betas)))
    y = (probs > 0.5).astype(int)

    return X, y, important_betas

## This is all preparation of the dummy data set
n, p, k = 500, 1000, 5

X, y, important_betas = _generate_dummy_classification_data(n=n, k=k)
base_estimator = Pipeline([
    ('scaler', StandardScaler()),
    ('model', LogisticRegression(penalty='l1'))
])

## Here stability selection is instantiated and run
selector = StabilitySelection(base_estimator=base_estimator, lambda_name='model__C',
                              lambda_grid=np.logspace(-5, -1, 50)).fit(X, y)

print(selector.get_support(indices=True))

Bootstrapping strategies

stability-selection uses bootstrapping without replacement by default (as proposed in the original paper), but does support different bootstrapping strategies. [Shah and Samworth] proposed complementary pairs bootstrapping, where the data set is bootstrapped in pairs, such that the intersection is empty but the union equals the original data set. StabilitySelection supports this through the bootstrap_func parameter.

This parameter can be:

A string, which must be one of
- 'subsample': For subsampling without replacement (default).
- 'complementary_pairs': For complementary pairs subsampling [2].
- 'stratified': For stratified bootstrapping in imbalanced classification.
A function that takes y, and a random state as inputs and returns a list of sample indices in the range (0, len(y)-1).

For example, the StabilitySelection call in the above example can be replaced with

selector = StabilitySelection(base_estimator=base_estimator,
                              lambda_name='model__C',
                              lambda_grid=np.logspace(-5, -1, 50),
                              bootstrap_func='complementary_pairs')
selector.fit(X, y)

to run stability selection with complementary pairs bootstrapping.

Feedback and contributing

Feedback and contributions are much appreciated. If you have any feedback, please post it on the issue tracker.

References

[1]: Meinshausen, N. and Buhlmann, P., 2010. Stability selection. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 72(4), pp.417-473.

[2] Shah, R.D. and Samworth, R.J., 2013. Variable selection with error control: another look at stability selection. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 75(1), pp.55-80.

Note that the project description data, including the texts, logos, images, and/or trademarks, for each open source project belongs to its rightful owner. If you wish to add or remove any projects, please contact us at [email protected].

Stars: ✭ 155

Visit Git Page 🔗Visit User Page 🔗Visit Issues Page (15) 🔗