-
Notifications
You must be signed in to change notification settings - Fork 73
Commit
This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository.
- Loading branch information
1 parent
a6c3421
commit 7795389
Showing
70 changed files
with
5,424 additions
and
8,583 deletions.
There are no files selected for viewing
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
@@ -1,4 +1,4 @@ | ||
[run] | ||
branch = True | ||
source = skrebate | ||
include = */skrebate/* | ||
[run] | ||
branch = True | ||
source = skrebate | ||
include = */skrebate/* |
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
|
@@ -70,3 +70,8 @@ testing.ipynb | |
|
||
*.prof | ||
/demo_scikitrebate.ipynb | ||
|
||
*.DS_Store | ||
.idea/ | ||
|
||
analysis_pipeline/skrebatewip |
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
@@ -1,5 +1,5 @@ | ||
doc-warnings: yes | ||
|
||
ignore-patterns: | ||
- __init__.py | ||
|
||
doc-warnings: yes | ||
|
||
ignore-patterns: | ||
- __init__.py | ||
|
This file was deleted.
Oops, something went wrong.
This file was deleted.
Oops, something went wrong.
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
@@ -1,20 +1,18 @@ | ||
language: python | ||
python: | ||
- "3.5" | ||
virtualenv: | ||
system_site_packages: true | ||
env: | ||
matrix: | ||
# let's start simple: | ||
- PYTHON_VERSION="3.5" LATEST="true" | ||
- PYTHON_VERSION="3.5" COVERAGE="true" LATEST="true" | ||
- PYTHON_VERSION="3.5" LATEST="true" | ||
install: source ./ci/.travis_install.sh | ||
script: bash ./ci/.travis_test.sh | ||
after_success: | ||
# Ignore coveralls failures as the coveralls server is not very reliable | ||
# but we don't want travis to report a failure in the github UI just | ||
# because the coverage report failed to be published. | ||
- if [[ "$COVERAGE" == "true" ]]; then coveralls || echo "failed"; fi | ||
cache: apt | ||
sudo: false | ||
language: python | ||
virtualenv: | ||
system_site_packages: true | ||
env: | ||
matrix: | ||
# let's start simple: | ||
- PYTHON_VERSION="2.7" LATEST="true" | ||
- PYTHON_VERSION="3.6" COVERAGE="true" LATEST="true" | ||
- PYTHON_VERSION="3.6" LATEST="true" | ||
install: source ./ci/.travis_install.sh | ||
script: bash ./ci/.travis_test.sh | ||
after_success: | ||
# Ignore coveralls failures as the coveralls server is not very reliable | ||
# but we don't want travis to report a failure in the github UI just | ||
# because the coverage report failed to be published. | ||
- if [[ "$COVERAGE" == "true" ]]; then coveralls || echo "failed"; fi | ||
cache: apt | ||
sudo: false |
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
@@ -1,21 +1,21 @@ | ||
The MIT License (MIT) | ||
|
||
Copyright (c) 2016 Randal S. Olson and Ryan J. Urbanowicz | ||
|
||
Permission is hereby granted, free of charge, to any person obtaining a copy | ||
of this software and associated documentation files (the "Software"), to deal | ||
in the Software without restriction, including without limitation the rights | ||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell | ||
copies of the Software, and to permit persons to whom the Software is | ||
furnished to do so, subject to the following conditions: | ||
|
||
The above copyright notice and this permission notice shall be included in all | ||
copies or substantial portions of the Software. | ||
|
||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR | ||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, | ||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE | ||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER | ||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, | ||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE | ||
SOFTWARE. | ||
The MIT License (MIT) | ||
Copyright (c) 2016 Randal S. Olson and Ryan J. Urbanowicz | ||
Permission is hereby granted, free of charge, to any person obtaining a copy | ||
of this software and associated documentation files (the "Software"), to deal | ||
in the Software without restriction, including without limitation the rights | ||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell | ||
copies of the Software, and to permit persons to whom the Software is | ||
furnished to do so, subject to the following conditions: | ||
The above copyright notice and this permission notice shall be included in all | ||
copies or substantial portions of the Software. | ||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR | ||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, | ||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE | ||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER | ||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, | ||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE | ||
SOFTWARE. |
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
@@ -1,2 +1,2 @@ | ||
include LICENSE | ||
include README.md | ||
include LICENSE | ||
include README.md |
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
@@ -1,120 +1,120 @@ | ||
Master status: [![Master Build Status](https://travis-ci.org/EpistasisLab/scikit-rebate.svg?branch=master)](https://travis-ci.org/EpistasisLab/scikit-rebate) | ||
[![Master Code Health](https://landscape.io/github/EpistasisLab/scikit-rebate/master/landscape.svg?style=flat)](https://landscape.io/github/EpistasisLab/scikit-rebate/master) | ||
[![Master Coverage Status](https://coveralls.io/repos/github/EpistasisLab/scikit-rebate/badge.svg?branch=master&service=github)](https://coveralls.io/github/EpistasisLab/scikit-rebate?branch=master) | ||
|
||
Development status: [![Development Build Status](https://travis-ci.org/EpistasisLab/scikit-rebate.svg?branch=development)](https://travis-ci.org/EpistasisLab/scikit-rebate) | ||
[![Development Code Health](https://landscape.io/github/EpistasisLab/scikit-rebate/development/landscape.svg?style=flat)](https://landscape.io/github/EpistasisLab/scikit-rebate/development) | ||
[![Development Coverage Status](https://coveralls.io/repos/github/EpistasisLab/scikit-rebate/badge.svg?branch=development&service=github)](https://coveralls.io/github/EpistasisLab/scikit-rebate?branch=development) | ||
|
||
Package information: ![Python 2.7](https://img.shields.io/badge/python-2.7-blue.svg) | ||
![Python 3.5](https://img.shields.io/badge/python-3.6-blue.svg) | ||
![License](https://img.shields.io/badge/license-MIT%20License-blue.svg) | ||
[![PyPI version](https://badge.fury.io/py/skrebate.svg)](https://badge.fury.io/py/skrebate) | ||
|
||
# scikit-rebate | ||
This package includes a scikit-learn-compatible Python implementation of ReBATE, a suite of [Relief-based feature selection algorithms](https://en.wikipedia.org/wiki/Relief_(feature_selection)) for Machine Learning. These Relief-Based algorithms (RBAs) are designed for feature weighting/selection as part of a machine learning pipeline (supervised learning). Presently this includes the following core RBAs: ReliefF, SURF, SURF\*, MultiSURF\*, and MultiSURF. Additionally, an implementation of the iterative TuRF mechanism and VLSRelief is included. **It is still under active development** and we encourage you to check back on this repository regularly for updates. | ||
|
||
These algorithms offer a computationally efficient way to perform feature selection that is sensitive to feature interactions as well as simple univariate associations, unlike most currently available filter-based feature selection methods. The main benefit of Relief algorithms is that they identify feature interactions without having to exhaustively check every pairwise interaction, thus taking significantly less time than exhaustive pairwise search. | ||
|
||
Certain algorithms require user specified run parameters (e.g. ReliefF requires the user to specify some 'k' number of nearest neighbors). | ||
|
||
Relief algorithms are commonly applied to genetic analyses, where epistasis (i.e., feature interactions) is common. However, the algorithms implemented in this package can be applied to almost any supervised classification data set and supports: | ||
|
||
* Feature sets that are discrete/categorical, continuous-valued or a mix of both | ||
|
||
* Data with missing values | ||
|
||
* Binary endpoints (i.e., classification) | ||
|
||
* Multi-class endpoints (i.e., classification) | ||
|
||
* Continuous endpoints (i.e., regression) | ||
|
||
Built into this code, is a strategy to 'automatically' detect from the loaded data, these relevant characteristics. | ||
|
||
Of our two initial ReBATE software releases, this scikit-learn compatible version primarily focuses on ease of incorporation into a scikit learn analysis pipeline. | ||
This code is most appropriate for scikit-learn users, Windows operating system users, beginners, or those looking for the most recent ReBATE developments. | ||
|
||
An alternative 'stand-alone' version of [ReBATE](https://github.com/EpistasisLab/ReBATE) is also available that focuses on improving run-time with the use of Cython for optimization. This implementation also outputs feature names and associated feature scores as a text file by default. | ||
|
||
## License | ||
|
||
Please see the [repository license](https://github.com/EpistasisLab/scikit-rebate/blob/master/LICENSE) for the licensing and usage information for scikit-rebate. | ||
|
||
Generally, we have licensed scikit-rebate to make it as widely usable as possible. | ||
|
||
## Installation | ||
|
||
scikit-rebate is built on top of the following existing Python packages: | ||
|
||
* NumPy | ||
|
||
* SciPy | ||
|
||
* scikit-learn | ||
|
||
All of the necessary Python packages can be installed via the [Anaconda Python distribution](https://www.continuum.io/downloads), which we strongly recommend that you use. We also strongly recommend that you use Python 3 over Python 2 if you're given the choice. | ||
|
||
NumPy, SciPy, and scikit-learn can be installed in Anaconda via the command: | ||
|
||
``` | ||
conda install numpy scipy scikit-learn | ||
``` | ||
|
||
Once the prerequisites are installed, you should be able to install scikit-rebate with a pip command: | ||
|
||
``` | ||
pip install skrebate | ||
``` | ||
|
||
Please [file a new issue](https://github.com/EpistasisLab/scikit-rebate/issues/new) if you run into installation problems. | ||
|
||
## Usage | ||
|
||
We have designed the Relief algorithms to be integrated directly into scikit-learn machine learning workflows. For example, the ReliefF algorithm can be used as a feature selection step in a scikit-learn pipeline as follows. | ||
|
||
```python | ||
import pandas as pd | ||
import numpy as np | ||
from sklearn.pipeline import make_pipeline | ||
from skrebate import ReliefF | ||
from sklearn.ensemble import RandomForestClassifier | ||
from sklearn.model_selection import cross_val_score | ||
|
||
genetic_data = pd.read_csv('https://github.com/EpistasisLab/scikit-rebate/raw/master/data/' | ||
'GAMETES_Epistasis_2-Way_20atts_0.4H_EDM-1_1.tsv.gz', | ||
sep='\t', compression='gzip') | ||
|
||
features, labels = genetic_data.drop('class', axis=1).values, genetic_data['class'].values | ||
|
||
clf = make_pipeline(ReliefF(n_features_to_select=2, n_neighbors=100), | ||
RandomForestClassifier(n_estimators=100)) | ||
|
||
print(np.mean(cross_val_score(clf, features, labels))) | ||
>>> 0.795 | ||
``` | ||
|
||
For more information on the Relief algorithms available in this package and how to use them, please refer to our [usage documentation](https://EpistasisLab.github.io/scikit-rebate/using/). | ||
|
||
## Contributing to scikit-rebate | ||
|
||
We welcome you to [check the existing issues](https://github.com/EpistasisLab/scikit-rebate/issues/) for bugs or enhancements to work on. If you have an idea for an extension to scikit-rebate, please [file a new issue](https://github.com/EpistasisLab/scikit-rebate/issues/new) so we can discuss it. | ||
|
||
Please refer to our [contribution guidelines](https://EpistasisLab.github.io/scikit-rebate/contributing/) prior to working on a new feature or bug fix. | ||
|
||
## Citing scikit-rebate | ||
|
||
If you use scikit-rebate in a scientific publication, please consider citing the following paper: | ||
|
||
Ryan J. Urbanowicz, Randal S. Olson, Peter Schmitt, Melissa Meeker, Jason H. Moore (2017). [Benchmarking Relief-Based Feature Selection Methods](https://arxiv.org/abs/1711.08477). *arXiv preprint*, under review. | ||
|
||
BibTeX entry: | ||
|
||
```bibtex | ||
@misc{Urbanowicz2017Benchmarking, | ||
author = {Urbanowicz, Ryan J. and Olson, Randal S. and Schmitt, Peter and Meeker, Melissa and Moore, Jason H.}, | ||
title = {Benchmarking Relief-Based Feature Selection Methods}, | ||
year = {2017}, | ||
howpublished = {arXiv e-print. https://arxiv.org/abs/1711.08477}, | ||
} | ||
``` | ||
Master status: [![Master Build Status](https://travis-ci.org/EpistasisLab/scikit-rebate.svg?branch=master)](https://travis-ci.org/EpistasisLab/scikit-rebate) | ||
[![Master Code Health](https://landscape.io/github/EpistasisLab/scikit-rebate/master/landscape.svg?style=flat)](https://landscape.io/github/EpistasisLab/scikit-rebate/master) | ||
[![Master Coverage Status](https://coveralls.io/repos/github/EpistasisLab/scikit-rebate/badge.svg?branch=master&service=github)](https://coveralls.io/github/EpistasisLab/scikit-rebate?branch=master) | ||
|
||
Development status: [![Development Build Status](https://travis-ci.org/EpistasisLab/scikit-rebate.svg?branch=development)](https://travis-ci.org/EpistasisLab/scikit-rebate) | ||
[![Development Code Health](https://landscape.io/github/EpistasisLab/scikit-rebate/development/landscape.svg?style=flat)](https://landscape.io/github/EpistasisLab/scikit-rebate/development) | ||
[![Development Coverage Status](https://coveralls.io/repos/github/EpistasisLab/scikit-rebate/badge.svg?branch=development&service=github)](https://coveralls.io/github/EpistasisLab/scikit-rebate?branch=development) | ||
|
||
Package information: ![Python 2.7](https://img.shields.io/badge/python-2.7-blue.svg) | ||
![Python 3.5](https://img.shields.io/badge/python-3.6-blue.svg) | ||
![License](https://img.shields.io/badge/license-MIT%20License-blue.svg) | ||
[![PyPI version](https://badge.fury.io/py/skrebate.svg)](https://badge.fury.io/py/skrebate) | ||
|
||
# scikit-rebate | ||
This package includes a scikit-learn-compatible Python implementation of ReBATE, a suite of [Relief-based feature selection algorithms](https://en.wikipedia.org/wiki/Relief_(feature_selection)) for Machine Learning. These Relief-Based algorithms (RBAs) are designed for feature weighting/selection as part of a machine learning pipeline (supervised learning). Presently this includes the following core RBAs: ReliefF, SURF, SURF\*, MultiSURF\*, and MultiSURF. Additionally, an implementation of the iterative TuRF mechanism and VLSRelief is included. **It is still under active development** and we encourage you to check back on this repository regularly for updates. | ||
|
||
These algorithms offer a computationally efficient way to perform feature selection that is sensitive to feature interactions as well as simple univariate associations, unlike most currently available filter-based feature selection methods. The main benefit of Relief algorithms is that they identify feature interactions without having to exhaustively check every pairwise interaction, thus taking significantly less time than exhaustive pairwise search. | ||
|
||
Certain algorithms require user specified run parameters (e.g. ReliefF requires the user to specify some 'k' number of nearest neighbors). | ||
|
||
Relief algorithms are commonly applied to genetic analyses, where epistasis (i.e., feature interactions) is common. However, the algorithms implemented in this package can be applied to almost any supervised classification data set and supports: | ||
|
||
* Feature sets that are discrete/categorical, continuous-valued or a mix of both | ||
|
||
* Data with missing values | ||
|
||
* Binary endpoints (i.e., classification) | ||
|
||
* Multi-class endpoints (i.e., classification) | ||
|
||
* Continuous endpoints (i.e., regression) | ||
|
||
Built into this code, is a strategy to 'automatically' detect from the loaded data, these relevant characteristics. | ||
|
||
Of our two initial ReBATE software releases, this scikit-learn compatible version primarily focuses on ease of incorporation into a scikit learn analysis pipeline. | ||
This code is most appropriate for scikit-learn users, Windows operating system users, beginners, or those looking for the most recent ReBATE developments. | ||
|
||
An alternative 'stand-alone' version of [ReBATE](https://github.com/EpistasisLab/ReBATE) is also available that focuses on improving run-time with the use of Cython for optimization. This implementation also outputs feature names and associated feature scores as a text file by default. | ||
|
||
## License | ||
|
||
Please see the [repository license](https://github.com/EpistasisLab/scikit-rebate/blob/master/LICENSE) for the licensing and usage information for scikit-rebate. | ||
|
||
Generally, we have licensed scikit-rebate to make it as widely usable as possible. | ||
|
||
## Installation | ||
|
||
scikit-rebate is built on top of the following existing Python packages: | ||
|
||
* NumPy | ||
|
||
* SciPy | ||
|
||
* scikit-learn | ||
|
||
All of the necessary Python packages can be installed via the [Anaconda Python distribution](https://www.continuum.io/downloads), which we strongly recommend that you use. We also strongly recommend that you use Python 3 over Python 2 if you're given the choice. | ||
|
||
NumPy, SciPy, and scikit-learn can be installed in Anaconda via the command: | ||
|
||
``` | ||
conda install numpy scipy scikit-learn | ||
``` | ||
|
||
Once the prerequisites are installed, you should be able to install scikit-rebate with a pip command: | ||
|
||
``` | ||
pip install skrebate | ||
``` | ||
|
||
Please [file a new issue](https://github.com/EpistasisLab/scikit-rebate/issues/new) if you run into installation problems. | ||
|
||
## Usage | ||
|
||
We have designed the Relief algorithms to be integrated directly into scikit-learn machine learning workflows. For example, the ReliefF algorithm can be used as a feature selection step in a scikit-learn pipeline as follows. | ||
|
||
```python | ||
import pandas as pd | ||
import numpy as np | ||
from sklearn.pipeline import make_pipeline | ||
from skrebate import ReliefF | ||
from sklearn.ensemble import RandomForestClassifier | ||
from sklearn.model_selection import cross_val_score | ||
|
||
genetic_data = pd.read_csv('https://github.com/EpistasisLab/scikit-rebate/raw/master/data/' | ||
'GAMETES_Epistasis_2-Way_20atts_0.4H_EDM-1_1.tsv.gz', | ||
sep='\t', compression='gzip') | ||
|
||
features, labels = genetic_data.drop('class', axis=1).values, genetic_data['class'].values | ||
|
||
clf = make_pipeline(ReliefF(n_features_to_select=2, n_neighbors=100), | ||
RandomForestClassifier(n_estimators=100)) | ||
|
||
print(np.mean(cross_val_score(clf, features, labels))) | ||
>>> 0.795 | ||
``` | ||
|
||
For more information on the Relief algorithms available in this package and how to use them, please refer to our [usage documentation](https://EpistasisLab.github.io/scikit-rebate/using/). | ||
|
||
## Contributing to scikit-rebate | ||
|
||
We welcome you to [check the existing issues](https://github.com/EpistasisLab/scikit-rebate/issues/) for bugs or enhancements to work on. If you have an idea for an extension to scikit-rebate, please [file a new issue](https://github.com/EpistasisLab/scikit-rebate/issues/new) so we can discuss it. | ||
|
||
Please refer to our [contribution guidelines](https://EpistasisLab.github.io/scikit-rebate/contributing/) prior to working on a new feature or bug fix. | ||
|
||
## Citing scikit-rebate | ||
|
||
If you use scikit-rebate in a scientific publication, please consider citing the following paper: | ||
|
||
Ryan J. Urbanowicz, Randal S. Olson, Peter Schmitt, Melissa Meeker, Jason H. Moore (2017). [Benchmarking Relief-Based Feature Selection Methods](https://arxiv.org/abs/1711.08477). *arXiv preprint*, under review. | ||
|
||
BibTeX entry: | ||
|
||
```bibtex | ||
@misc{Urbanowicz2017Benchmarking, | ||
author = {Urbanowicz, Ryan J. and Olson, Randal S. and Schmitt, Peter and Meeker, Melissa and Moore, Jason H.}, | ||
title = {Benchmarking Relief-Based Feature Selection Methods}, | ||
year = {2017}, | ||
howpublished = {arXiv e-print. https://arxiv.org/abs/1711.08477}, | ||
} | ||
``` |
Binary file not shown.
Oops, something went wrong.