A package for converting time series data from e.g. electronic health records into wide format data.

Project description

Timeseriesflattener

python versions

Time series from e.g. electronic health records often have a large number of variables, are sampled at irregular intervals and tend to have a large number of missing values. Before this type of data can be used for prediction modelling with machine learning methods such as logistic regression or XGBoost, the data needs to be reshaped.

In essence, the time series need to be flattened so that each prediction time is represented by a set of predictor values and an outcome value. These predictor values can be constructed by aggregating the preceding values in the time series within a certain time window.

timeseriesflattener aims to simplify this process by providing an easy-to-use and fully-specified pipeline for flattening complex time series.

🔧 Installation

To get started using timeseriesflattener simply install it using pip by running the following line in your terminal:

pip install timeseriesflattener

⚡ Quick start

import numpy as np
import pandas as pd

if __name__ == "__main__":

    # Load a dataframe with times you wish to make a prediction
    prediction_times_df = pd.DataFrame(
        {
            "id": [1, 1, 2],
            "date": ["2020-01-01", "2020-02-01", "2020-02-01"],
        },
    )
    # Load a dataframe with raw values you wish to aggregate as predictors
    predictor_df = pd.DataFrame(
        {
            "id": [1, 1, 1, 2],
            "date": [
                "2020-01-15",
                "2019-12-10",
                "2019-12-15",
                "2020-01-02",
            ],
            "value": [1, 2, 3, 4],
        },
    )
    # Load a dataframe specifying when the outcome occurs
    outcome_df = pd.DataFrame({"id": [1], "date": ["2020-03-01"], "value": [1]})

    # Specify how to aggregate the predictors and define the outcome
    from timeseriesflattener.feature_spec_objects import OutcomeSpec, PredictorSpec
    from timeseriesflattener.resolve_multiple_functions import maximum, mean

    predictor_spec = PredictorSpec(
        values_df=predictor_df,
        lookbehind_days=30,
        fallback=np.nan,
        entity_id_col_name="id",
        resolve_multiple_fn=mean,
        feature_name="test_feature",
    )
    outcome_spec = OutcomeSpec(
        values_df=outcome_df,
        lookahead_days=31,
        fallback=0,
        entity_id_col_name="id",
        resolve_multiple_fn=maximum,
        feature_name="test_outcome",
        incident=False,
    )

    # Instantiate TimeseriesFlattener and add the specifications
    from timeseriesflattener import TimeseriesFlattener

    ts_flattener = TimeseriesFlattener(
        prediction_times_df=prediction_times_df,
        entity_id_col_name="id",
        timestamp_col_name="date",
        n_workers=1,
        drop_pred_times_with_insufficient_look_distance=False,
    )
    ts_flattener.add_spec([predictor_spec, outcome_spec])
    df = ts_flattener.get_df()
    df

Output:

	id	date	prediction_time_uuid	pred_test_feature_within_30_days_mean_fallback_nan	outc_test_outcome_within_31_days_maximum_fallback_0_dichotomous
0	1	2020-01-01 00:00:00	1-2020-01-01-00-00-00	2.5	0
1	1	2020-02-01 00:00:00	1-2020-02-01-00-00-00	1	1
2	2	2020-02-01 00:00:00	2-2020-02-01-00-00-00	4	0

📖 Documentation

Documentation
🎓 Tutorial	Simple and advanced tutorials to get you started using `timeseriesflattener`
🎛 API References	The detailed reference for timeseriesflattener's API. Including function documentation
🙋 FAQ	Frequently asked question
🗺️ Roadmap	Kanban board for the roadmap for the project

💬 Where to ask questions

Type
🚨 Bug Reports	GitHub Issue Tracker
🎁 Feature Requests & Ideas	GitHub Issue Tracker
👩‍💻 Usage Questions	GitHub Discussions
🗯 General Discussion	GitHub Discussions

🎓 Projects

PSYCOP projects which use timeseriesflattener. Note that some of these projects have yet to be published and are thus private.

Project	Publications
Type 2 Diabetes		Prediction of type 2 diabetes among patients with visits to psychiatric hospital departments
Cancer		Prediction of Cancer among patients with visits to psychiatric hospital departments
COPD		Prediction of Chronic obstructive pulmonary disease (COPD) among patients with visits to psychiatric hospital departments
Forced admissions		Prediction of forced admissions of patients to the psychiatric hospital departments. Encompasses two seperate projects: 1. Prediciting at time of discharge for inpatient admissions. 2. Predicting day before outpatient admissions.
Coercion		Prediction of coercion among patients admittied to the hospital psychiatric department. Encompasses predicting mechanical restraint, sedative medication and manual restraint 48 hours before coercion occurs.

Project details

Release history Release notifications | RSS feed

2.4.0

Sep 27, 2024

2.3.0

Sep 27, 2024

2.2.6

May 23, 2024

2.2.5

May 22, 2024

2.2.4

May 17, 2024

2.2.3

May 7, 2024

2.2.2

May 3, 2024

2.2.1

May 2, 2024

2.2.0

Apr 30, 2024

2.1.2

Apr 18, 2024

2.1.1

Apr 18, 2024

2.1.0

Feb 27, 2024

2.0.2

Feb 27, 2024

2.0.1

Feb 26, 2024

2.0.0

Feb 26, 2024

1.36.2

Feb 23, 2024

1.36.1

Feb 22, 2024

1.36.0

Feb 22, 2024

1.35.0

Feb 22, 2024

1.34.0

Feb 22, 2024

1.33.0

Feb 22, 2024

1.32.0

Feb 22, 2024

1.31.3

Feb 22, 2024

1.31.2

Feb 19, 2024

1.31.1

Feb 19, 2024

1.31.0

Feb 19, 2024

1.30.0

Feb 19, 2024

1.29.0

Feb 19, 2024

1.28.0

Feb 19, 2024

1.27.0

Feb 16, 2024

1.26.0

Feb 16, 2024

1.25.1

Feb 16, 2024

1.25.0

Feb 16, 2024

1.24.0

Feb 15, 2024

1.23.0

Feb 14, 2024

1.22.0

Feb 14, 2024

1.21.1

Feb 14, 2024

1.21.0

Feb 14, 2024

1.20.1

Feb 13, 2024

1.20.0

Feb 13, 2024

1.19.0

Feb 13, 2024

1.18.1

Feb 13, 2024

1.18.0

Feb 13, 2024

1.17.0

Feb 13, 2024

1.16.0

Feb 12, 2024

1.15.0

Feb 12, 2024

1.14.0

Feb 12, 2024

1.13.0

Feb 12, 2024

1.12.0

Feb 12, 2024

1.11.0

Feb 9, 2024

1.10.0

Jan 25, 2024

1.9.1

Jan 23, 2024

1.9.0

Jan 18, 2024

1.8.0

Nov 24, 2023

1.7.0

Oct 20, 2023

1.6.1

Oct 5, 2023

1.6.0

Aug 9, 2023

1.5.2

Aug 2, 2023

1.5.1

Aug 1, 2023

1.5.0

Aug 1, 2023

1.4.0

Jul 12, 2023

1.3.1

Jun 30, 2023

1.3.0

Jun 29, 2023

1.2.1

Jun 20, 2023

1.2.0

Jun 20, 2023

1.0.0

Jun 15, 2023

0.27.0

May 19, 2023

0.26.0

May 4, 2023

0.25.1

May 2, 2023

0.25.0

Apr 26, 2023

0.24.0

Apr 20, 2023

0.23.11

Mar 28, 2023

0.23.10

Mar 28, 2023

0.23.9

Mar 28, 2023

0.23.8

Mar 28, 2023

0.23.7

Mar 24, 2023

0.23.6

Mar 20, 2023

0.23.5

Mar 20, 2023

0.23.4

Mar 20, 2023

0.23.3

Mar 10, 2023

0.23.2

Mar 1, 2023

0.23.1

Feb 24, 2023

0.23.0

Feb 9, 2023

This version

0.22.1

Dec 19, 2022

0.22.0

Dec 15, 2022

0.21.0

Dec 14, 2022

0.20.3

Dec 13, 2022

0.20.2

Dec 9, 2022

0.20.1

Dec 9, 2022

0.20.0

Dec 8, 2022

0.19.1

Dec 8, 2022

0.19.0

Dec 8, 2022

0.18.0

Dec 8, 2022

0.17.0

Dec 8, 2022

0.16.0

Dec 7, 2022

0.15.0

Dec 6, 2022

0.14.0

Dec 6, 2022

0.13.0

Dec 6, 2022

0.12.1

Dec 2, 2022

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

timeseriesflattener-0.22.1.tar.gz (31.0 kB view details)

Uploaded Dec 19, 2022 Source

Built Distribution

timeseriesflattener-0.22.1-py3-none-any.whl (31.9 kB view details)

Uploaded Dec 19, 2022 Python 3

File details

Details for the file timeseriesflattener-0.22.1.tar.gz.

File metadata

Download URL: timeseriesflattener-0.22.1.tar.gz
Upload date: Dec 19, 2022
Size: 31.0 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.8.0 pkginfo/1.9.2 readme-renderer/37.3 requests/2.28.1 requests-toolbelt/0.10.1 urllib3/1.26.13 tqdm/4.64.1 importlib-metadata/5.2.0 keyring/23.13.1 rfc3986/2.0.0 colorama/0.4.6 CPython/3.9.16

File hashes

Hashes for timeseriesflattener-0.22.1.tar.gz
Algorithm	Hash digest
SHA256	`95b1cad3d0230bd116727775293c95048579b87679aff307d6f0bedeb50411d5`
MD5	`34f1c908bcb925d54c529a11746e0d94`
BLAKE2b-256	`6d430fe84f6b795da7d5271699c224ce527374e528912c9dc4d1808649d6ae67`

See more details on using hashes here.

File details

Details for the file timeseriesflattener-0.22.1-py3-none-any.whl.

File metadata

Download URL: timeseriesflattener-0.22.1-py3-none-any.whl
Upload date: Dec 19, 2022
Size: 31.9 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.8.0 pkginfo/1.9.2 readme-renderer/37.3 requests/2.28.1 requests-toolbelt/0.10.1 urllib3/1.26.13 tqdm/4.64.1 importlib-metadata/5.2.0 keyring/23.13.1 rfc3986/2.0.0 colorama/0.4.6 CPython/3.9.16

File hashes

Hashes for timeseriesflattener-0.22.1-py3-none-any.whl
Algorithm	Hash digest
SHA256	`9ade8bc275ab994f53373f74fbe45a6c08b30f61764876201477e67f7ce5995d`
MD5	`6454462e0259378d9a54a31ef37dea7f`
BLAKE2b-256	`ac85ea4371bbe50ce94751836e8897f1d94e9434ef1b57a0197fbb694d09478e`