Resume where you left off
Online at https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/02/python-pandas-feature-engineering/
Topics
Time-series feature engineering in pandas turns raw telemetry into actionable signals. We will compute rolling statistics, annotate seasonal trends, and document the math so collaborators can audit each transformation.
Prerequisites
- Python 3.10+
- pandas 2.2+
- A dataset with timestamped metrics (e.g., request latency in milliseconds)
Tip: Store raw data in Parquet to preserve schema metadata and improve I/O.
Rolling aggregations
1
2
3
4
5
6
7
8
9
10
11
12
import pandas as pd
latency = pd.read_parquet("data/api-latency.parquet").set_index("timestamp")
window = latency.rolling("7D", min_periods=3)
features = pd.DataFrame(
{
"latency_p50": window["p50"].median(),
"latency_p95": window["p95"].max(),
"error_rate": window["errors"].sum() / window["requests"].sum(),
}
)
features = features.dropna()
Inline math keeps the feature derivation transparent. The error rate feature is simply $\hat{p} = \frac{\text{errors}}{\text{requests}}$ evaluated over the rolling window.
Seasonal decomposition
Use statsmodels to separate long-term trend from seasonality:
1
2
3
4
5
6
from statsmodels.tsa.seasonal import STL
stl = STL(features["latency_p95"], period=14)
result = stl.fit()
features["latency_trend"] = result.trend
features["latency_seasonal"] = result.seasonal
Documenting each derived column in the post ensures the DataLog theme renders callouts, syntax highlighting, and inline equations without extra plugins.
Share the output
- Publish the generated CSV in
_datasets/with provenance metadata. - Attach the notebook run to
_notebooks/for reviewers to replay calculations. - Link to dashboards or alerts that consume the engineered features.
By combining equations with syntax-highlighted code snippets, the article remains reproducible while showcasing how the theme renders mathematical context alongside pandas transformations.
Your private highlights
Kept in this browser only and never sent anywhere. These are your notes, not comments. With text selected, Alt+Shift+H highlights it and Alt+Shift+N adds a note.
This browser is not letting the site keep data, so saving, progress and highlights are off.
Select a passage of the article to highlight it.
Your reading data
Embed interactive plots, widgets, and demos using <figure>, <iframe>, or <div class="interactive-embed"> containers. Ensure each embed includes descriptive captions for accessibility.
© 2024 Diogo Ribeiro. Text and figures under CC BY 4.0.
How to cite
Use the quick export buttons to save citations for reference managers or copy the formatted text directly.
Diogo Ribeiro (2024). Feature Engineering with pandas Window Functions. DataLog | Data Science & Research Theme. https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/02/python-pandas-feature-engineering/.
BibTeX
RIS
EndNote
Open science & reproducibility badges
These badges highlight the transparency practices applied to this work. Hover or focus on each badge to learn more about the criteria.
- Open Data Dataset and code repository published with permissive license. Public repository, DOI issued, README with reproduction steps.
- Reproducible Workflow Containerized environment and automated tests provided. Continuous integration pipeline with reproducibility checks.
- Transparent Peer Review Peer review reports archived with DOI and linked to article. Open peer review statement and archived reports on Zenodo.