Resume where you left off
Online at https://diogoribeiro7.github.io/analytics-blog-jekyll/tutorials/2024/02/01/python-data-wrangling-foundations/
Topics
“Clean data is the foundation for every successful analysis.”
Why tidy data matters
- Consistent column naming and typing enables reproducible pipelines.
- Explicit missing data handling prevents silent downstream errors.
- Vectorized operations in pandas deliver fast, readable transformations.
Loading and inspecting data
1
2
3
4
5
6
7
8
import pandas as pd
def load_sales_data(path: str) -> pd.DataFrame:
"""Load the monthly sales CSV and parse dates."""
return pd.read_csv(path, parse_dates=["order_date"])
sales = load_sales_data("data/monthly_sales.csv")
print(sales.head())
The load_sales_data helper enforces date parsing and sets the tone for reusable
functions.
Cleaning column names
1
2
3
4
import janitor
sales = janitor.clean_names(sales)
sales.columns
pyjanitor.clean_names standardizes casing and spacing. Track these helpers in
utils/cleaning.py to share with teammates.
Handling missing values
- Use
DataFrame.info()to surface unexpected null columns. - Apply domain-driven imputations when appropriate.
- Preserve the original column when imputing to maintain auditability.
1
2
3
4
5
6
from sklearn.impute import SimpleImputer
import numpy as np
imputer = SimpleImputer(strategy="median")
sales["discount_filled"] = imputer.fit_transform(sales[["discount"]])
sales["discount_was_missing"] = np.where(sales["discount"].isna(), 1, 0)
Deriving tidy features
Create narrow columns that answer single analytical questions.
1
2
3
4
5
6
7
8
sales = (
sales.assign(
revenue=lambda df: df.quantity * df.unit_price,
order_month=lambda df: df.order_date.dt.to_period("M"),
)
.query("status == 'completed'")
.rename(columns={"customer_segment": "segment"})
)
Validating assumptions with tests
1
2
3
4
import pytest
def test_only_completed_orders():
assert sales.status.unique().tolist() == ["completed"]
Automated data tests catch regressions when upstream schemas drift.
Takeaways
- Encapsulate IO, cleaning, and feature engineering in composable functions.
- Version notebooks alongside unit tests to guard scientific integrity.
- Document design decisions inline so collaborators understand trade-offs.
Open the companion notebook through Binder to explore the exercises hands-on.
Your private highlights
Kept in this browser only and never sent anywhere. These are your notes, not comments. With text selected, Alt+Shift+H highlights it and Alt+Shift+N adds a note.
This browser is not letting the site keep data, so saving, progress and highlights are off.
Select a passage of the article to highlight it.
Your reading data
-
Explore the notebook in Binder
Launch the companion notebook to interactively run each transformation.
© 2024 Diogo Ribeiro. Text and figures under CC BY 4.0.
How to cite
Use the quick export buttons to save citations for reference managers or copy the formatted text directly.
Diogo Ribeiro (2024). Python Data Wrangling Foundations. DataLog | Data Science & Research Theme. https://diogoribeiro7.github.io/analytics-blog-jekyll/tutorials/2024/02/01/python-data-wrangling-foundations/.
BibTeX
RIS
EndNote
Open science & reproducibility badges
These badges highlight the transparency practices applied to this work. Hover or focus on each badge to learn more about the criteria.
- Open Data Dataset and code repository published with permissive license. Public repository, DOI issued, README with reproduction steps.
- Reproducible Workflow Containerized environment and automated tests provided. Continuous integration pipeline with reproducibility checks.
- Transparent Peer Review Peer review reports archived with DOI and linked to article. Open peer review statement and archived reports on Zenodo.
