Python Data Wrangling Foundations

“Clean data is the foundation for every successful analysis.”

Saved articles
Hides the site's navigation and the panels around the article; press Escape to leave.

Online at https://diogoribeiro7.github.io/analytics-blog-jekyll/tutorials/2024/02/01/python-data-wrangling-foundations/

Topics

“Clean data is the foundation for every successful analysis.”

Why tidy data matters

  • Consistent column naming and typing enables reproducible pipelines.
  • Explicit missing data handling prevents silent downstream errors.
  • Vectorized operations in pandas deliver fast, readable transformations.

Loading and inspecting data

1
2
3
4
5
6
7
8
import pandas as pd

def load_sales_data(path: str) -> pd.DataFrame:
    """Load the monthly sales CSV and parse dates."""
    return pd.read_csv(path, parse_dates=["order_date"])

sales = load_sales_data("data/monthly_sales.csv")
print(sales.head())

The load_sales_data helper enforces date parsing and sets the tone for reusable functions.

Cleaning column names

1
2
3
4
import janitor

sales = janitor.clean_names(sales)
sales.columns

pyjanitor.clean_names standardizes casing and spacing. Track these helpers in utils/cleaning.py to share with teammates.

Handling missing values

  1. Use DataFrame.info() to surface unexpected null columns.
  2. Apply domain-driven imputations when appropriate.
  3. Preserve the original column when imputing to maintain auditability.
1
2
3
4
5
6
from sklearn.impute import SimpleImputer
import numpy as np

imputer = SimpleImputer(strategy="median")
sales["discount_filled"] = imputer.fit_transform(sales[["discount"]])
sales["discount_was_missing"] = np.where(sales["discount"].isna(), 1, 0)

Deriving tidy features

Create narrow columns that answer single analytical questions.

1
2
3
4
5
6
7
8
sales = (
    sales.assign(
        revenue=lambda df: df.quantity * df.unit_price,
        order_month=lambda df: df.order_date.dt.to_period("M"),
    )
    .query("status == 'completed'")
    .rename(columns={"customer_segment": "segment"})
)

Validating assumptions with tests

1
2
3
4
import pytest

def test_only_completed_orders():
    assert sales.status.unique().tolist() == ["completed"]

Automated data tests catch regressions when upstream schemas drift.

Takeaways

  • Encapsulate IO, cleaning, and feature engineering in composable functions.
  • Version notebooks alongside unit tests to guard scientific integrity.
  • Document design decisions inline so collaborators understand trade-offs.

Open the companion notebook through Binder to explore the exercises hands-on.

Your private highlights

Kept in this browser only and never sent anywhere. These are your notes, not comments. With text selected, Alt+Shift+H highlights it and Alt+Shift+N adds a note.

Select a passage of the article to highlight it.

    Your reading data

    • Explore the notebook in Binder

      Launch the companion notebook to interactively run each transformation.

      Launch demo

    © 2024 Diogo Ribeiro. Text and figures under CC BY 4.0.

    How to cite

    Use the quick export buttons to save citations for reference managers or copy the formatted text directly.

    Diogo Ribeiro (2024). Python Data Wrangling Foundations. DataLog | Data Science & Research Theme. https://diogoribeiro7.github.io/analytics-blog-jekyll/tutorials/2024/02/01/python-data-wrangling-foundations/.

    BibTeX

    RIS

    EndNote

    Open science & reproducibility badges

    These badges highlight the transparency practices applied to this work. Hover or focus on each badge to learn more about the criteria.

    • Open Data Dataset and code repository published with permissive license. Public repository, DOI issued, README with reproduction steps.
    • Reproducible Workflow Containerized environment and automated tests provided. Continuous integration pipeline with reproducibility checks.
    • Transparent Peer Review Peer review reports archived with DOI and linked to article. Open peer review statement and archived reports on Zenodo.

    Related posts

    Loading mathematical content