Posts by Year

Every article, filterable by year, category and tag.

595 Total Posts
595 Showing
10 Years
13 Categories

2026 (142 posts)

January

February

March

April

May

June

July

  • Cost-Sensitive Learning for Rare Event Prediction

    Rare event models should be optimized for decisions, not only class balance. Cost-sensitive learning connects model thresholds to real operational ...

  • Counterfactual Evaluation for Decision Policies

    Counterfactual evaluation helps teams estimate how a new decision policy might perform before deploying it to users, patients, customers, or operat...

  • Anomaly Detection in Sensor Streams

    Sensor anomaly detection works best when statistical signals, domain constraints, and alert workflows are designed together.

  • Prevalence Shift and Base-Rate Drift in Machine Learning

    Prevalence shift occurs when the base rate of the outcome changes, breaking thresholds, workloads, and probability interpretation even when the mod...

  • Uplift Modeling for Targeted Interventions

    Uplift modeling estimates treatment effect heterogeneity so interventions can target the people, assets, or cases most likely to benefit.

  • Forecast Combination: Why Averaging Usually Wins

    Choosing the best model is the obvious strategy. Averaging several is usually better, and the reason is not that the average is smarter but that it...

  • Regime-Switching Models for Time Series

    A single model fitted across a recession and an expansion describes neither. Regime-switching models allow the dynamics themselves to change, with ...

  • The Largest Eigenvalue of Noise Is Not Evidence of a Factor

    When the number of variables is not small relative to the sample size, sample covariance eigenvalues spread dramatically even if the true covarianc...

  • Nowcasting with Mixed-Frequency Data

    The quantity you care about arrives quarterly and two months late. Related indicators arrive daily. Nowcasting is the problem of estimating the pre...

  • Interrupted Time Series and Causal Impact

    A intervention happened at a known date and you need its effect. There is no control group, only the series itself before and after, and the counte...

  • Forecast Value Added: Is Your Process Helping?

    Forecasting processes accumulate steps: a statistical model, a planner override, a consensus meeting. Each is assumed to improve the number. FVA is...

  • LLM Quantization Is Not Just Using Fewer Bits

    Saying that a model is 4-bit tells you how weights are represented, not how they were calibrated, which tensors were quantized, what arithmetic the...

  • Modelling Count Time Series

    Daily incident counts are integers, non-negative, often small, and correlated with yesterday. ARIMA assumes none of that and Poisson regression ass...

  • Long Memory and Fractional Integration in Time Series

    Standard practice offers two options: the series is stationary, or you difference it. Some series are genuinely in between, and forcing them either...

  • Dynamic Time Warping and Time Series Clustering

    Two series can trace an identical shape while one runs slightly ahead of the other. Point-by-point distance calls them dissimilar; dynamic time war...

August

September

October

  • Economic Data Have Two Dates

    A GDP value describes one quarter but may become known months later and change again after that. Financial and economic analyses need to preserve b...

2025 (107 posts)

January

  • Nonlinear Growth Models in Macroeconomics

    Nonlinear growth models offer a richer and more realistic framework for understanding macroeconomic development over time. This article explores th...

  • Differential Equations in Growth Models

    Differential equations are essential in modeling economic growth, providing insight into long-term trends and the impact of policy changes on macro...

  • Service Level Is Not One Metric

    "Service level" sounds like one number. In inventory and logistics it is a family of different probabilities and ratios, each weighting shortages, ...

  • Bayesian State Space Models in Macroeconometrics

    Explore the critical role of Bayesian state space models in macroeconometric analysis, with a focus on linear Gaussian models, dimension reduction,...

February

March

April

May

June

  • Why SMOTE Isn't Always the Answer

    SMOTE generates synthetic samples to rebalance datasets, but using it blindly can create unrealistic data and biased models.

  • Entropy Minimisation Can Make the Wrong Answer More Confident

    Entropy minimisation encourages decisive predictions on unlabelled data. That can be useful when decision boundaries should avoid high-density regi...

  • Model Deployment: Best Practices and Tips

    Deploying machine learning models to production requires planning and robust infrastructure. Here are key practices to ensure success.

  • Hyperparameter Tuning Strategies

    Hyperparameter tuning can drastically improve model performance. Explore common search strategies and tools.

  • A Gentle Introduction to Neural Networks

    Neural networks power many modern AI applications. This article introduces their basic structure and training process.

  • Crafting Time Series Features for Better Models

    Learn specialized feature engineering techniques to make time series data more predictive for machine learning models.

  • Why Data Scientists Need Math and Statistics

    Mastering mathematics and statistics is essential for understanding data science algorithms and avoiding common pitfalls when building models.

  • Exploratory Data Analysis: A Beginner's Guide

    Discover the essential steps of Exploratory Data Analysis (EDA) and how to gain insights from your data before building models.

  • Replication Is More Than Getting the Same p-Value Twice

    Repeating p < 0.05 is a poor definition of replication. Under modest power, an exact repeat of a real effect may often fail to cross the same thres...

  • Least Angle Regression: A Gentle Dive into LARS

    Least Angle Regression, or LARS, is an efficient regression algorithm designed for high-dimensional data. It provides a pathwise approach to linear...

July

August

September

October

November

December

2024 (201 posts)

January

  • Probabilistic Programming and MCMC

    Probabilistic programming separates model specification from inference, but inference still depends on diagnostics, geometry, and numerical stability.

  • The Normal Distribution: Why the Bell Curve Appears

    The normal distribution is important because of Gaussian models, additive noise, asymptotic approximations, and the central limit theorem—not becau...

  • Marina Viazovska and the E8 Sphere-Packing Proof

    Marina Viazovska solved the eight-dimensional sphere-packing problem by constructing the exact auxiliary function needed to prove optimality of the...

February

  • Climate Financial Risk Beyond Traditional VaR

    Climate financial risk is better treated as scenario-conditioned loss analysis than as a simple extension of short-horizon market VaR.

  • Cold Days Still Belong in a Warming Climate

    A warmer climate can still produce a freezing morning. The question is how the range and frequency of temperatures change, rather than whether cold...

  • Paths of Combinatorics and Probability

    Dive into the intersection of combinatorics and probability, exploring how these fields work together to solve problems in mathematics, data scienc...

  • Ethics of AI and Sensing in Older-Adult Care

    Ethical technology for older adults requires consent, privacy, proportional monitoring, accessibility, contestability, and attention to decision-ma...

  • Ergodicity: Time Averages, Invariant Sets, and Mixing

    Ergodicity is a precise property of a measure-preserving dynamical system. It is not merely a temporary regime, nor is chaos sufficient for ergodic...

  • Spectral Clustering: The Graph Is the Model

    Spectral clustering converts a similarity graph into an eigenvector embedding and then clusters that embedding. The graph construction is the model.

  • Clustering Is a Model of Similarity

    Clustering does not discover a unique hidden partition of data. It produces groups relative to a representation, similarity measure, algorithm, and...

  • Topological Data Analysis: Shape Across Scales

    Topological data analysis summarizes multiscale shape through constructions such as persistent homology and Mapper, but topology does not make high...

  • Customer Lifetime Value: Expected Future Contribution

    Customer lifetime value is an expected discounted future contribution under assumptions about activity, retention, margin, censoring, and intervent...

March

  • Forecast Accuracy Is Not Inventory Performance

    Forecasting metrics evaluate predictions. Supply chains pay for decisions. A model can achieve a lower RMSE and still produce higher stockout and i...

  • A Technical History of Artificial Intelligence

    The history of artificial intelligence is often told as a straight line from ancient automata to modern language models. That makes a good story bu...

May

June

July

August

September

October

November

December

2023 (38 posts)

January

  • Walking the Mathematical Path

    Dive into the fascinating world of pedestrian behavior through mathematical models like the Social Force Model. Learn how these models inform urban...

  • Error Terms in Linear and Logistic Regression

    Linear and logistic regression encode randomness differently. The distinction is between an additive disturbance model and a Bernoulli conditional ...

February

  • Advanced Statistical Methods for Efficient A/B Testing

    An in-depth exploration of sequential testing and its application in A/B testing. Understand the statistical underpinnings, advantages, limitations...

March

  • Chi-Square Test: Testing Categorical Data

    The Chi-Square Test is a powerful tool for analyzing relationships in categorical data. Learn its principles and practical applications.

May

July

August

  • Ethics in Data Science

    Ethics in data science is a question of governance, measurement, rights, incentives, and accountability, not a checklist added after a model is built.

  • Applying R Functions on Rolling Windows with runner

    A practical guide to rolling computations in R with runner, including fixed-size and time-indexed windows, lags, custom evaluation points, grouped ...

  • The Life and Mathematics of Paul Erdős

    Paul Erdős transformed twentieth-century combinatorics, number theory, graph theory, and probabilistic mathematics through an unusually collaborati...

  • Demystifying Data Science

    Data science is not a catalogue of algorithms. It is a discipline for turning imperfect observations into defensible descriptions, predictions, and...

  • Gaussian Processes for Time-Series Analysis in Python

    Dive into Gaussian Processes for time-series analysis using Python, combining flexible modeling with Bayesian inference for trends, seasonality, an...

September

  • Sample Size: Power, Precision, and Design

    Sample size should be derived from the estimand, design, effect size, uncertainty target, and error rates, not from a universal rule that more data...

  • Data Communication: Preserve the Evidence

    Good data communication preserves the structure of the evidence: the estimand, denominator, uncertainty, assumptions, and distinction between descr...

  • Rolling Windows in Signal Processing

    Rolling windows are local operators whose statistical meaning depends on window width, alignment, overlap, sampling rate, leakage control, and the ...

  • Traffic and Pedestrian Flow as Dynamical Systems

    Traffic and pedestrian flow can sometimes be modeled with conservation laws and continuum approximations, but the analogy with fluids has limits th...

  • The Risks and Limits of Artificial Intelligence

    A sober analysis of AI risk requires separating present operational harms, labor-market effects, security risks, model limitations, environmental c...

  • Binary Classification: Probabilities Before Labels

    Binary classification is not just choosing an algorithm. It requires a clear target, calibrated probabilities, realistic validation, and thresholds...

October

  • Mann-Kendall Trend Test: Assumptions and Pitfalls

    The Mann-Kendall test detects monotone association with time, but serial dependence, seasonality, ties, change points, and irregular sampling must ...

  • Natural Language Processing: Models, Tasks, and Evaluation

    Modern NLP ranges from classical token-based models to pretrained transformers and language models. The core challenges remain representation, eval...

  • Coverage Probability in Statistical Inference

    Coverage probability is a property of an interval procedure under repeated sampling. Nominal coverage, conditional coverage, prediction coverage, a...

November

  • Mathematics for Machine Learning: What Each Tool Is For

    Machine learning rests on linear algebra, calculus, probability, statistics, and optimization, but each mathematical tool matters because of the st...

  • Mann-Whitney U Test: What It Actually Tests

    The Mann-Whitney U test compares pairwise ordering between two independent distributions. It is not automatically a test of medians or a fallback w...

  • Linear Probability Models vs Logistic Regression

    The linear probability model and logistic regression target the same conditional event probability through different functional forms. The choice d...

December

  • Data Engineering: Reliable Data Systems

    Data engineering is the design of reliable data systems: ingestion, storage, contracts, transformations, orchestration, lineage, quality, and serving.

  • Renewable Energy Optimization Under Uncertainty

    Renewable-energy optimization is a constrained stochastic control problem involving generation, storage, transmission, demand, uncertainty, and mar...

  • Managing Data Science Under Uncertainty

    Data science and engineering both contain uncertainty. The useful management distinction is between discovery risk, implementation risk, and operat...

2022 (22 posts)

January

February

  • Optimizing Staff Scheduling with Linear Programming

    Discover how linear programming and Python's PuLP library can efficiently solve staff scheduling challenges, minimizing costs while meeting operati...

March

May

July

August

September

October

November

December

2021 (24 posts)

January

February

  • Traffic Crash KDE: Density Is Not Risk

    Kernel density estimation can map concentrations of traffic crashes, but crash density is not crash risk unless exposure and road-network geometry ...

March

April

May

June

July

August

September

October

November

December

2020 (47 posts)

January

February

March

April

May

June

July

August

  • Understanding Markov Chain Monte Carlo (MCMC)

    This article delves into the fundamentals of Markov Chain Monte Carlo (MCMC), its applications, and its significance in solving complex, high-dimen...

September

October

November

  • Data Visualization Best Practices

    Discover best practices for creating clear and compelling data visualizations that communicate insights effectively.

  • Bayesian Inference Explained

    Explore the fundamentals of Bayesian inference and how prior beliefs combine with data to form posterior conclusions.

  • A Primer on Simple Linear Regression

    Understand how simple linear regression models the relationship between two variables using a single predictor.

December

2019 (12 posts)

December

2016 (1 posts)

July

  • Probability Distributions as Statistical Models

    Probability distributions are models for random variables, not labels attached to datasets. Their parameters, support, tail behavior, and mean-vari...

2015 (1 posts)

July

Loading mathematical content