Articles, newest first
Two Hundred Past Experiments Know More Than Your Next One
The experiment reports a 4.2 percent lift. The last two hundred experiments in the same programme had effects spread around a percentage point. Both facts ar...
Read articleQuantile Regression: Predicting the Range, Not the Average
The model predicts 36 minutes and 40 percent of deliveries take longer. The customer did not ask for the mean. They asked when the parcel would arrive, and t...
Read articleRAG, LoRA, and Fine-Tuning Solve Different LLM Problems
RAG, LoRA and fine-tuning are often presented as competing ways to improve an LLM. That framing is wrong. They intervene at different parts of the system, so...
Read articleCortisol Is a Dynamic Signal, Not a Diagnosis
Cortisol is essential physiology. Pathological cortisol excess is real, but social media often turns a dynamic hormone into a catch-all explanation for belly...
Read articleThe Winner's Curse in Model Selection: Why the Best Validation Score Is Too Good
A hundred hyperparameter configurations are compared on a thousand validation cases. The winner scores 82.4 percent. Its true accuracy is 80 percent, and the...
Read articleThe Inspection Paradox Is Length-Biased Sampling
The interval seen at a random time is not distributed like an interval chosen at random from the event sequence. Long intervals occupy more time and are ther...
Read articleSlice-Based Model Evaluation: Finding the Failures Average Metrics Hide
Slice-based evaluation exposes where a machine learning model fails by breaking aggregate performance into meaningful subgroups, conditions, and operational ...
Read articleCluster-Randomised Experiments: When You Randomise Stores and Analyse Customers
Twenty stores are randomised, ten to each arm, and the 4,000 customers are compared with a t-test. With no true effect and only 5 percent of the outcome vari...
Read articleHow to Fine-Tune an LLM Without Fooling Yourself
Fine-tuning an LLM is easy to start and surprisingly easy to do badly. The difficult parts are defining the behaviour to change, constructing the dataset, pr...
Read articleData Drift and Fairness: Monitoring Equity When Populations Change
A fair model at launch can become unfair in production when populations, behavior, policies, or measurement systems change.
Read article






