Articles, newest first
Offline Change-Point Detection: Segmenting a Series After the Fact
Sequential detection asks whether something has changed as of now. The retrospective question is different: given two years of a sensor's history, where did ...
Read articleCorrelation Does Not Determine Joint Tail Risk
Dependence is more than correlation. Two multivariate models can agree on every marginal distribution and on a global rank-correlation measure while disagree...
Read articleAnnotator Disagreement Sets the Ceiling: What Label Noise Does to Every Number You Report
Two annotators agree on 82 percent of items. A model that predicts the truth perfectly will score 90 percent against their labels, and two models five points...
Read articleA t-SNE or UMAP Plot Is Not Evidence That Clusters Exist
A two-dimensional embedding is a model of selected relationships in the original data, not a neutral photograph of high-dimensional geometry. t-SNE and UMAP ...
Read articleSelective Prediction in Machine Learning: When Models Should Abstain
Selective prediction gives machine learning systems a third option: predict when confidence is adequate and abstain when the cost of being wrong is too high.
Read articleRecurrent Failures: Why Time to First Failure Throws Away Two Thirds of the Data
Four hundred machines, 880 failures over three years. The time-to-first-failure analysis uses 299 of them, reports a rate a quarter too low, and cannot say t...
Read articleUrine pH Is Not Blood pH
Diet can alter renal acid load and urine pH. That does not mean ordinary food meaningfully “acidifies the blood” in a healthy person. Blood pH is tightly reg...
Read articleSynthetic Control: Evaluating an Intervention on One Unit
One plant got the new maintenance regime. Before-after says it did nothing, because demand was rising. Comparing against the other plants says it did too lit...
Read articlePercentile Metrics: Why p95 Latency Is Harder to Move and Harder to Measure
The change makes 97 percent of requests five percent faster and the remaining three percent slightly more likely to hit the slow path. The median improves by...
Read articleRepresentation Learning for Tabular Data: Beyond Manual Feature Engineering
Representation learning for tabular data is not about replacing feature engineering blindly. It is about learning useful structure while respecting the const...
Read article








