Articles, newest first
Learning Curves: Deciding Whether More Data Will Help
The request arrives as a budget line: ten thousand more labels, at two euros each. Whether they are worth it is not a matter of opinion. The learning curve s...
Read articleParasites Are Diagnosed by Species, Not by Symptom Lists
Parasitic infections are real and can be serious. The online parasite cleanse reverses the order of clinical parasitology: common symptoms become the diagnos...
Read articleLoRA Is a Low-Rank Model of the Fine-Tuning Update
LoRA is usually described as a memory-saving trick. More fundamentally, it assumes that a useful fine-tuning update can be represented in a low-dimensional s...
Read articleOne Factor at a Time Is Not an Experiment: Factorial Designs and Interactions
The team tests each of four process settings on its own, sees two of them make things worse, keeps the baseline, and never learns that the two "harmful" chan...
Read articlePseudo-Label Confidence Is Not the Same as Correctness
Self-training promotes model predictions into training labels. The usual safeguard is confidence thresholding, but confidence is produced by the same model t...
Read articleWasserstein Distance Is Geometry, Not Just Another Divergence
Wasserstein distance does not compare probability densities point by point. It asks how much probability mass must move, how far it must travel, and what tra...
Read articleActive Learning for Machine Learning: Getting More Value from Fewer Labels
Active learning improves machine learning by choosing which examples to label, not merely by asking for more labeled data.
Read articleRegression Discontinuity: Estimating an Effect From the Rule That Assigns It
Customers with a risk score of 600 or more get the credit line; those below do not. Comparing everyone above with everyone below gives an effect of 8.5. The ...
Read articleIntermittent Demand: Forecasting a Series That Is Mostly Zeros
Eighty-two percent of the weeks have no demand. Forecasting zero every week scores the lowest mean absolute error of any method tested, and delivers a fill r...
Read articleRegression to the Mean: The Improvement You Did Not Cause
Pick the worst ten machines, sites or agents, intervene, and watch them improve. Most of that improvement was going to happen anyway, and there is a formula ...
Read article








