Weighted least squares and a linear fit
Suppose an experiment should follow a straight line. Ordinary least squares chooses the slope and intercept that minimize the sum of squared vertical residuals. Weighted least squares gives different observations different influence when their uncertainties differ.
The residual at a point is the measured y minus the model's prediction at that x. A positive residual places the measurement above the fitted line.
From residuals to weights
If x uncertainty is negligible and the y measurements have independent standard uncertainties σᵢ, the weighted objective is:
Each weight is the inverse variance, 1/σᵢ². A point with half the standard uncertainty has four times the weight. Its influence on an individual parameter also depends on where it lies in x; a point far from the center of the data can be especially influential on the slope.
If every uncertainty is equal, minimizing this objective gives the same slope and intercept as ordinary least squares. Multiplying all uncertainties by the same factor also leaves those estimates unchanged. The uncertainty convention still matters for the reported parameter uncertainties.
The NIST introduction to weighted least squares discusses why unequal measurement variability motivates weighting.
A small example
Use these illustrative observations, all with negligible x uncertainty:
| x | y | δy |
|---|---|---|
| 0 | 1.0 | 0.1 |
| 1 | 2.0 | 0.1 |
| 2 | 4.0 | 1.0 |
The third observation is less precise. The weighted fit has slope approximately 1.02857 and intercept 0.99048. An unweighted fit instead has slope 1.5 and intercept 0.83333. The weighted line stays closer to the first two measurements because their stated variances are smaller.
To try this in curve.fit, enter the values in x, y and δy, choose Linear, set δx to None and δy to Table, then fit. For the unweighted comparison, change δy to None and fit again. The earlier result remains in its own Fit tab. Application Help explains data entry and saved results.
Parameter uncertainties are a separate decision
For a model with N observations and p free parameters, the usual residual degrees of freedom are N − p:
Absolute propagates the entered standard uncertainties. Nominal, curve.fit's default, additionally estimates a common scale from the residual scatter. Both use the same weighted objective and give the same slope and intercept. See measurement uncertainties for how to choose the interpretation.
With only three points and two free parameters, this example has one residual degree of freedom. That makes an estimate of scatter fragile. A near-perfect line through a very small dataset is weak evidence about the reliability of its uncertainty estimate.
When x has uncertainty too
The objective above treats x as fixed. Selecting x uncertainties changes the fitting problem: curve.fit uses orthogonal distance regression rather than ordinary vertical least squares. It weighs discrepancies in both coordinates. The SciPy ODR documentation describes this distinction.
The Pearson–York example includes x and y uncertainties and is useful for exploring that case. It is not a y-only weighted-least-squares benchmark with its original settings.
What to check after fitting
Look for curvature, groups of residuals on one side of zero, and points whose influence is inconsistent with their uncertainty estimates. Check whether the selected range can distinguish slope from intercept. An apparently precise line can still extrapolate poorly beyond the observed range.