On 1 March I opened four pull requests against pyGAM, the Python library for generalized additive models. All four are still open.
The p-values
The one I spent longest on changes how p-values and effective degrees of freedom are computed for smooth terms. A smooth term is penalised, so it uses fewer degrees of freedom than it has basis functions, and a test that ignores the penalty gets the reference distribution wrong. Wood's 2013 approach tests the term through a rank-r truncated eigendecomposition of its coefficient covariance, with r set by the term's effective degrees of freedom. The pull request implements that.
The other three
- Overflow in link functions, wrong weights in some distributions, and a boolean bug in the B-spline basis.
- Restoring the scikit-learn estimator contract (
score,clone,predict_proba, tags,decision_function), so pyGAM models work inside scikit-learn tooling. - Categorical partial dependence, gridsearch memory use, packaging and reproducibility fixes.
What I took from it: statistical libraries fail quietly. A wrong p-value does not raise an exception. It gets copied into a results table.