Between 25 and 28 February I opened 11 pull requests against pgmpy, the Python library for probabilistic graphical models and causal inference. Reading code closely enough to fix it is the fastest way I know to learn how a library really works.
What I found
- Graph logic. I rewrote
minimal_dseparatoron the moralized ancestral graph, the standard construction for finding a separating set. - Silent data loss. State names were lost when a model was fit, and writing a model to the UAI format and reading it back lost its variable names.
- Inference plumbing.
querydid not forward its keyword arguments, the average treatment effect came back as NaN on an empty adjustment set, andApproxInferencebuilt its state space from the samples it drew rather than from the model, which misses any state that happens not to appear in the samples. - Model utilities.
simulate()dropped valid parent edges when a CPD was replaced for an intervention, and a dynamic Bayesian network'sget_evidencereturnedNone.
Eight of the eleven have since been closed, and three are still open.
What I would do differently
Several of the pull requests bundle two or three related fixes. A maintainer reviewing a library in their spare time would rather get one fix per request, each with a test that fails before the change and passes after it. Bundling made my week faster and their review slower, which is the wrong trade.