Calibrate against several LC setups at once (v4.3.0) - #117
Calibrate against several LC setups at once (v4.3.0)#117RobbinBouwmeester wants to merge 1 commit into
Conversation
Calibration ranks all 6,543 heads of the multitask model and keeps one, so a gradient that sits between trained setups is described by the closest single one and the rest of the matrix is discarded. MultiHeadRidgeCalibration keeps that ranking, calibrates the 80 best heads individually with SplineTransformerCalibration, and fits a ridge from those calibrated estimates onto the observed retention times. Each head then contributes an estimate already in the unit of the reference and the ridge decides how much to trust it. On the eight PRIDE setups no DeepLC model was trained on it lowered the held-out error on all eight, median 13 % relative to the gradient (0.01248 to 0.01090 MAE/span): 0.7 % on the setup that pools fractions into one run, 7.5 to 17 % on four others, 28 to 38 % on the three where a single head fitted worst, the largest gain on the smallest reference (230 peptidoforms). Fitting is no slower than the current path (median 1.0 s against 2.3 s) because the head ranking is vectorised, and prediction is unchanged since the full matrix is computed anyway. Opt in by passing the calibration; the default is untouched. Calibration gains a uses_all_heads flag, which is what tells core to hand over the whole matrix instead of one column. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
|
||
| from importlib.metadata import version | ||
|
|
||
| from deeplc.calibration import MultiHeadRidgeCalibration |
There was a problem hiding this comment.
Does this explicitly need to be part of the public API? As part of deeplc.calibration, it would be. But I don't think an explicit mention in __all__ is needed. It's not consistent with the other items in __all__.
| The default path keeps one setup head: it ranks all heads by Pearson correlation to the | ||
| reference and fits a spline on the winner. A gradient that no trained setup matches exactly | ||
| is then described by the closest single setup, and the rest of the matrix is discarded. | ||
|
|
There was a problem hiding this comment.
We should probably compile some instructions for LLMs in how to document code. This initial paragraph can be quite confusing. I expected an description of this class, and instead got one of the previous MT calibration method 😅 Generally, it could also be a bit more concise.
There was a problem hiding this comment.
In general, it's also clear that Claude writes comments, docstrings, and PR messages from the context of the conversation that led to the work.
I had a short chat with Claude about it to come up with some instructions 😅
Documentation should be self-contained, not conversation-dependent.
Comments, docstrings, and PR descriptions must be understandable to a reader with no access to the conversation or reasoning that produced the code. Do not assume the reader knows what was discussed, tried, or rejected beforehand.
Comments and docstrings should describe the current state and behavior of the code: what it does and how it is used. They should not read as a log of the development process. Historical context (why a previous approach was replaced, what motivated a design choice) belongs in commit messages, not in comments or docstrings, unless it is strictly necessary to understand the current implementation, in which case it should be brief.
PR descriptions may include more of this process-oriented information: the reasoning behind the change, alternatives considered, and relevant historical context. This is the appropriate place for a log of the development process, distinct from the code documentation itself.
| #: Whether ``fit`` and ``transform`` take the whole ``(n, n_heads)`` prediction matrix of a | ||
| #: multitask model instead of a single series. False for every calibration here except | ||
| #: :class:`MultiHeadRidgeCalibration`. | ||
| uses_all_heads: bool = False |
There was a problem hiding this comment.
I'd also advocate for instructing Claude to only use inline comments to describe code when not obvious from the code itself. To me, this line feels quite self-explanatory and does not require three lines of comments to explain itself :P
| On the eight PRIDE setups that no DeepLC model was trained on, this lowered the held-out | ||
| error on all eight: a median of 13 % relative to the observed gradient (0.01248 to 0.01090 | ||
| MAE/span), from 1 % on a setup whose retention times are not a single gradient to 38 % on the | ||
| smallest reference of 230 peptidoforms. The cost is ``n_heads`` spline fits and one ridge on | ||
| the reference; prediction is unchanged, because the full matrix is computed either way. |
There was a problem hiding this comment.
Could also be left out? Benchmarking results can go in the PR message, but do not really need to be part of a docstring.
| accepted and treated as a single head, so a single-task model still works. | ||
|
|
||
| """ | ||
| from sklearn.linear_model import RidgeCV |
There was a problem hiding this comment.
Should not be a lazy import; move to head of module.
What
Calibration ranks all 6,543 heads of the multitask model and keeps exactly one (
_best_correlating_head), so a gradient that sits between trained setups is described by the closest single setup and the other 6,542 columns are thrown away.MultiHeadRidgeCalibrationkeeps that same ranking, calibrates the 80 best heads individually withSplineTransformerCalibration, and fits a ridge regression from those calibrated estimates onto the observed retention times. Each head contributes an estimate already in the reference's unit; the ridge decides how much to trust each one.Opt in per call — the default path is untouched:
Calibration.uses_all_heads(new,Falseeverywhere else) is what tellscalibrate/predict_and_calibrateto hand over the whole(n, n_heads)matrix instead of one column.Result on the eight held-out PRIDE setups (Figure 1b corpus)
Reference and test peptides are disjoint; relative MAE is MAE / observed gradient span. Measured through the public API, not a notebook reimplementation.
8 of 8 improved, median 0.01248 → 0.01090 (13 %). PXD082486 pools several fractions into one run, so its retention times are not a single gradient and nothing helps there; PXD080314 is the wheat-gluten out-of-domain case.
Cost
Not slower than the current path: median calibration fit 2.3 s → 1.0 s, because
_best_correlating_headloops in Python over 6,543 columns withnp.corrcoefper column while the new ranking is vectorised. (That suggests a separate small win: vectorising the existing helper.) Prediction is unchanged — the full matrix is computed either way.How 80 was chosen
A sweep over k = 1, 2, 3, 5, 10, 20, 40, 80, 160, 320, 640, 1280, 2000 on the same corpus: median relative MAE 0.01248 (k=1) → 0.01137 (10) → 0.01090 (80, flat to 320) → 0.01125 (2000). Also compared, and rejected:
The class also never fits more weights than half the reference supports, which matters for the 230-peptidoform case.
Verification
pytest tests: 141 passed, including 12 new tests intests/test_multihead_calibration.py(beats a single head when the target mixes two setups, records the best head, guards on unfitted use and mismatched head counts, 1-D input, weight cap, empty input, and the twocoreintegration paths).ruff check/ruff format --check: clean.docs/source/models.rstupdated.🤖 Generated with Claude Code