Updated the memory size estimate in contactmap - #1700
Open
amjjbonvin wants to merge 8 commits into
Open
amjjbonvin wants to merge 8 commits into
amjjbonvin wants to merge 8 commits into
Conversation
The `contactmap` memory gate skipped the module on machines with ample memory. Three problems compounded: - `get_necessary_memory()` looked up `models[0].file_name`, a bare basename, while `_run()` executes with cwd set to the contactmap step folder and the models live in the previous one. `os.path.getsize()` therefore raised on every real run and a bare `except Exception` silently fell back to a 10000 atom guess -- 0.745 Gb, multiplied by `ncores`. The estimate never looked at the input files at all. - The `file_size // 10` heuristic assumed 10 bytes per atom. A PDB ATOM record is 81 bytes and `extract_pdb_dt()` skips hydrogens, so the real figure is ~100-160 bytes per heavy atom. Memory is quadratic in the atom count, making this a 100-250x overestimate. - The requirement was scaled by `ncores` even though only `min(ncores, n_jobs)` distance matrices are ever live at once. Now the largest model is located by `stat` (via `rel_path`) and only that one is parsed, by a new `count_heavy_atoms()` helper that mirrors the filtering of `extract_pdb_dt()`. The peak accounts for `squareform()` allocating the full matrix while the condensed `pdist()` array is still alive. Lookup failures are logged instead of being swallowed. `get_available_memory()` now reports total physical memory: psutil's `available` excludes the reclaimable page cache, so an idle 32 Gb macOS host reports 14-17 Gb and the gate flips between runs on identical input. The existing tests only asserted `memory > 0` and so passed straight through the broken fallback path. They now assert the exact expected value, and a new test reproduces the two-step run directory layout that triggered the bug. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
rvhonorato
reviewed
Sep 23, 2026
Member
Author
|
Thanks Rodrigo - will change that one back to available to be safe (especially in the docker container for the server)
|
Updated docstring and implementation of get_available_memory to return available memory instead of total memory.
VGPReys
reviewed
Sep 25, 2026
VGPReys
reviewed
Sep 25, 2026
When the memory estimate exceeded what the host had available, the module skipped itself entirely and produced no contact maps at all. But the shortfall is usually one of parallelism, not of feasibility: each job holds one NxN distance matrix, so running fewer of them side by side lowers the requirement proportionally. The gate now computes how many jobs the available memory can feed and hands that number to the execution engine as `ncores`. Only when a single job does not fit is there no way to proceed, and that is the one case left that still skips the module. The reduction goes down to what actually fits rather than straight to 1, so affordable parallelism is not thrown away: with 8 cores requested, 0.5Gb per job and 1.6Gb free, 3 jobs run in parallel. The reduced count is passed to `get_engine()` through a copy of the params rather than by mutating `self.params`. `params.cfg` is written before `run()` so mutating would not have corrupted it, but keeping `self.params` as the record of what the user asked for is preferable -- the reduction is a local scheduling decision, not a change of intent. Tests cover the four outcomes (untouched, reduced to an intermediate count, reduced to 1, skipped) plus the params-not-mutated guarantee. Verified that the reduction reaches the `Scheduler` and that a real run on the degraded path still produces its heatmaps, chordcharts and TSVs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do and why?
Fixes #1699. The
[contactmap]memory gate was skipping the module on hostswith plenty of memory — the antibody-antigen tutorial reported
needs 37.25Gb has 16.83Gbon a 32 GB laptop.That 37.25 Gb figure never looked at the input files at all. Three separate
problems compounded in
get_necessary_memory()/get_available_memory()(
src/haddock/libs/libutil.py):The file lookup always failed. The estimate called
os.path.getsize(models[0].file_name), butPDBFile.file_nameis a barebasename (
libontology.py:49) while_run()executes with the cwd set tothe contactmap step folder (
modules/__init__.py:251) and the models livein the previous step folder.
getsize()therefore raised on every realrun, and a bare
except Exceptionsilently substituted a hard-coded 10000atom guess. That is where the number came from:
10000² × 8 / 1024³ = 0.745 Gb, timesncores = 50, is exactly 37.25 Gb.The size→atoms heuristic was off by an order of magnitude.
file_size // 10assumes 10 bytes per atom. A PDB ATOM record is 81 bytes,and
extract_pdb_dt()skips hydrogens, so the matrix only ever covers heavyatoms. Measured on
tests/golden_data, the real figure is 99–158 bytes perheavy atom — a 10–16× overestimate of the atom count, which because memory
is quadratic becomes a 100–250× overestimate of memory.
The scaling assumed more parallelism than exists. The per-model cost was
multiplied by
ncores, but onlymin(ncores, n_jobs)distance matrices areever live at once, and
n_jobsis the number of clusters (ortopXmodelswhen unclustered) — typically far below
ncores.What it does now
rel_pathinstead offile_name, so the filesare actually found.
stat(cheap) — matching what the docstringalways claimed but the code never did — and only that one file is parsed, by
a new
count_heavy_atoms()helper that mirrorsextract_pdb_dt()'s filterexactly. Reading one file costs a few ms and removes the guesswork entirely.
compute_distance_matrix()callingsquareform(pdist(coords)): the condensedN(N-1)/2array is still alivewhile
squareformallocates the fullN²matrix, so the peak is ~1.5× thefinal matrix. This was previously unaccounted for.
min(ncores, len(contact_jobs)). The checkmoved to just after the job list is built, which is the first point where
the job count is known.
except Exception. That silence is why this went unnoticed.get_available_memory()now reports total physical memory rather thanpsutil'savailable. On macOSavailableexcludes the reclaimable pagecache and compressed pages, so an idle 32 GB host reports 14–17 GB — meaning
the old gate could flip between runs on identical input.
Effect
(Every "before" is identical because the estimate never read the files.)
The denominator goes from a fluctuating 14–17 Gb to a stable 32 Gb.
The gate is kept rather than removed — the
NxNmatrix is still quadratic andcan still exhaust memory on a large enough system. It should now trip only when
that is genuinely about to happen.
How was this tested?
pytest tests/— 1796 passed, 5 skippedpytest integration_tests/(contactmap and util subset) — 14 passed, 1 skippedruff checkandruff format --checkclean on all changed files(the 5 pre-existing
E721/F401warnings intests/test_module_contmap.pyare unchanged from
main)The existing memory tests only asserted
memory > 0, so they passed straightthrough the broken fallback path and caught none of this. They now assert the
exact expected value against a reference formula. New tests:
test_count_heavy_atoms— the count must equal whatextract_pdb_dt()actually puts into the matrix, so the two cannot drift apart.
test_get_necessary_memory— exact match for the largest of severalmodels, plus a
< 0.1 Gbbound that would have failed on the old code.test_get_necessary_memory_uses_rel_path— reproduces the real two-step rundirectory layout, asserts that the bare
file_namedoes not resolve fromthe step folder, and that the estimate is correct anyway. This is the
regression test for the actual bug.
Verified manually against a simulated run directory as well: with cwd at
2_contactmap/and the model in1_prev/, the old lookup fails and the newone returns 36.3 MB.
AI assistance
Diagnosis and implementation were done with Claude Code.
Checklist
CHANGELOG.mdupdated for user-facing changesRelated issues
Closes #1699
Notes for reviewers