Repository navigation
Stop three tests failing on every pull request - #164
Conversation
…emote host Three tests fail on every pull request for reasons unrelated to the changes under review, which blocks anything from merging on a green check. Two API tests asked for "the first" Vermont trap image and then asserted that the pipeline returned detections. The helper took that image from an unsorted glob, so which image it got depended on directory order, which differs between a developer checkout and a fresh CI checkout. One of the three Vermont images is an empty frame where nothing clears the detector's 0.80 score threshold, and CI consistently picked it, so the pipeline correctly returned zero detections and the assertions failed. The helper now sorts, and both tests name the image they need. The source image URL test fetched a thumbnail from a third-party host that now rejects arbitrary thumbnail widths with an HTTP 400. It now serves the image from the local test HTTP server the repository already uses elsewhere, which exercises the same download path without depending on the public internet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HcXFHJRXMrsHPX7xz9ZifF
|
Warning Review limit reachedNext included review available in 34 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (3)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
mihow
left a comment
There was a problem hiding this comment.
Claude says: Reviewed in place of the CodeRabbit pass that was rate-limited. This looks good to merge. It is posted as a comment review because it comes from the author's account, which cannot approve its own pull request.
What's solid
- Each fix targets the actual cause. Sorted selection removes the dependence on directory order, the two pipeline tests name an image that has detections, and
test_urlserves a generated image from the existing local test server (free port, context manager) while exercising the same download-and-open path. - The unequal image dimensions in
test_urlare a nice touch: they would catch a mixed-up image, not only a failed download. filenamesis added as the last parameter, so existing callers are unaffected. The six callers intrapdata/antenna/tests/now also receive sorted images; I checked that each tolerates an image without detections (len(response.detections) >= 0, or an explicit empty check), so the deterministic order does not introduce a new failure mode.
Worth knowing before merge
- The green checks predate the object store move. They ran on 2026-08-12. The CI workflow has no model cache, and the insect-orders pipeline these tests use downloads its weights from the previous object store host, which now returns 403 for both the classifier and the detector weights; the URLs in #168 return 200. A re-run of this pull request on its own would likely fail on those downloads. The two pull requests complement each other: #168's CI currently fails only
TestSourceImageSchema::test_url, which this one fixes. Suggest merging this first and then updating #168's branch, which should come back green. - Description wording. "Three tests have been failing on every pull request" holds for
test_url, but the two order-dependent tests fail intermittently rather than every time; both passed in #168's latest CI run. The fix is still the right one, since it removes the flakiness. It may also be worth a line pointing at #168, because the remote-weights dependency mentioned under "Notes for reviewers" has since become a real failure.
Smaller cleanups (low priority)
Three inline suggestions: an existence check for explicitly named fixtures, an assertion that can no longer fail, and passing arguments by keyword. None of them block merging.
| if filenames is None: | ||
| paths = sorted(images_dir.glob("*.jpg"))[:num] | ||
| else: | ||
| paths = [images_dir / filename for filename in filenames] |
There was a problem hiding this comment.
Claude says: Small robustness suggestion: an explicitly named file is not checked for existence here, so if the fixture is ever renamed or removed, the test would likely fail as a 404 from the test server inside the pipeline rather than as a clear missing-fixture error. For example:
paths = [images_dir / filename for filename in filenames]
missing = [path.name for path in paths if not path.exists()]
assert not missing, f"Test images not found in {images_dir}: {missing}"It may also be worth saying in the docstring that num is ignored when filenames is given.
| """ | ||
| test_images = self.get_test_images(num=1) | ||
| test_images = self.get_test_images(filenames=[IMAGE_WITH_DETECTIONS]) | ||
| assert test_images, "No test images found" |
There was a problem hiding this comment.
Claude says: With an explicit filename this assertion can no longer fail, because the helper always returns one URL per filename. It could be dropped, or made meaningful again by the existence check suggested in utils.py. The same applies at line 191.
The wrapper in test_api.py and the helper chain in utils.py passed subdir, num and filenames positionally. Passing them by keyword keeps these calls correct if another optional parameter is added to the helpers later. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012iwV2sD1vVtL5xEUN8Rgf3
|
The remaining CI failures are due to the new model URLs. Merging this, then will work on #168 |
Summary
Three tests fail on pull requests for reasons unrelated to what those pull requests change. One fails on every run: it downloads a thumbnail from Wikimedia, which now refuses arbitrary thumbnail widths and returns an HTTP 400 pointing at its allowed-sizes policy. The other two fail intermittently: they test whichever sample image the filesystem happens to list first, and one of the sample images is an empty frame with nothing for the detector to find. This makes all three reliable, by serving the thumbnail from a local test server and by having tests pick their images deterministically. No application code changes; the fixes are entirely in the tests and their helpers.
None of this was caused by an open pull request. The last test run on
main(2026-04-14) passed. The Wikimedia change happened after that, andmainhas not had a test run since, so its green status is out of date rather than evidence that the suite still passes.The two intermittent failures were not network problems at all. The two API tests asked for "the first" trap image and then checked that the ML pipeline found insects in it. The helper picked that image from an unsorted directory listing, so which image a test got depended on filesystem ordering, which differs between checkouts and between runs. When the listing put the empty Vermont frame first, the pipeline correctly returned no detections and the test had nothing to assert on. Whenever it put a different image first, the same tests passed, which is why they looked flaky rather than broken.
List of Changes
test_logits_in_classification_responseandtest_config_num_classification_predictionspass an explicit filename via a newIMAGE_WITH_DETECTIONSconstant intrapdata/api/tests/utils.py.get_test_image_urlssorts the glob result and accepts an optionalfilenamesargument, added last so existing callers are unaffected. The helpers passsubdir,numandfilenamesto each other by keyword.test_urlwrites an image to a temporary directory and serves it throughStaticFileTestServer, which the repository already uses for this purpose. The same download-and-open code path is exercised, and the image has unequal dimensions so a mixed-up image would be caught, not only a failed download.Verification
When this pull request was opened (2026-08-12), run against this branch with the GPU disabled, matching the CPU-only CI runners:
IMAGE_WITH_DETECTIONSat the empty frame reproduced the exact CI failure (AssertionError: No detections found in response), confirming the image choice is what these tests turn on.Rechecked on 2026-09-11, during review:
TestSourceImageSchema::test_url, with 54 other tests passing on Python 3.10 and 3.12.Notes for reviewers
🤖 Generated with Claude Code