Skip to content

[Draft] Make processing service support model retraining - #167

Draft
mohamedelabbas1996 wants to merge 8 commits into
mainfrom
feat/bioclip-antenna-integration
Draft

mohamedelabbas1996 wants to merge 8 commits into
mainfrom
feat/bioclip-antenna-integration

Conversation

@mohamedelabbas1996

@mohamedelabbas1996 mohamedelabbas1996 commented Sep 3, 2026 •

Copy link
Copy Markdown

Summary

Antenna can now retrain a classifier head from species that people have verified, and this is the half that does the training. A head sits on a frozen backbone, so retraining it is fitting a small matrix over embeddings that already exist rather than another pass over the images — seconds, not weeks.

Antenna prepares the training set, writes it to storage and hands over a URL. This downloads it, fits a new head, scores it against the head currently in service on the same held-out rows, saves it, and offers it as its own pipeline. It never replaces the running head: promoting one is a separate, deliberate step, because an automatic swap would let a single bad run quietly degrade every later classification.

The Antenna side is RolnickLab/antenna#1407.

List of Changes

# What it does How
1 A model can declare that its head can be retrained trainable on InferenceBaseClass, defaulting to False, mirrored into the algorithm config Antenna reads
2 Serves BioCLIP 2.5 classifiers whose head can be swapped trapdata/ml/models/bioclip.py: frozen ViT-H/14 backbone, linear head read from the Hub or from disk
3 A classification carries the embedding that produced it features on the classification response, so a caller can keep the vector without running the backbone twice
4 Retrains a head from a prepared dataset POST /train, and trapdata/ml/training.py
5 Keeps species the old head knew but the new data does not cover warm_start and restore_untrained_classes
6 Scores the new head against the one in service score_incumbent and evaluate, on the same held-out rows
7 Every head it produces becomes selectable trapdata/ml/trained_heads.py registers each saved head as its own pipeline, without a restart
8 Reports the outcome back to the Antenna job signed-token callback, so training can outlast the request that started it
9 Uploads the head so it outlives this service's disk optional upload URL in the request; the response says where it landed

Detailed Description

Why the dataset arrives as a URL

A project with a few hundred thousand verified labels runs to hundreds of megabytes. Sending that in one request is fragile and has to start over if the connection drops. A URL is small to hand over and cheap to retry.

Why nothing is promoted automatically

/train reports promote and the reason for it, and stops there. The head it was trained from stays exactly where it was; a retrain adds a choice rather than replacing one. A held-out set under MIN_MEANINGFUL_TEST_ROWS (30) is reported as too small to tell two heads apart, so a smoke test is not mistaken for evidence.

Why the species list comes from the dataset

Antenna sends the project's taxa list inside the dataset, and it wins over whatever happens to be in the rows. A species with no verified crops yet must still be something the head can predict, or the pipeline goes blind to it the moment it is retrained.

Why the head is uploaded

It used to be written only to this service's own disk, under a cache directory, so a cleared cache or a rebuilt container lost it while Antenna still reported the version as existing. The upload is optional: a caller that sends no URL gets exactly the previous behaviour, and a failed upload is logged rather than raised, since the training succeeded and the head is still on local disk.

Verification

  • test_head_retraining.py, test_bioclip_classifier.py and test_trainable_flag.py cover fitting, the declared species list winning over the rows, an unsupported head type being refused, a saved head being discoverable and servable, the three label-map shapes, and the upload including its failure path.
  • Run end to end against a live Antenna: a job started from the UI prepared 96 verified crops over 17 species, this service trained and saved a head, uploaded it, and Antenna registered the new version with the location recorded.

Known gaps

  1. A stored head is never fetched back. The weights are durable and their location is recorded, but a service that has lost its disk copy does not re-download them; that needs head discovery at startup.
  2. The incumbent is looked up by key. A head registered under a different key than the one Antenna asks for is reported as having no baseline, so promote can only ever be false for it. Worth deciding whether to fall back to the parent recorded in the training info.
  3. A classifier with neither a local head directory nor a published repository crashes the service at startup, with an error from the Hub client rather than a clear message. It should skip that pipeline and say why.
  4. Only linear heads are supported. Anything else is refused rather than quietly served as a linear head.

Antenna is gaining the ability to retrain a classifier head from species people
have verified, but only some models are worth retraining: a head over a frozen
backbone is cheap, a detector or a model trained end to end is not.

A model now declares whether its head can be retrained, and the worker reports
that in the algorithm config it registers with Antenna. The flag is per algorithm
rather than per service, because a service usually hosts several pipelines and
only one of them has a retrainable head.

Nothing is trainable unless it says so, and nothing here trains anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnKz4AS4iFrYj1GrgkZQbq
@coderabbitai

coderabbitai Bot commented Sep 3, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@mohamedelabbas1996 mohamedelabbas1996 changed the title [Draft] Report which algorithms can be retrained [Draft] Tell Antenna which models can have their classifier head retrained Sep 3, 2026
@mohamedelabbas1996 mohamedelabbas1996 changed the title [Draft] Tell Antenna which models can have their classifier head retrained [Draft] Make processing service support model retraining Sep 3, 2026
mohamedelabbas1996 and others added 2 commits September 17, 2026 13:00
Main moved model weights onto the new object-store endpoint, which is what the
branch's CI was failing on: the old URL now answers 403. Both algorithm response
builders take that fix and keep the `trainable` flag alongside it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnKz4AS4iFrYj1GrgkZQbq
…verified labels

A classifier head over a frozen backbone is cheap to retrain: the backbone never
changes, so a crop's embedding never changes either, and fitting a new head is a
small matrix over stored embeddings rather than another pass over the images. This
adds such a classifier and the endpoint that retrains it.

The classifier is an ordinary model here — BioCLIP 2.5 with a linear head, written
against InferenceBaseClass like every other — so it inherits batching, devices and
the API response shape rather than reimplementing them. Its forward pass returns the
embedding alongside the logits, and that embedding now travels with the
classification it produced, because it is what a future head must be fit on.

`POST /train` takes a dataset Antenna has already prepared, fits a head, and scores
it against the head currently in service on the same held-out rows. It never swaps
the running head: an automatic swap would let one bad run quietly degrade every
later classification. Each saved head is offered as its own pipeline, so it can be
selected like any other, and the head it was trained from stays exactly where it was.

Heads are discovered from disk at startup, so a restart does not lose them. Two
things write a label map — a head published on the Hub, and a head retrained here,
which stores its labels beside the counts and metrics of the run — so the loader
accepts both shapes; a head that can be trained but not served is no use.

Moved from antenna's processing_services/, where it had been a fork of the example
backend carrying its own copies of the schemas and base classes this repo already
defines. See RolnickLab/antenna#1407.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnKz4AS4iFrYj1GrgkZQbq
mohamedelabbas1996 added a commit to mohamedelabbas1996/antenna that referenced this pull request Sep 17, 2026
The service was added under processing_services/, which the README describes as
somewhere to keep a local demo backend copied from `example` — and that is exactly
what it was, a fork of `example` with BioCLIP bolted on. Two thirds of it was that
scaffolding: a second Algorithm base class, a second copy of fifteen schema classes
already defined in the companion repo, and four demo classifiers that had nothing
to do with BioCLIP.

It now lives in ami-data-companion, which is where the team's real inference code,
model loading and weight handling already are, and where the trainable flag it
depends on was added. The classifier is written against that repo's
InferenceBaseClass rather than a parallel one, so it is a model alongside the
others instead of a service beside the service. See RolnickLab/ami-data-companion#167.

Co-Authored-By: Claude <noreply@anthropic.com>
mohamedelabbas1996 and others added 2 commits September 17, 2026 14:27
Running the loop from Antenna against this service turned up three things the move
had left broken.

The classifier needs open_clip to build its backbone, and that was never a
dependency here, so describing the pipeline failed at startup with a bare import
error. It is declared now.

The published head maps a class index to a record holding the species name and its
iNaturalist id, not to a plain string. Reading it as a string gave every class a
dict for a name, which surfaced much later and far from the cause, as an unhashable
key while warm-starting from the current head. One helper now normalises the three
shapes a label map arrives in: that record form, a plain index-to-name map, and the
labels a retrain writes beside its counts and metrics.

Retrained heads defaulted to /data/bioclip-service, a path that only existed on the
machine the standalone service ran on. They now go beside the downloaded weights,
in their own directory so a training run cannot overwrite the head being served.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnKz4AS4iFrYj1GrgkZQbq
… move

Checking the moved service against what arrived turned up two things that had not
made the crossing.

The Panama head is a second published classifier over the same frozen backbone, and
its export keeps the label vocabulary inside the npz rather than in a file beside
it. That vocabulary is longer than the number of output classes — a species with no
training data can never be predicted, so it is not a class — which is why the head's
`classes` array indexes into it rather than lining up with it. Reading them as
parallel arrays would mislabel every class without failing, so it has its own loader
and its own test.

`scripts/export_logreg_head.py` is what turns a trained sklearn probe into the two
files a head is served from, and the classifier's own docstring points at it. It had
been left behind, which would have left no documented way to publish a new head. The
SSH tunnel script that serves this on a remote GPU box came with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnKz4AS4iFrYj1GrgkZQbq
…pipelines

Adds two species classifiers built on a frozen BioCLIP 2.5 ViT-H/14 backbone with a
linear logistic-regression head on top, one over the Newfoundland species list and one
over the Panama list, and registers them as the bioclip_2_5_newfoundland and
bioclip_2_5_panama pipelines. The head is an sklearn LogisticRegression exported to a
single Linear layer, so a softmax over its output reproduces sklearn's predict_proba.

The backbone is frozen, so a crop's embedding never changes. The forward pass returns
that embedding alongside the logits, and a classification now carries it as `features`,
so whatever consumes the classification can keep the vector that produced it without
running the backbone over the same crop twice.

Only serving is included here. Retraining a head from verified labels builds on this and
is a separate change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnKz4AS4iFrYj1GrgkZQbq
Brings the standalone BioCLIP 2.5 + LogReg serving change (#174) in as the base this
retraining work builds on, so the model that gets retrained is the one that PR reviews and
merges on its own. The serving code was already here; this records the relationship and
adds the serving-only registration tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnKz4AS4iFrYj1GrgkZQbq
mohamedelabbas1996 added a commit to mohamedelabbas1996/antenna that referenced this pull request Sep 21, 2026
The service was added under processing_services/, which the README describes as
somewhere to keep a local demo backend copied from `example` — and that is exactly
what it was, a fork of `example` with BioCLIP bolted on. Two thirds of it was that
scaffolding: a second Algorithm base class, a second copy of fifteen schema classes
already defined in the companion repo, and four demo classifiers that had nothing
to do with BioCLIP.

It now lives in ami-data-companion, which is where the team's real inference code,
model loading and weight handling already are, and where the trainable flag it
depends on was added. The classifier is written against that repo's
InferenceBaseClass rather than a parallel one, so it is a model alongside the
others instead of a service beside the service. See RolnickLab/ami-data-companion#167.

Co-Authored-By: Claude <noreply@anthropic.com>
…'s disk

A retrained head was written only to this service's own disk, under a cache
directory. A cleared cache, a rebuilt container or a replaced host lost it, and
with more than one replica a head trained on one was not servable by another.
Antenna recorded that a version existed but not where its weights were, so
promoting one later could mean promoting something already gone.

The training request may now carry a URL to upload the head to, and the response
reports where it landed so the caller can record it against the version. A caller
that sends no URL keeps exactly the current behaviour.

The upload is reported rather than raised on failure, like the result callback
beside it: the training itself succeeded, and the head is still on local disk, so
a failed upload costs durability rather than the ability to serve it.

Re-download is deliberately not included. The weights are durable and their
location is recorded, but fetching them back when a service has lost its disk
copy needs head discovery at startup, which is its own change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnKz4AS4iFrYj1GrgkZQbq
@mohamedelabbas1996
mohamedelabbas1996 force-pushed the feat/bioclip-antenna-integration branch from 456ffa1 to ac65025 Compare September 28, 2026 20:43
mohamedelabbas1996 added a commit to mohamedelabbas1996/antenna that referenced this pull request Sep 28, 2026
…e models live

The service was added under processing_services/, which the README describes as
somewhere to keep a local demo backend copied from `example` — and that is exactly
what it was, a fork of `example` with BioCLIP bolted on. Two thirds of it was that
scaffolding: a second Algorithm base class, a second copy of fifteen schema classes
already defined in the companion repo, and four demo classifiers that had nothing
to do with BioCLIP.

It now lives in ami-data-companion, which is where the team's real inference code,
model loading and weight handling already are, and where the trainable flag it
depends on was added. The classifier is written against that repo's
InferenceBaseClass rather than a parallel one, so it is a model alongside the
others instead of a service beside the service. See RolnickLab/ami-data-companion#167.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CnKz4AS4iFrYj1GrgkZQbq
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant