Skip to content

Readiness report prose contradicts its own data: healthy indexes called leftovers at phase 0, wrong drift direction at phase 3 #37638

Description

@fabrizzio-dotCMS

Found during a lab run of the ES→OpenSearch 3 migration, by Jamie Mauro.

Description

The readiness report's numbers are right; the sentences it wraps around them are not. Two instances,
both in MigrationReadinessService.evaluate(), both reassuring the reader about something other than
what is actually true.

1. Phase 0 calls healthy indexes "left over from an earlier migration attempt"

Rolled back to phase 0, the summary read:

Note: 2 indices from an earlier migration attempt are left over on OpenSearch; dual-write will
overwrite them on the next crawl/reindex.

Those were the current, in-sync OpenSearch indexes, built over the preceding 22 hours at phase 1 and
verified at 3 and 5 documents — matching Elasticsearch exactly. The report sees OpenSearch indexes
while at phase 0, concludes they must be debris, and says so.

An operator who drops to phase 0 to troubleshoot — which the guide presents as a safe move — is told
their good indexes are junk.

The second clause is wrong on its own terms too: dual-write does not backfill. Only a reindex
would repair those indexes if they genuinely were stale. Promising that dual-write will overwrite them
contradicts the rule the rest of the migration is built on.

The code is at MigrationReadinessService.java:~102. The comment above it shows the intent — "anything
still counted here is NOT a yet-to-be-built counterpart" — but a temporary rollback from a dual-write
phase produces exactly that state and is not distinguished from real debris.

2. Phase 3 volunteers the wrong drift direction

With OpenSearch missing two documents that Elasticsearch had:

"outOfSyncCount": 2,
"summary": "Phase 3 (OpenSearch only) — the final phase, nothing to advance to.
            No index shows Elasticsearch behind OpenSearch; still verify before any downgrade."

The actual drift was OpenSearch behind Elasticsearch. The summary is technically true — it reports
on the other direction — but it volunteers a reassurance about the direction that does not matter
here and says nothing about the gap that does.

At the terminal phase, where OpenSearch is the only engine serving traffic, a bare outOfSyncCount: 2
is the sole hint that live search is missing content. The code branches only on esBehindAnywhere
(MigrationReadinessService.java:~82-87), so the opposite condition has no sentence of its own.

Why this matters

The operator documentation tells the reader to trust this report over document counts and the admin
UI. Both of these send a reader in the wrong direction while the underlying data was correct all
along — which is the worst shape for a diagnostic to fail in, because the numbers next to the prose
look authoritative.

Acceptance Criteria

  • At phase 0, OpenSearch indexes that are in sync with their Elasticsearch counterparts are not
    described as leftovers from an earlier attempt.
  • No summary claims dual-write will repair a stale index; the repair is a reindex.
  • At phase 3, a summary that mentions drift direction addresses the direction actually present —
    and OpenSearch being behind Elasticsearch gets its own sentence, since that is the one affecting
    live traffic.
  • Tests cover the phase-0-after-rollback state and the phase-3 OS-behind state.

Additional Context

Same lab run as #37635, #37636 and #37637. Related in kind to #37635, where the report's prose reports
health while its content map is empty — this issue is the same failure mode with the data present.

Activity

  1. github-actions commented on Sep 24, 2026

    @github-actions
    Contributor
  2. fabrizzio-dotCMS commented on Sep 24, 2026

    @fabrizzio-dotCMS
    MemberAuthor

    QA Note — how to test this fix

    The fix is in #37735. Use the migration test stack from the tester guide (docs/backend/OPENSEARCH_MIGRATION_TESTER_GUIDE.md): dotCMS on 8082, Elasticsearch on 9200, OpenSearch on 9201.

    This fix only changes the text of the readiness report's summary. The numbers, the blockers and the safeToAdvance / safeToRollback answers are the same as before. Every check below reads the summary sentence.

    ⚠️ Prerequisite — the image must include #37723

    Images built from main between 2026-09-22 and the base-image republish cannot talk to OpenSearch at all (#37722). Before testing, confirm the image has the fix:

    docker exec <dotcms-container> /java/bin/java --list-modules | grep jdk.net

    If that prints nothing, stop: the report will hang.

    How to take the report

    curl -s -u admin:admin http://localhost:8082/api/v1/index/migration/readiness | jq '.verdict'

    The user needs CMS Admin plus the role in DOT_OS_MIGRATION_INDEX_VISIBILITY_ROLE_KEY, as the tester guide explains. Change phases the way the tester guide does, then restart dotCMS.

    Setup

    1. Start at Phase 1 with some content in the site. Run a full reindex (System → Maintenance → Index) and wait for it to finish.
    2. Take the report. Both content entries should say IN_SYNC, and the summary should say everything is in sync.

    Test 1 — rolling back to Phase 0 with nothing changed

    1. Set Phase 0 and restart. Do not edit any content.
    2. Take the report.
    3. Expected: outOfSyncCount is 0, and the summary says there is nothing to reconcile yet and it is safe to advance. It must not mention indices "left over from an earlier migration attempt".

    Test 2 — Phase 0 after content changed (the main case)

    1. Still at Phase 0, create and publish one piece of content (any type).
    2. Take the report.
    3. Expected:
      • outOfSyncCount is 2, the content entries say COUNT_DRIFT, and OpenSearch shows one document fewer than Elasticsearch.
      • safeToAdvance is still true.
      • The summary says OpenSearch holds different counts from Elasticsearch, which is what a rollback leaves behind once content changes. It says dual-write never backfills and that a full reindex is needed after advancing to Phase 1.
      • The summary must not say "left over", "earlier migration attempt" or "dual-write will overwrite them".
    4. Before the fix the summary said: "2 indices from an earlier migration attempt are left over on OpenSearch; dual-write will overwrite them on the next crawl/reindex."

    Test 3 — the Phase 0 advice is correct

    1. Without reindexing, set Phase 1 and restart. Take the report.
    2. Expected: safeToAdvance is false, with one blocker per content index asking for a full reindex. This confirms what the Phase 0 summary said would happen. Each blocker's note should say the copy is missing content because of "a reindex that never finished, or content changed while this engine received no writes". It must no longer say "it was never fully rebuilt".

    Test 4 — Phase 3 with OpenSearch behind Elasticsearch

    1. Leave the state from Test 2 as it is, with OpenSearch one document behind (do not reindex). Set Phase 3 and restart.
    2. Take the report.
    3. Expected:
      • outOfSyncCount is 2, and the summary asks for a full reindex.
      • A separate sentence says OpenSearch holds fewer documents than Elasticsearch in 2 indices. It adds that deleted content produces the same effect, and that the database comparison tells the two apart.
      • The summary must not say "No index shows Elasticsearch behind OpenSearch".
    4. Before the fix the summary asked for the reindex, then ended with "No index shows Elasticsearch behind OpenSearch; still verify before any downgrade."

    Test 5 — Phase 3 with equal counts

    1. Run a full reindex at Phase 3 and wait for it to finish, then take the report.
    2. Expected: outOfSyncCount is 0, and the summary no longer mentions OpenSearch holding fewer documents. Because Elasticsearch stops receiving writes at Phase 3, it may now be behind OpenSearch. In that case the summary shows the existing "WARNING: OpenSearch holds content Elasticsearch does not" sentence, which is correct.
  3. self-assigned this
    on Oct 2, 2026
  4. zJaaal commented on Oct 5, 2026

    @zJaaal
    Member

    QA Passed

    Tested the fix (#37735, 43267ce9b2) against the QA note's five tests. All pass.

    Environment

    Setup — Phase 1 + full reindex: both entries IN_SYNC, ES/OS/DB agree at 3 live and 5 working.

    Test Scenario Result
    1 Phase 0, nothing changed ✅ outOfSyncCount: 0, safe to advance. No "left over" / "earlier migration attempt"
    2 Phase 0, one contentlet published ✅ outOfSyncCount: 2, COUNT_DRIFT, safeToAdvance: true
    3 Phase 1, no reindex ✅ safeToAdvance: false, one blocker per content index
    4 Phase 3, OpenSearch behind ✅ outOfSyncCount: 2, correct drift direction stated
    5 Phase 3, after full reindex ✅ outOfSyncCount: 0, drift sentence gone

    Test 2 — DB 4 live / 6 working, ES matching, OS at 3 / 5. Summary:

    Phase 0 (Elasticsearch only). Safe to advance to Phase 1, but OpenSearch already holds copies that do not match. 2 indices on OpenSearch hold a different document count than Elasticsearch, which is what a rollback from a dual-write phase leaves behind once content changes: Phase 0 writes Elasticsearch only. Dual-write mirrors new writes and never backfills, so after advancing to Phase 1 run a full reindex; until then the Phase 1 report will block on them.

    Zero occurrences of "left over", "earlier migration attempt", "dual-write will overwrite".

    Test 3 — blockers confirm the Phase 0 advice was accurate. Note now reads "a reindex that never finished, or content changed while this engine received no writes (a rollback to Phase 0, for one)". "never fully rebuilt" gone (0 occurrences). Percentages correct: 5 of 6 (83.33%), 3 of 4 (75.00%).

    Test 4 — summary:

    Phase 3 (OpenSearch only) — the final phase, nothing to advance to. 2 indices have an OpenSearch copy that is missing, could not be measured, or holds materially less than the database; run a full reindex. OpenSearch holds fewer documents than Elasticsearch in 2 indices. Elasticsearch stopped receiving writes at the cutover, so content deleted or unpublished since the cutover also shows up this way; the database comparison is what tells the two apart, and it flags the indices counted above.

    "No index shows Elasticsearch behind OpenSearch" gone (0 occurrences). The deleted-vs-missing ambiguity is handled by pointing at the DB comparison, and the DB counts (4/6 vs OS 3/5) confirm it classified correctly.

    Test 5 — reindex brought OS to 4 / 6, matching ES and the DB. outOfSyncCount: 0, "fewer documents" gone.

  5. removed their assignment
    on Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions