Repository navigation
Readiness report prose contradicts its own data: healthy indexes called leftovers at phase 0, wrong drift direction at phase 3 #37638
Description
Activity
github-actions commented
on Sep 24, 2026 on Sep 24, 2026 – with GitHub ActionsContributorMore actionsPRs linked to this issue
QA Note — how to test this fix
The fix is in #37735. Use the migration test stack from the tester guide (
docs/backend/OPENSEARCH_MIGRATION_TESTER_GUIDE.md): dotCMS on8082, Elasticsearch on9200, OpenSearch on9201.This fix only changes the text of the readiness report's
summary. The numbers, the blockers and thesafeToAdvance/safeToRollbackanswers are the same as before. Every check below reads the summary sentence.⚠️ Prerequisite — the image must include #37723Images built from
mainbetween 2026-09-22 and the base-image republish cannot talk to OpenSearch at all (#37722). Before testing, confirm the image has the fix:docker exec <dotcms-container> /java/bin/java --list-modules | grep jdk.net
If that prints nothing, stop: the report will hang.
How to take the report
curl -s -u admin:admin http://localhost:8082/api/v1/index/migration/readiness | jq '.verdict'
The user needs CMS Admin plus the role in
DOT_OS_MIGRATION_INDEX_VISIBILITY_ROLE_KEY, as the tester guide explains. Change phases the way the tester guide does, then restart dotCMS.Setup
- Start at Phase 1 with some content in the site. Run a full reindex (System → Maintenance → Index) and wait for it to finish.
- Take the report. Both content entries should say
IN_SYNC, and the summary should say everything is in sync.
Test 1 — rolling back to Phase 0 with nothing changed
- Set Phase 0 and restart. Do not edit any content.
- Take the report.
- Expected:
outOfSyncCountis0, and the summary says there is nothing to reconcile yet and it is safe to advance. It must not mention indices "left over from an earlier migration attempt".
Test 2 — Phase 0 after content changed (the main case)
- Still at Phase 0, create and publish one piece of content (any type).
- Take the report.
- Expected:
outOfSyncCountis2, the content entries sayCOUNT_DRIFT, and OpenSearch shows one document fewer than Elasticsearch.safeToAdvanceis stilltrue.- The summary says OpenSearch holds different counts from Elasticsearch, which is what a rollback leaves behind once content changes. It says dual-write never backfills and that a full reindex is needed after advancing to Phase 1.
- The summary must not say "left over", "earlier migration attempt" or "dual-write will overwrite them".
- Before the fix the summary said: "2 indices from an earlier migration attempt are left over on OpenSearch; dual-write will overwrite them on the next crawl/reindex."
Test 3 — the Phase 0 advice is correct
- Without reindexing, set Phase 1 and restart. Take the report.
- Expected:
safeToAdvanceisfalse, with one blocker per content index asking for a full reindex. This confirms what the Phase 0 summary said would happen. Each blocker's note should say the copy is missing content because of "a reindex that never finished, or content changed while this engine received no writes". It must no longer say "it was never fully rebuilt".
Test 4 — Phase 3 with OpenSearch behind Elasticsearch
- Leave the state from Test 2 as it is, with OpenSearch one document behind (do not reindex). Set Phase 3 and restart.
- Take the report.
- Expected:
outOfSyncCountis2, and the summary asks for a full reindex.- A separate sentence says OpenSearch holds fewer documents than Elasticsearch in 2 indices. It adds that deleted content produces the same effect, and that the database comparison tells the two apart.
- The summary must not say "No index shows Elasticsearch behind OpenSearch".
- Before the fix the summary asked for the reindex, then ended with "No index shows Elasticsearch behind OpenSearch; still verify before any downgrade."
Test 5 — Phase 3 with equal counts
- Run a full reindex at Phase 3 and wait for it to finish, then take the report.
- Expected:
outOfSyncCountis0, and the summary no longer mentions OpenSearch holding fewer documents. Because Elasticsearch stops receiving writes at Phase 3, it may now be behind OpenSearch. In that case the summary shows the existing "WARNING: OpenSearch holds content Elasticsearch does not" sentence, which is correct.
- linked a pull request that will close this issuefix(opensearch): readiness summary states the drift that is actually present (#37638) #37735
on Sep 24, 2026 - added a commit that references this issue
on Sep 25, 2026 QA Passed
Tested the fix (#37735,
43267ce9b2) against the QA note's five tests. All pass.Environment
- Image
dotcms/dotcms:trunk, revision11940bb79(2026-10-03), confirmed to contain the fix commit - Prerequisite OpenSearch requests hang in the Docker image: httpcore5 5.4.3 needs jdk.net, which the java-base runtime omits #37722 verified:
jdk.net@25.0.4.1present in the runtime - Lab
docker/docker-compose-examples/single-node-os-migration, clean volumes, target OpenSearch 3.8.0 over TLS - arm64
Setup — Phase 1 + full reindex: both entries
IN_SYNC, ES/OS/DB agree at 3 live and 5 working.Test Scenario Result 1 Phase 0, nothing changed ✅ outOfSyncCount: 0, safe to advance. No "left over" / "earlier migration attempt"2 Phase 0, one contentlet published ✅ outOfSyncCount: 2,COUNT_DRIFT,safeToAdvance: true3 Phase 1, no reindex ✅ safeToAdvance: false, one blocker per content index4 Phase 3, OpenSearch behind ✅ outOfSyncCount: 2, correct drift direction stated5 Phase 3, after full reindex ✅ outOfSyncCount: 0, drift sentence goneTest 2 — DB 4 live / 6 working, ES matching, OS at 3 / 5. Summary:
Phase 0 (Elasticsearch only). Safe to advance to Phase 1, but OpenSearch already holds copies that do not match. 2 indices on OpenSearch hold a different document count than Elasticsearch, which is what a rollback from a dual-write phase leaves behind once content changes: Phase 0 writes Elasticsearch only. Dual-write mirrors new writes and never backfills, so after advancing to Phase 1 run a full reindex; until then the Phase 1 report will block on them.
Zero occurrences of "left over", "earlier migration attempt", "dual-write will overwrite".
Test 3 — blockers confirm the Phase 0 advice was accurate. Note now reads "a reindex that never finished, or content changed while this engine received no writes (a rollback to Phase 0, for one)". "never fully rebuilt" gone (0 occurrences). Percentages correct: 5 of 6 (83.33%), 3 of 4 (75.00%).
Test 4 — summary:
Phase 3 (OpenSearch only) — the final phase, nothing to advance to. 2 indices have an OpenSearch copy that is missing, could not be measured, or holds materially less than the database; run a full reindex. OpenSearch holds fewer documents than Elasticsearch in 2 indices. Elasticsearch stopped receiving writes at the cutover, so content deleted or unpublished since the cutover also shows up this way; the database comparison is what tells the two apart, and it flags the indices counted above.
"No index shows Elasticsearch behind OpenSearch" gone (0 occurrences). The deleted-vs-missing ambiguity is handled by pointing at the DB comparison, and the DB counts (4/6 vs OS 3/5) confirm it classified correctly.
Test 5 — reindex brought OS to 4 / 6, matching ES and the DB.
outOfSyncCount: 0, "fewer documents" gone.- Image
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Found during a lab run of the ES→OpenSearch 3 migration, by Jamie Mauro.
Description
The readiness report's numbers are right; the sentences it wraps around them are not. Two instances,
both in
MigrationReadinessService.evaluate(), both reassuring the reader about something other thanwhat is actually true.
1. Phase 0 calls healthy indexes "left over from an earlier migration attempt"
Rolled back to phase 0, the summary read:
Those were the current, in-sync OpenSearch indexes, built over the preceding 22 hours at phase 1 and
verified at 3 and 5 documents — matching Elasticsearch exactly. The report sees OpenSearch indexes
while at phase 0, concludes they must be debris, and says so.
An operator who drops to phase 0 to troubleshoot — which the guide presents as a safe move — is told
their good indexes are junk.
The second clause is wrong on its own terms too: dual-write does not backfill. Only a reindex
would repair those indexes if they genuinely were stale. Promising that dual-write will overwrite them
contradicts the rule the rest of the migration is built on.
The code is at
MigrationReadinessService.java:~102. The comment above it shows the intent — "anythingstill counted here is NOT a yet-to-be-built counterpart" — but a temporary rollback from a dual-write
phase produces exactly that state and is not distinguished from real debris.
2. Phase 3 volunteers the wrong drift direction
With OpenSearch missing two documents that Elasticsearch had:
The actual drift was OpenSearch behind Elasticsearch. The summary is technically true — it reports
on the other direction — but it volunteers a reassurance about the direction that does not matter
here and says nothing about the gap that does.
At the terminal phase, where OpenSearch is the only engine serving traffic, a bare
outOfSyncCount: 2is the sole hint that live search is missing content. The code branches only on
esBehindAnywhere(
MigrationReadinessService.java:~82-87), so the opposite condition has no sentence of its own.Why this matters
The operator documentation tells the reader to trust this report over document counts and the admin
UI. Both of these send a reader in the wrong direction while the underlying data was correct all
along — which is the worst shape for a diagnostic to fail in, because the numbers next to the prose
look authoritative.
Acceptance Criteria
described as leftovers from an earlier attempt.
and OpenSearch being behind Elasticsearch gets its own sentence, since that is the one affecting
live traffic.
Additional Context
Same lab run as #37635, #37636 and #37637. Related in kind to #37635, where the report's prose reports
health while its content map is empty — this issue is the same failure mode with the data present.