Skip to content

Build read aloud chunks from the text selection OCR instead of a separate request - #1602

Merged
cdrini merged 3 commits into
internetarchive:masterfrom
cdrini:tts-chunks-from-text-layer
Oct 1, 2026
Merged

cdrini merged 3 commits into
internetarchive:masterfrom
cdrini:tts-chunks-from-text-layer

Conversation

@cdrini

@cdrini cdrini commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Feature/refactor. Read aloud used to request each page's text a second time from BookReaderGetTextWrapper.php (paragraphs mode), even though the text selection plugin had already fetched the same page's djvu xml. PageChunk.fetch now builds its chunks from the text selection plugin's OCR (getPageText), and only falls back to the page chunk URL when that plugin is missing or disabled, has no OCR for the page, or fails to load it.

Technical

  • PageChunk._chunkOcrPage is a JS port of the server's extract_paragraphs. A chunk ends at the first word ending in "." once it has more than 25 words, or at the end of a line once it has more than 50. It returns one rect per line, or per part of a line when a chunk ends mid-line, and skips header/footer paragraphs (x-role). We likely want to tweak/improve some of this logic.
  • It returns the same [text, ...rects] shape as the endpoint, so _fromTextWrapperResponse (hyphen removal, rect fixing) still applies unchanged.
  • Words the text layer drops as unpositionable (at 0,0 or without coords) are skipped too, so chunks don't depend on whether the text layer has rendered yet.
  • The server version resets its line box with the fields in the wrong order after a sentence ends mid-line, so the next rect comes out with right=maxsize, top=-1. That's the "ridiculously wide first rect" _fixChunkRects works around. The JS port doesn't have this bug.
  • Pages buffered ahead by read aloud now go through the text selection plugin's page cache (10 pages), so the text layer reuses them when it renders instead of fetching them again.

Testing

  • New unit tests for _chunkOcrPage and the fetch fallbacks.
  • Compared against the live endpoint on pages 5, 20, 40, 100 and 150 of adventureofsherl0000unse (most with headers/footers): the text was identical, and the rects were identical except the ones the server corrupts as described above.
  • To verify: open a book, start read aloud, and check in the network tab that there are no BookReaderGetTextWrapper.php requests without mode=djvu_xml, and that highlighting and page turning still follow along.

Screenshots/Videos

N/A, no visual change.

🤖 Generated with Claude Code

…rate request

Co-Authored-By: Drini Cami <cdrini@gmail.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 64.11%. Comparing base (47837c8) to head (0487b44).

Additional details and impacted files
@@            Coverage Diff             @@
##           master    #1602      +/-   ##
==========================================
+ Coverage   63.84%   64.11%   +0.27%     
==========================================
  Files          69       69              
  Lines        6253     6290      +37     
  Branches     1386     1398      +12     
==========================================
+ Hits         3992     4033      +41     
+ Misses       2222     2218       -4     
  Partials       39       39              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Co-Authored-By: Drini Cami <cdrini@gmail.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* @param {Element} ocrPage djvu xml `OBJECT` element
* @return {Array<[String, ...DJVURect[]]>}
*/
static _chunkOcrPage(ocrPage) {

@cdrini cdrini Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a JS translation of the python code. Having this code here means we can make a number of improvements to how this works much more easily!

Comment thread src/plugins/tts/PageChunk.js Outdated

@bfalling bfalling left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM (one suggestion re: ===

Co-Authored-By: Drini Cami <cdrini@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@cdrini
cdrini merged commit e4c8e6b into internetarchive:master Oct 1, 2026
10 checks passed
@cdrini
cdrini deleted the tts-chunks-from-text-layer branch October 1, 2026 16:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants