Skip to content

Add a dev tools area with benchmark, memory-leak, streaming, parser, and visual-diff tools - #453

Draft
sophiedeziel wants to merge 22 commits into
developfrom
feature/benchmark-page
Draft

sophiedeziel wants to merge 22 commits into
developfrom
feature/benchmark-page

Conversation

@sophiedeziel

@sophiedeziel sophiedeziel commented Sep 6, 2026 •

Copy link
Copy Markdown
Collaborator

Disclaimer

As this kind of toolset is what I always dreamed of (for this project), this is a POC that is 100% vibe-coded and probably mostly slop.

Code is cheap, the tools are still valuable.

Summary

This PR adds a devtools/ area for internal development tooling, served alongside the demo by the existing dev server but not deployed with it. It ships five tools aimed at the 3.0.0 release gates in #402, built on a small shared infrastructure layer (devtools/lib/) that handles version discovery on jsDelivr, per-version import maps (each release loads the three/lil-gui versions it was published against), sandboxed iframe runners, and a 2.x/3.x API compat shim.

Run npm run dev and open http://localhost:8080/devtools/. Every page has a light/dark toggle top right, persisted across pages and reloads (previews and captures stay dark in both themes so numbers and pixel diffs never depend on the page theme).

Dev tools landing page

Dev tools landing page, light mode

Debugger

Steps through a gcode file like a debugger. Click a row's gutter to set a breakpoint, then step forward or back one command at a time, continue to the next breakpoint, or use Run to first to jump past the slicer's metadata comment preamble (on Benchy it lands on command 1,149, the first M73 after PrusaSlicer's comments). The preview renders the toolpath live as you step, and a State panel shows the interpreter's State object (position, positionShift, tool, units, isHomed) with the keys that changed on the last step highlighted. Stepping back re-executes from the start, which stays fast: benchy's 163,474 commands re-execute in about 40 ms. The loaded file and breakpoints survive page reloads via localStorage, with corrupt stored values discarded safely. A collapsed Instantiation settings editor exposes the constructor options as editable JSON with live validation, curated presets, Reset to default, and persistence on every edit; applying rebuilds the session while keeping breakpoints and position, and a rejected settings object falls back to the last-known-good ones. The page is full width and the two columns are resizable by dragging the separator.

It doubles as a parser inspector: summary stats and a command-type histogram (both collapsed by default after loading), plus a filterable command list that virtual-scrolls through all of Benchy's 163k commands with no pagination. Full inspection targets 3.x builds since 2.x does not export the parser; older versions degrade to a friendly notice.

Two more findings came out of building it: constructing a preview without buildVolume throws in 3.x (the SceneManager constructor reads this.buildVolume.x unconditionally), and Mach3-style bare S900/F140 lines currently parse as gcodes s900/f140.

Debugger

Benchmark

Compares two library versions head to head on one of the demo gcode files. Any published npm version can face any other, or the local dist/ build. Each version runs sequentially in a fresh sandboxed iframe, runs are interleaved A/B/A/B, and the table reports medians for parse time, geometry build and render time, time to first render, FPS while orbiting, peak JS heap, triangles, and draw calls. Results can be copied as CSV or as Markdown ready to paste into #402.

Benchmark

Memory-leak tester

Runs repeated create, load, render, dispose cycles against a fresh canvas each time to mimic a framework remount, then charts the JS heap per cycle along with the renderer.info geometry and texture counts read after dispose() — any nonzero residual means dispose leaked GPU resources. The verdict compares the median heap between the first and second half of the steady-state cycles, which stays calm through GC sawtooth and allocator plateau jumps instead of crying wolf on healthy builds. The local build currently passes with a flat heap and flat counters.

Memory-leak tester

Streaming equivalence checker

Parses the same file twice in one version, once as a whole string and once as a ReadableStream chunked at a configurable size, then diffs the parser and job stats plus a rendered-triangle fingerprint. Small chunk sizes maximize the odds of splitting a command across a chunk boundary, which is exactly where streaming bugs live.

It caught a real bug on its first run: parser.lineCount was one higher on the whole-string path for every newline-terminated file (a phantom empty trailing line from split), and digging in revealed the streamed path also fabricated lines on chunks without newlines. Both sides are now fixed in src/ on this PR, the ingestion-equivalence suite asserts line and command counts across all four modes, and the checker reports all metrics green.

Streaming equivalence checker

Visual diff

Renders the same file in two versions with a pinned camera and target, captures both canvases in the same frame (required on 3.x, which uses preserveDrawingBuffer: false), and pixel-diffs the captures with an adjustable threshold. Changed pixels light up magenta over a dimmed grayscale base. Local vs local diffs at exactly 0.00%, and 2.18.0 vs local on Benchy shows a real 9.09% difference in scene composition relative to the origin.

Visual diff

Also in this PR

The rollup UMD output was missing a global mapping for lil-gui, so rollup warned on every build and guessed lilGui, which is not the global that lil-gui's UMD build actually exposes (lil). One line in rollup.config.mjs fixes both the warning and the broken guess.

The CodeQL alerts raised on this PR are addressed: interpolated values in the benchmark and streaming-equivalence tables are HTML-escaped, and the debugger's postMessage dispatches guard message types with Object.hasOwn.

A debugger-screenshots skill (in .agents/) documents how to capture before/after PR proof shots from the Debugger with custom gcode, following the pattern PR #439 established.

A local multi-angle code review of the branch was applied before marking ready: benchmark and visual-diff runners now dispose their previews so a finished side cannot contaminate the other side's measurements, breakpoints survive re-loads, pre-load edits persist, the debugger parses once and steps with exact full renders via the new public executeCommands API, the default baseline tracks the newest published stable instead of hardcoded 2.x, presets pass their real colors through, theme variables fall back to dark without JS, and shared helpers/CSS were extracted into devtools/lib.

Related to #402.

Assisted by Claude Code - Fable 5

@sophiedeziel sophiedeziel added 3.0 Targeted for the 3.0 release meta a non-functional change, doesn't require a new release labels Sep 6, 2026
@github-actions

github-actions Bot commented Sep 6, 2026 •

Copy link
Copy Markdown

Visit the preview URL for this PR (updated for commit 62a50b2):

https://gcode-preview--pr453-feature-benchmark-pa-uqxm5awh.web.app

(expires Wed, 07 Oct 2026 17:37:17 GMT)

🔥 via Firebase Hosting GitHub Action 🌎

Sign: 59bd114ae4847b32c2bba0b68620b9069a3e3531

Comment thread devtools/benchmark/bench.js Fixed
Comment thread devtools/debugger/runner.js Fixed
Comment thread devtools/streaming-equivalence/main.js Fixed
Comment thread devtools/debugger/runner.js Fixed
@sophiedeziel sophiedeziel added poc proof-of-concept. not intended te be merged into develop marinating Left to simmer — needs more thought before acting labels Sep 6, 2026
sophiedeziel added a commit that referenced this pull request Sep 6, 2026
Addresses the CodeQL DOM-text-reinterpreted-as-HTML alerts on PR #453:
version labels and runner-reported metric values now go through an
escapeHtml helper before landing in innerHTML.

Assisted by Claude Code - Fable 5
sophiedeziel added a commit that referenced this pull request Sep 6, 2026
…tools

The debugger gains a collapsed Instantiation settings editor: the full
constructor options object as editable JSON with live validation,
curated presets, Reset to default, and persistence on every edit.
Applying rebuilds the session on a fresh canvas while preserving
breakpoints and the current command; a constructor rejection falls back
to the last-known-good settings. The page is now full width and the
commands/panel columns are resizable by dragging a themed separator
(double-click resets, split persisted).

The whole suite gains a light/dark toggle on every page, built on
shared semantic theme variables in devtools/lib (dark matches the
previous look exactly). Preview canvases deliberately stay dark in both
themes so benchmark numbers and visual-diff captures are
theme-independent; the memory-leak chart re-reads its colors and
redraws on toggle.

Also hardens both debugger postMessage dispatches with Object.hasOwn
guards, addressing the CodeQL unvalidated-dynamic-method-call alert on
PR #453.

Assisted by Claude Code - Fable 5
@sophiedeziel sophiedeziel mentioned this pull request Sep 6, 2026
7 of 10 tasks
benchmark.html compares two library versions (any published npm version
via jsDelivr, or the local dist build) on one of the demo gcode files.
Each version runs sequentially in a fresh sandboxed iframe with its own
import map, resolving the three/lil-gui versions that release was
published against, so runs cannot contaminate each other.

Measures the #402 release-gate metrics — parse time, time to first
render, FPS while orbiting, peak JS heap — plus geometry build time,
triangles and draw calls, with interleaved runs and a median-based
comparison table.

Assisted by Claude Code - Fable 5
Silences rollup's 'Missing global variable name' warning and fixes the
guessed global: lil-gui's UMD build exposes 'lil', not 'lilGui'.

Assisted by Claude Code - Fable 5
Copy buttons export the results table: CSV with raw numbers (units in
the row labels), Markdown formatted for pasting into issue #402.

Fixes blank viewports on CNC files: the runner used one fixed settings
block (renderExtrusion on, renderTravel off), so pure-travel toolpaths
like mach3/easel built zero geometry. The bench now merges each
preset's display settings (renderTravel, camera, build volume, line
widths) over the demo defaults, sent identically to both versions.

Assisted by Claude Code - Fable 5
devtools/ is a home for internal development tooling, served alongside
the demo (new /devtools live-server mount) but not deployed with it.
The benchmark is its first tool, moved to devtools/benchmark/ and
reaching the demo's presets, gcodes, styles, and the local dist build
through root-absolute URLs. Includes a small landing page and a README
describing the area.

Assisted by Claude Code - Fable 5
Version discovery/import-map resolution, the sandboxed iframe runner
protocol, the 2.x/3.x preview compat shim, and demo-preset access move
into devtools/lib so upcoming tools can reuse them. The benchmark is
refactored on top with no behavior change.

Assisted by Claude Code - Fable 5
…or, visual-diff

All built on the shared devtools/lib infrastructure and listed on the
devtools landing page:

- memory-leak: repeated create/load/render/dispose cycles with a live
  heap chart and renderer.info counters; the verdict compares median
  steady-state heap between window halves, robust to GC sawtooth and
  allocator plateau jumps.
- streaming-equivalence: parses the same file whole vs as a chunked
  ReadableStream and diffs job stats plus a triangle fingerprint.
  Already caught a real divergence: parser.lineCount is one higher on
  the whole-string path for newline-terminated files (split('\n')
  counts a phantom empty trailing line; the streamed count is correct).
- parser-inspector: summary stats, command-type histogram, and a paged,
  filterable command table for presets or pasted gcode (3.x builds).
- visual-diff: renders one file in two versions with a pinned camera,
  captures both canvases in-frame, and pixel-diffs them with an
  adjustable threshold. local-vs-local diffs at exactly 0%.

Assisted by Claude Code - Fable 5
The demo stylesheet sets overflow: hidden on html/body for its
fullscreen layout, which every devtools page inherits; each page now
overrides it. Screenshots of every tool are committed for the README
and the PR description.

Assisted by Claude Code - Fable 5
Set breakpoints in the command table's gutter, step forward and back one
command at a time, or continue to the next breakpoint. The runner iframe
is now visible so the toolpath renders live as you step, and a State
panel shows the interpreter's State object with last-step changes
highlighted. Stepping executes command slices via interpreter.execute;
back-steps rebuild via preview.clear() plus one re-execution from the
start (163k benchy commands re-execute in ~40 ms).

Assisted by Claude Code - Fable 5
…Run to first

Summary and Command types collapse by default once a file loads. The
commands table drops pagination for a virtual-scrolled list that fills
the available viewport height (row height calibrated at runtime to
avoid sub-pixel drift across 163k rows). A new Run to first button
seeks past the slicer's metadata comment preamble to the first real
command.

Assisted by Claude Code - Fable 5
…loads

One localStorage key holds the selected version, the source (preset key
or pasted text, capped at 1 MB), and the breakpoint list. On open the
page restores the selections, auto-loads the file, and re-applies the
breakpoints. Corrupt or wrong-shaped stored values are removed and the
page starts fresh; failed saves are swallowed so persistence can never
break a session.

Assisted by Claude Code - Fable 5
Assisted by Claude Code - Fable 5
The tool outgrew its name: stepping, breakpoints, the live preview, and
State inspection are now the core, with parser inspection as supporting
panels. Directory moves to devtools/debugger, page and cards renamed,
localStorage key updated to match.

Assisted by Claude Code - Fable 5
Assisted by Claude Code - Fable 5
Addresses the CodeQL DOM-text-reinterpreted-as-HTML alerts on PR #453:
version labels and runner-reported metric values now go through an
escapeHtml helper before landing in innerHTML.

Assisted by Claude Code - Fable 5
…tools

The debugger gains a collapsed Instantiation settings editor: the full
constructor options object as editable JSON with live validation,
curated presets, Reset to default, and persistence on every edit.
Applying rebuilds the session on a fresh canvas while preserving
breakpoints and the current command; a constructor rejection falls back
to the last-known-good settings. The page is now full width and the
commands/panel columns are resizable by dragging a themed separator
(double-click resets, split persisted).

The whole suite gains a light/dark toggle on every page, built on
shared semantic theme variables in devtools/lib (dark matches the
previous look exactly). Preview canvases deliberately stay dark in both
themes so benchmark numbers and visual-diff captures are
theme-independent; the memory-leak chart re-reads its colors and
redraws on toggle.

Also hardens both debugger postMessage dispatches with Object.hasOwn
guards, addressing the CodeQL unvalidated-dynamic-method-call alert on
PR #453.

Assisted by Claude Code - Fable 5
Documents how to drive the devtools Debugger to produce before/after
screenshots for pull request descriptions, following the PR #439
pattern: synthetic gcode pasted into the debugger, identical framing on
both sides, images published on a throwaway pr<NUM>-assets branch, and
the two-column table layout for the PR body.

Assisted by Claude Code - Fable 5
Review feedback: the skill assumed the chrome-devtools MCP tools; it
now speaks in page-context JavaScript and element ids so any browser
automation (or a manual browser) works, scopes the MCP-only gotcha
accordingly, and drops the invented dark-theme convention in favor of
'use the same theme for both shots'.

Assisted by Claude Code - Fable 5
@sophiedeziel
sophiedeziel force-pushed the feature/benchmark-page branch from 3b0ada4 to 710e60b Compare September 7, 2026 16:36
parseGCode no longer fabricates a phantom empty line from a trailing
newline, and splitChunk now keeps the boundary newline in the complete
part so streamed chunk boundaries produce identical lines to a
whole-string parse (previously no-newline chunks inflated the count).
The ingestion-equivalence suite now asserts lineCount and command-count
equality across all four modes, including trailing-newline, no-newline,
and blank-line inputs.

executeCommands(commands) is now public: it executes already-parsed
commands into the current job without rendering and is re-entrant
(Interpreter.execute brackets each batch with resumeLastPath/
finishPath), which the devtools debugger relies on for stepping. Found
by the devtools streaming-equivalence checker on this branch.

Assisted by Claude Code - Fable 5
Measurement integrity: benchmark and visual-diff runners dispose their
preview once metrics/captures are taken, so a finished side's render
loop can no longer contaminate the other side's FPS and heap numbers
(failed runs also remove their iframe). The memory-leak tool now reads
geometry/texture counters after dispose() so they are real residual-
leak detectors instead of constants, and its verdict note fires on any
nonzero residual.

Robustness: the shared runner-frame dispatcher ignores malformed
messages instead of freezing until timeout, dark theme variables are
duplicated under :root so pages render styled even if theme.js fails,
the memory-leak table builds cells via textContent, and the copy
buttons restore their true labels on rapid clicks.

Correctness: the default baseline is now the newest published stable
(latestStable) instead of hardcoded 2.x, and presetSettings passes the
full preset through (minus app-only keys) with colors mapped to
extrusionColor, so tools render per-preset colors - visual-diff was
previously blind to color regressions by construction.

Dedup: shared page helpers (el/setStatus/escapeHtml/median/
runWithButton) in lib/page.js, the 2.x/3.x parse fork in
preview-compat.parseInto, the iframe scaffold exported as frameSrcdoc,
and the five-way page-shell CSS consolidated into theme.css with
visual-diff renamed onto the shared classes.

Assisted by Claude Code - Fable 5
Breakpoints survive a second Load (carried through the reset and
re-applied), split and settings edits made before the first load now
persist (storage split into a small hot key plus a source key, with
one-time migration from the old blob and per-key corruption handling),
and a post-load runner error resets the command list cleanly instead of
stranding it on dead placeholders.

Loading parses once via the standalone Parser and fills the job through
the new public executeCommands (with an interpreter fallback for older
builds), cutting benchy load from ~563 ms to ~138 ms. Steps always do a
full render (12-19 ms on benchy) so the preview exactly reflects the
current command; the renderProgressive step path could silently drop
segments appended to an already-rendered path. Row-height calibration
runs once per load instead of forcing a reflow on every render.

The page adopts the shared lib helpers and theme.css shell, and the
debugger-screenshots skill now names the Run to first button id.

Assisted by Claude Code - Fable 5
Screenshots retaken after the review fixes: presets render their real
colors now, command counts reflect the lineCount fix, and the streaming
checker shows the all-green verdict that the library fix earns. The
chart only reallocates its bitmap when the size changes and no longer
redraws when the result payload matches what the progress messages
already drew; the README describes the after-dispose residual counters.

Assisted by Claude Code - Fable 5
@remcoder remcoder added 3.1+ Targeted for the 3.1 release or later and removed 3.0 Targeted for the 3.0 release labels Sep 10, 2026
@remcoder remcoder mentioned this pull request Sep 14, 2026
1 of 3 tasks
@remcoder

remcoder commented Sep 14, 2026 •

Copy link
Copy Markdown
Member

in #524 I updated the benchmark page with a sweep mode, so it can run the benchmark against a number of builds.
might be a bit rough..

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

3.1+ Targeted for the 3.1 release or later marinating Left to simmer — needs more thought before acting meta a non-functional change, doesn't require a new release poc proof-of-concept. not intended te be merged into develop

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants