Add a dev tools area with benchmark, memory-leak, streaming, parser, and visual-diff tools - #453
Draft
sophiedeziel wants to merge 22 commits into
Draft
sophiedeziel wants to merge 22 commits into
sophiedeziel wants to merge 22 commits into
Conversation
|
Visit the preview URL for this PR (updated for commit 62a50b2): https://gcode-preview--pr453-feature-benchmark-pa-uqxm5awh.web.app (expires Wed, 07 Oct 2026 17:37:17 GMT) 🔥 via Firebase Hosting GitHub Action 🌎 Sign: 59bd114ae4847b32c2bba0b68620b9069a3e3531 |
sophiedeziel
added a commit
that referenced
this pull request
Sep 6, 2026
Addresses the CodeQL DOM-text-reinterpreted-as-HTML alerts on PR #453: version labels and runner-reported metric values now go through an escapeHtml helper before landing in innerHTML. Assisted by Claude Code - Fable 5
sophiedeziel
added a commit
that referenced
this pull request
Sep 6, 2026
…tools The debugger gains a collapsed Instantiation settings editor: the full constructor options object as editable JSON with live validation, curated presets, Reset to default, and persistence on every edit. Applying rebuilds the session on a fresh canvas while preserving breakpoints and the current command; a constructor rejection falls back to the last-known-good settings. The page is now full width and the commands/panel columns are resizable by dragging a themed separator (double-click resets, split persisted). The whole suite gains a light/dark toggle on every page, built on shared semantic theme variables in devtools/lib (dark matches the previous look exactly). Preview canvases deliberately stay dark in both themes so benchmark numbers and visual-diff captures are theme-independent; the memory-leak chart re-reads its colors and redraws on toggle. Also hardens both debugger postMessage dispatches with Object.hasOwn guards, addressing the CodeQL unvalidated-dynamic-method-call alert on PR #453. Assisted by Claude Code - Fable 5
7 of 10 tasks
benchmark.html compares two library versions (any published npm version via jsDelivr, or the local dist build) on one of the demo gcode files. Each version runs sequentially in a fresh sandboxed iframe with its own import map, resolving the three/lil-gui versions that release was published against, so runs cannot contaminate each other. Measures the #402 release-gate metrics — parse time, time to first render, FPS while orbiting, peak JS heap — plus geometry build time, triangles and draw calls, with interleaved runs and a median-based comparison table. Assisted by Claude Code - Fable 5
Silences rollup's 'Missing global variable name' warning and fixes the guessed global: lil-gui's UMD build exposes 'lil', not 'lilGui'. Assisted by Claude Code - Fable 5
Copy buttons export the results table: CSV with raw numbers (units in the row labels), Markdown formatted for pasting into issue #402. Fixes blank viewports on CNC files: the runner used one fixed settings block (renderExtrusion on, renderTravel off), so pure-travel toolpaths like mach3/easel built zero geometry. The bench now merges each preset's display settings (renderTravel, camera, build volume, line widths) over the demo defaults, sent identically to both versions. Assisted by Claude Code - Fable 5
devtools/ is a home for internal development tooling, served alongside the demo (new /devtools live-server mount) but not deployed with it. The benchmark is its first tool, moved to devtools/benchmark/ and reaching the demo's presets, gcodes, styles, and the local dist build through root-absolute URLs. Includes a small landing page and a README describing the area. Assisted by Claude Code - Fable 5
Version discovery/import-map resolution, the sandboxed iframe runner protocol, the 2.x/3.x preview compat shim, and demo-preset access move into devtools/lib so upcoming tools can reuse them. The benchmark is refactored on top with no behavior change. Assisted by Claude Code - Fable 5
…or, visual-diff
All built on the shared devtools/lib infrastructure and listed on the
devtools landing page:
- memory-leak: repeated create/load/render/dispose cycles with a live
heap chart and renderer.info counters; the verdict compares median
steady-state heap between window halves, robust to GC sawtooth and
allocator plateau jumps.
- streaming-equivalence: parses the same file whole vs as a chunked
ReadableStream and diffs job stats plus a triangle fingerprint.
Already caught a real divergence: parser.lineCount is one higher on
the whole-string path for newline-terminated files (split('\n')
counts a phantom empty trailing line; the streamed count is correct).
- parser-inspector: summary stats, command-type histogram, and a paged,
filterable command table for presets or pasted gcode (3.x builds).
- visual-diff: renders one file in two versions with a pinned camera,
captures both canvases in-frame, and pixel-diffs them with an
adjustable threshold. local-vs-local diffs at exactly 0%.
Assisted by Claude Code - Fable 5
The demo stylesheet sets overflow: hidden on html/body for its fullscreen layout, which every devtools page inherits; each page now overrides it. Screenshots of every tool are committed for the README and the PR description. Assisted by Claude Code - Fable 5
Set breakpoints in the command table's gutter, step forward and back one command at a time, or continue to the next breakpoint. The runner iframe is now visible so the toolpath renders live as you step, and a State panel shows the interpreter's State object with last-step changes highlighted. Stepping executes command slices via interpreter.execute; back-steps rebuild via preview.clear() plus one re-execution from the start (163k benchy commands re-execute in ~40 ms). Assisted by Claude Code - Fable 5
…Run to first Summary and Command types collapse by default once a file loads. The commands table drops pagination for a virtual-scrolled list that fills the available viewport height (row height calibrated at runtime to avoid sub-pixel drift across 163k rows). A new Run to first button seeks past the slicer's metadata comment preamble to the first real command. Assisted by Claude Code - Fable 5
…loads One localStorage key holds the selected version, the source (preset key or pasted text, capped at 1 MB), and the breakpoint list. On open the page restores the selections, auto-loads the file, and re-applies the breakpoints. Corrupt or wrong-shaped stored values are removed and the page starts fresh; failed saves are swallowed so persistence can never break a session. Assisted by Claude Code - Fable 5
Assisted by Claude Code - Fable 5
The tool outgrew its name: stepping, breakpoints, the live preview, and State inspection are now the core, with parser inspection as supporting panels. Directory moves to devtools/debugger, page and cards renamed, localStorage key updated to match. Assisted by Claude Code - Fable 5
Assisted by Claude Code - Fable 5
Assisted by Claude Code - Fable 5
Addresses the CodeQL DOM-text-reinterpreted-as-HTML alerts on PR #453: version labels and runner-reported metric values now go through an escapeHtml helper before landing in innerHTML. Assisted by Claude Code - Fable 5
…tools The debugger gains a collapsed Instantiation settings editor: the full constructor options object as editable JSON with live validation, curated presets, Reset to default, and persistence on every edit. Applying rebuilds the session on a fresh canvas while preserving breakpoints and the current command; a constructor rejection falls back to the last-known-good settings. The page is now full width and the commands/panel columns are resizable by dragging a themed separator (double-click resets, split persisted). The whole suite gains a light/dark toggle on every page, built on shared semantic theme variables in devtools/lib (dark matches the previous look exactly). Preview canvases deliberately stay dark in both themes so benchmark numbers and visual-diff captures are theme-independent; the memory-leak chart re-reads its colors and redraws on toggle. Also hardens both debugger postMessage dispatches with Object.hasOwn guards, addressing the CodeQL unvalidated-dynamic-method-call alert on PR #453. Assisted by Claude Code - Fable 5
Documents how to drive the devtools Debugger to produce before/after screenshots for pull request descriptions, following the PR #439 pattern: synthetic gcode pasted into the debugger, identical framing on both sides, images published on a throwaway pr<NUM>-assets branch, and the two-column table layout for the PR body. Assisted by Claude Code - Fable 5
Review feedback: the skill assumed the chrome-devtools MCP tools; it now speaks in page-context JavaScript and element ids so any browser automation (or a manual browser) works, scopes the MCP-only gotcha accordingly, and drops the invented dark-theme convention in favor of 'use the same theme for both shots'. Assisted by Claude Code - Fable 5
sophiedeziel
force-pushed
the
feature/benchmark-page
branch
from
September 7, 2026 16:36
3b0ada4 to
710e60b
Compare
parseGCode no longer fabricates a phantom empty line from a trailing newline, and splitChunk now keeps the boundary newline in the complete part so streamed chunk boundaries produce identical lines to a whole-string parse (previously no-newline chunks inflated the count). The ingestion-equivalence suite now asserts lineCount and command-count equality across all four modes, including trailing-newline, no-newline, and blank-line inputs. executeCommands(commands) is now public: it executes already-parsed commands into the current job without rendering and is re-entrant (Interpreter.execute brackets each batch with resumeLastPath/ finishPath), which the devtools debugger relies on for stepping. Found by the devtools streaming-equivalence checker on this branch. Assisted by Claude Code - Fable 5
Measurement integrity: benchmark and visual-diff runners dispose their preview once metrics/captures are taken, so a finished side's render loop can no longer contaminate the other side's FPS and heap numbers (failed runs also remove their iframe). The memory-leak tool now reads geometry/texture counters after dispose() so they are real residual- leak detectors instead of constants, and its verdict note fires on any nonzero residual. Robustness: the shared runner-frame dispatcher ignores malformed messages instead of freezing until timeout, dark theme variables are duplicated under :root so pages render styled even if theme.js fails, the memory-leak table builds cells via textContent, and the copy buttons restore their true labels on rapid clicks. Correctness: the default baseline is now the newest published stable (latestStable) instead of hardcoded 2.x, and presetSettings passes the full preset through (minus app-only keys) with colors mapped to extrusionColor, so tools render per-preset colors - visual-diff was previously blind to color regressions by construction. Dedup: shared page helpers (el/setStatus/escapeHtml/median/ runWithButton) in lib/page.js, the 2.x/3.x parse fork in preview-compat.parseInto, the iframe scaffold exported as frameSrcdoc, and the five-way page-shell CSS consolidated into theme.css with visual-diff renamed onto the shared classes. Assisted by Claude Code - Fable 5
Breakpoints survive a second Load (carried through the reset and re-applied), split and settings edits made before the first load now persist (storage split into a small hot key plus a source key, with one-time migration from the old blob and per-key corruption handling), and a post-load runner error resets the command list cleanly instead of stranding it on dead placeholders. Loading parses once via the standalone Parser and fills the job through the new public executeCommands (with an interpreter fallback for older builds), cutting benchy load from ~563 ms to ~138 ms. Steps always do a full render (12-19 ms on benchy) so the preview exactly reflects the current command; the renderProgressive step path could silently drop segments appended to an already-rendered path. Row-height calibration runs once per load instead of forcing a reflow on every render. The page adopts the shared lib helpers and theme.css shell, and the debugger-screenshots skill now names the Run to first button id. Assisted by Claude Code - Fable 5
Screenshots retaken after the review fixes: presets render their real colors now, command counts reflect the lineCount fix, and the streaming checker shows the all-green verdict that the library fix earns. The chart only reallocates its bitmap when the size changes and no longer redraws when the result payload matches what the progress messages already drew; the README describes the after-dispose residual counters. Assisted by Claude Code - Fable 5
Member
|
in #524 I updated the benchmark page with a sweep mode, so it can run the benchmark against a number of builds. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Disclaimer
As this kind of toolset is what I always dreamed of (for this project), this is a POC that is 100% vibe-coded and probably mostly slop.
Code is cheap, the tools are still valuable.
Summary
This PR adds a
devtools/area for internal development tooling, served alongside the demo by the existing dev server but not deployed with it. It ships five tools aimed at the 3.0.0 release gates in #402, built on a small shared infrastructure layer (devtools/lib/) that handles version discovery on jsDelivr, per-version import maps (each release loads the three/lil-gui versions it was published against), sandboxed iframe runners, and a 2.x/3.x API compat shim.Run
npm run devand open http://localhost:8080/devtools/. Every page has a light/dark toggle top right, persisted across pages and reloads (previews and captures stay dark in both themes so numbers and pixel diffs never depend on the page theme).Debugger
Steps through a gcode file like a debugger. Click a row's gutter to set a breakpoint, then step forward or back one command at a time, continue to the next breakpoint, or use Run to first to jump past the slicer's metadata comment preamble (on Benchy it lands on command 1,149, the first M73 after PrusaSlicer's comments). The preview renders the toolpath live as you step, and a State panel shows the interpreter's State object (position, positionShift, tool, units, isHomed) with the keys that changed on the last step highlighted. Stepping back re-executes from the start, which stays fast: benchy's 163,474 commands re-execute in about 40 ms. The loaded file and breakpoints survive page reloads via localStorage, with corrupt stored values discarded safely. A collapsed Instantiation settings editor exposes the constructor options as editable JSON with live validation, curated presets, Reset to default, and persistence on every edit; applying rebuilds the session while keeping breakpoints and position, and a rejected settings object falls back to the last-known-good ones. The page is full width and the two columns are resizable by dragging the separator.
It doubles as a parser inspector: summary stats and a command-type histogram (both collapsed by default after loading), plus a filterable command list that virtual-scrolls through all of Benchy's 163k commands with no pagination. Full inspection targets 3.x builds since 2.x does not export the parser; older versions degrade to a friendly notice.
Two more findings came out of building it: constructing a preview without
buildVolumethrows in 3.x (the SceneManager constructor readsthis.buildVolume.xunconditionally), and Mach3-style bareS900/F140lines currently parse as gcodess900/f140.Benchmark
Compares two library versions head to head on one of the demo gcode files. Any published npm version can face any other, or the local
dist/build. Each version runs sequentially in a fresh sandboxed iframe, runs are interleaved A/B/A/B, and the table reports medians for parse time, geometry build and render time, time to first render, FPS while orbiting, peak JS heap, triangles, and draw calls. Results can be copied as CSV or as Markdown ready to paste into #402.Memory-leak tester
Runs repeated create, load, render, dispose cycles against a fresh canvas each time to mimic a framework remount, then charts the JS heap per cycle along with the
renderer.infogeometry and texture counts read after dispose() — any nonzero residual means dispose leaked GPU resources. The verdict compares the median heap between the first and second half of the steady-state cycles, which stays calm through GC sawtooth and allocator plateau jumps instead of crying wolf on healthy builds. The local build currently passes with a flat heap and flat counters.Streaming equivalence checker
Parses the same file twice in one version, once as a whole string and once as a ReadableStream chunked at a configurable size, then diffs the parser and job stats plus a rendered-triangle fingerprint. Small chunk sizes maximize the odds of splitting a command across a chunk boundary, which is exactly where streaming bugs live.
It caught a real bug on its first run:
parser.lineCountwas one higher on the whole-string path for every newline-terminated file (a phantom empty trailing line fromsplit), and digging in revealed the streamed path also fabricated lines on chunks without newlines. Both sides are now fixed insrc/on this PR, the ingestion-equivalence suite asserts line and command counts across all four modes, and the checker reports all metrics green.Visual diff
Renders the same file in two versions with a pinned camera and target, captures both canvases in the same frame (required on 3.x, which uses
preserveDrawingBuffer: false), and pixel-diffs the captures with an adjustable threshold. Changed pixels light up magenta over a dimmed grayscale base. Local vs local diffs at exactly 0.00%, and 2.18.0 vs local on Benchy shows a real 9.09% difference in scene composition relative to the origin.Also in this PR
The rollup UMD output was missing a global mapping for
lil-gui, so rollup warned on every build and guessedlilGui, which is not the global that lil-gui's UMD build actually exposes (lil). One line inrollup.config.mjsfixes both the warning and the broken guess.The CodeQL alerts raised on this PR are addressed: interpolated values in the benchmark and streaming-equivalence tables are HTML-escaped, and the debugger's postMessage dispatches guard message types with Object.hasOwn.
A
debugger-screenshotsskill (in.agents/) documents how to capture before/after PR proof shots from the Debugger with custom gcode, following the pattern PR #439 established.A local multi-angle code review of the branch was applied before marking ready: benchmark and visual-diff runners now dispose their previews so a finished side cannot contaminate the other side's measurements, breakpoints survive re-loads, pre-load edits persist, the debugger parses once and steps with exact full renders via the new public
executeCommandsAPI, the default baseline tracks the newest published stable instead of hardcoded 2.x, presets pass their real colors through, theme variables fall back to dark without JS, and shared helpers/CSS were extracted intodevtools/lib.Related to #402.
Assisted by Claude Code - Fable 5