Skip to content

Hoist the XCOMP_BV read into MSHV vCPU creation - #1912

Merged
jprendes merged 1 commit into
hyperlight-dev:mainfrom
ludfjig:mshv-xcomp-bv-cache
Oct 9, 2026
Merged

jprendes merged 1 commit into
hyperlight-dev:mainfrom
ludfjig:mshv-xcomp-bv-cache

Conversation

@ludfjig

@ludfjig ludfjig commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Currently, every MSHV vCPU reset calls get_xsave to copy XCOMP_BV into the zeroed XSAVE area. This PR hoists that read into vCPU creation, and every reset reuses the value. This saves one hypercall per Sandbox::restore().

This is safe because XCOMP_BV depends only on the partition's XSAVE features, which are fixed at partition initialization and unaffected by guest state or snapshots.

Benchmarks

mshv3: AMD EPYC 7763, Linux 6.6 MSHV.

Benchmark Before After Change
snapshots/restore/default 87.9 µs 66.0 µs −25%
snapshots/restore/small 88.6 µs 66.5 µs −25%
snapshots/restore/medium 99.4 µs 77.9 µs −22%
guest_calls/call_with_restore/default 156.8 µs 133.9 µs −15%
guest_calls/call_with_restore/medium 165.5 µs 140.6 µs −15%
sandboxes/create_initialized/default 1.48 ms 1.53 ms +3%
sandboxes/sandbox_from_snapshot/default 598 µs 624 µs +4%

Each restore saves one get_xsave hypercall, about 22 µs. Each vCPU creation makes that hypercall once.

Copilot AI balanced review requested due to automatic review settings October 8, 2026 23:52
@ludfjig
ludfjig requested a review from danbugs as a code owner October 8, 2026 23:52
@ludfjig ludfjig added the kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. label Oct 8, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The focused optimization preserves XSAVE reset behavior while removing the repeated hypercall.

0 open findings

What changed in this PR

Caches MSHV’s partition-fixed XSAVE format during vCPU creation, avoiding a hypercall on every reset.

Changes:

  • Store XCOMP_BV once and reuse it during XSAVE reset.
  • Restrict XSAVE_MIN_SIZE to WHP.
  • Document the restore performance improvement.
File Description
src/​hyperlight_host/​src/​hypervisor/​virtual_machine/​mshv/​x86_64.rs Caches and reuses XCOMP_BV.
src/​hyperlight_host/​src/​hypervisor/​virtual_machine/​mod.rs Narrows the XSAVE minimum-size constant to WHP.
CHANGELOG.md Records the MSHV optimization.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@ludfjig ludfjig changed the title Read XCOMP_BV once when creating an MSHV vCPU Hoist the XCOMP_BV read into MSHV vCPU creation Oct 9, 2026
vCPU reset writes a zeroed XSAVE area that carries the partition's
XCOMP_BV. The hypervisor derives XCOMP_BV from the partition's XSAVE
features, which are fixed when the partition is created. MshvVm::new
reads it once, and every reset reuses it.

This removes one get_xsave hypercall from every restore. On mshv3
(AMD EPYC 7763, Linux 6.6 MSHV), snapshots/restore/default drops from
88 us to 66 us (-25%). Sandbox creation pays the read once, about +3%
for create_initialized/default.

Signed-off-by: Ludvig Liljenberg <4257730+ludfjig@users.noreply.github.com>
@ludfjig
ludfjig force-pushed the mshv-xcomp-bv-cache branch from 8b9317c to fc81cac Compare October 9, 2026 00:33
@hyperlight-gh-bot

Copy link
Copy Markdown

Benchmark Results

Measured commit: fc81cac1a8a5
Baseline commit: 164728cf08d0

kvm / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 816.92 ns (➖ 1.09x slower)
vec_bytes 574.73 ns (➖ 1.03x faster)
373.34 µs (➖ 1.00x faster)

payload_allocation

slot_pool_segmented
262144 536.13 ns (➖ 1.03x slower)
65536 143.94 ns (➖ 1.00x slower)

sandboxes

create_initialized_and_drop
medium 81.38 ms (➖ 1.01x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
7.75 ns (➖ 1.00x faster) 7.71 ns (➖ 1.00x faster) 7.76 ns (➖ 1.00x slower)

snapshot_files

load_snapshot_unverified
small 92.51 µs (➖ 1.00x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.33 µs (➖ 1.04x slower) 7.37 µs (➖ 1.05x slower)
65536 2.07 µs (➖ 1.07x slower) 1.92 µs (➖ 1.08x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 6.30 µs (➖ 1.00x faster) 6.24 µs (➖ 1.06x faster)
8192 1.09 µs (➖ 1.03x slower) 1.05 µs (➖ 1.00x faster)
262144 28.19 µs (➖ 1.05x slower)
kvm / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 722.00 ns (➖ 1.07x slower)
vec_bytes 507.90 ns (➖ 1.02x faster)
662.61 µs (➖ 1.03x faster)

payload_allocation

slot_pool_segmented
262144 500.37 ns (➖ 1.02x faster)
65536 135.52 ns (➖ 1.02x faster)

sandboxes

create_initialized_and_drop
medium 76.24 ms (➖ 1.04x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
7.04 ns (➖ 1.01x slower) 6.90 ns (➖ 1.01x faster) 6.90 ns (➖ 1.01x faster)

snapshot_files

load_snapshot_unverified
small 45.36 µs (➖ 1.02x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.53 µs (➖ 1.00x faster) 7.54 µs (➖ 1.00x faster)
65536 2.11 µs (➖ 1.01x faster) 2.12 µs (➖ 1.00x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.05 µs (➖ 1.02x faster) 7.11 µs (➖ 1.00x faster)
8192 729.98 ns (➖ 1.04x faster) 729.95 ns (➖ 1.00x slower)
262144 30.10 µs (➖ 1.02x faster)
mshv3 / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.06 µs (➖ 1.11x slower)
vec_bytes 714.37 ns (➖ 1.02x slower)
300.74 µs (➖ 1.09x slower)

payload_allocation

slot_pool_segmented
262144 782.93 ns (➖ 1.08x slower)
65536 193.73 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 61.24 ms (➖ 1.01x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.73 ns (➖ 1.00x faster) 10.03 ns (➖ 1.02x slower) 9.79 ns (➖ 1.01x slower)

snapshot_files

load_snapshot_unverified
small 86.57 µs (➖ 1.00x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 10.43 µs (➖ 1.12x faster) 9.69 µs (➖ 1.01x slower)
65536 2.46 µs (➖ 1.01x faster) 2.24 µs (➖ 1.12x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.44 µs (➖ 1.15x faster) 8.21 µs (➖ 1.18x faster)
8192 1.26 µs (➖ 1.08x faster) 1.27 µs (➖ 1.03x slower)
262144 35.45 µs (➖ 1.08x faster)
mshv3 / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.01 µs (➖ 1.09x slower)
vec_bytes 672.16 ns (➖ 1.03x slower)
770.09 µs (➖ 1.13x slower)

payload_allocation

slot_pool_segmented
262144 619.09 ns (➖ 1.00x faster)
65536 169.55 ns (➖ 1.03x slower)

sandboxes

create_initialized_and_drop
medium 67.08 ms (➖ 1.06x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.53 ns (➖ 1.01x faster) 8.75 ns (➖ 1.00x faster) 8.78 ns (➖ 1.00x slower)

snapshot_files

load_snapshot_unverified
small 44.30 µs (➖ 1.00x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 8.00 µs (➖ 1.01x slower) 7.96 µs (➖ 1.00x slower)
65536 2.31 µs (➖ 1.03x slower) 2.28 µs (➖ 1.00x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.62 µs (➖ 1.00x slower) 7.47 µs (➖ 1.01x faster)
8192 915.81 ns (➖ 1.02x faster) 920.84 ns (➖ 1.03x slower)
262144 39.83 µs (➖ 1.05x faster)
hyperv-ws2025 / amd (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.15 µs (➖ 1.02x slower)
vec_bytes 756.13 ns (➖ 1.17x faster)
2.20 ms (➖ 1.05x faster)

payload_allocation

slot_pool_segmented
262144 791.35 ns (➖ 1.00x slower)
65536 229.76 ns (➖ 1.01x faster)

sandboxes

create_initialized_and_drop
medium 96.70 ms (➖ 1.04x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.55 ns (➖ 1.15x faster) 10.63 ns (➖ 1.07x faster) 10.55 ns (➖ 1.05x faster)

snapshot_files

load_snapshot_unverified
small 682.74 µs (➖ 1.13x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 8.68 µs (➖ 1.01x slower) 8.70 µs (➖ 1.00x slower)
65536 2.56 µs (➖ 1.01x slower) 2.55 µs (➖ 1.00x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 9.17 µs (➖ 1.07x slower) 8.86 µs (➖ 1.00x slower)
8192 1.37 µs (➖ 1.07x faster) 1.38 µs (➖ 1.00x faster)
262144 38.10 µs (➖ 1.02x slower)
hyperv-ws2025 / intel (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.21 µs (➖ 1.00x faster)
vec_bytes 786.21 ns (➖ 1.03x faster)
3.27 ms (➖ 1.04x slower)

payload_allocation

slot_pool_segmented
262144 742.96 ns (➖ 1.00x faster)
65536 215.31 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 105.30 ms (➖ 1.09x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.92 ns (➖ 1.06x slower) 10.07 ns (➖ 1.01x slower) 10.63 ns (➖ 1.01x slower)

snapshot_files

load_snapshot_unverified
small 543.02 µs (➖ 1.19x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.79 µs (➖ 1.02x slower) 7.86 µs (➖ 1.01x slower)
65536 2.35 µs (➖ 1.02x slower) 2.34 µs (➖ 1.00x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.43 µs (➖ 1.02x faster) 8.51 µs (➖ 1.01x slower)
8192 1.14 µs (➖ 1.11x faster) 1.11 µs (➖ 1.01x slower)
262144 47.96 µs (➖ 1.04x slower)

Reported by cargo ci bench-report --candidate run:37865422069 --config-file bench_report.toml.

@ludfjig ludfjig added the ready-for-review PR is ready for (re-)review label Oct 9, 2026
@jprendes
jprendes merged commit 9f7cfba into hyperlight-dev:main Oct 9, 2026
63 checks passed
@github-actions github-actions Bot removed the ready-for-review PR is ready for (re-)review label Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kind/enhancement For PRs adding features, improving functionality, docs, tests, etc.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants