Skip to content

perf(serde): Compile per-shape JSON and XML serde plans with caching - #3363

Open
rohangavankar wants to merge 37 commits into
aws:masterfrom
rohangavankar:serde-cache
Open

rohangavankar wants to merge 37 commits into
aws:masterfrom
rohangavankar:serde-cache

Conversation

@rohangavankar

@rohangavankar rohangavankar commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Compile per-shape JSON and XML serialize/deserialize instructions once, cache them on
the shape, and reuse them on later calls. No public API change. Wire output stays
byte-identical. Serialization runs in the build step and parsing runs as the handler
result resolves.

This is the JSON + XML half of the serde-plan work (shared ShapePlanCache foundation

  • JSON encode/decode + XML encode/decode). HTTP bindings and Query serialization are a
    separate workstream and are not in this PR.

What's included

  • Foundation: ShapePlanCache slot registry. Plan cache on AbstractModel / ShapeMap
    keyed off a monotonic generation, with invalidation on shape mutation. Operation
    context params use the same generation check.
  • JSON: JsonEncodePlanProvider / JsonDecodePlanProvider, wired into JsonBody and
    JsonParser.
  • XML: XmlEncodePlanProvider / XmlDecodePlanProvider, wired into XmlBody and
    XmlParser. AWS Query response parsing also uses XmlParser.

Correctness

  • tests/Api: 2,251 tests pass. tests/S3 and tests/S3Control fail the same set of
    tests as the base commit when run without credentials on the same host, and no test
    fails only on this branch.
  • XmlWireCompatibilityTest pins XML encode and decode output for root-name precedence,
    namespaces, multiple attributes, empty, flattened and nested collections, blobs,
    timestamps, special floats and unions. All 35 cases also pass against
    upstream master without the plan code.
  • PHPCS 3.13.4 with phpcs.xml.dist: 0 errors across the 22 changed src/ files.

Performance

Baseline is the PR's base commit 4a8bc490c (3.398.3). Candidate is the PR head.
Host: x86 m7i.xlarge, PHP 8.1.34, OPcache on, Xdebug not loaded. Baseline and
candidate ran alternately on the same host, two runs each. Run-to-run drift of
identical code was under 0.3%.

Whole operation (serde corpus, serde_benchmark.php, 10,000 iterations)

Weighted p50 change. JIT on = opcache.jit=tracing, 64M buffer.

Group Cases JIT on p50 JIT on p90 JIT off p50
awsJson1 deserialize 6 -49.7% -48.0% -45.7%
awsJson1 serialize 12 -24.9% -24.4% -24.3%
awsQuery deserialize 3 -4.4% -4.7% -6.8%
awsQuery serialize 6 -1.1% -1.2% +0.5%
restJson1 deserialize 7 -1.0% -0.8% -0.9%
restJson1 serialize 1 +1.0% +0.2% -0.8%
restXml deserialize 7 -2.3% -2.4% -2.7%
restXml serialize 1 -0.6% +0.1% -0.4%
rpcv2Cbor deserialize 6 -0.4% -0.4% +0.3%
rpcv2Cbor serialize 11 +0.6% +0.3% -0.2%

Query serialization and CBOR have no plans in this PR and act as controls. The 10
minimal _Baseline cases moved between -6.5% and +1.0%. REST-JSON and REST-XML
whole-operation time is dominated by HTTP-binding work, which is the separate workstream.

Body codec only (focused benchmarks, 50,000 iterations, --items=24)

Warm p50 change, SmallNoList / NestedLarge / MapHeavy:

Direction JIT on JIT off
JSON encode -46% / -38% / -33% -47% / -35% / -33%
JSON decode -15% / -42% / -45% -16% / -32% / -40%
XML encode -40% / -45% / -42% -40% / -46% / -46%
XML decode -11% / -20% / -15% -13% / -22% / -19%

Output hashes match the base commit in every case.

First use and retained memory (JIT off, median of 2 runs)

First call on a fresh model, which includes plan compilation:

Direction SmallNoList NestedLarge MapHeavy
JSON encode 3.95 -> 7.16 us 60.6 -> 48.7 us 14.5 -> 12.7 us
JSON decode 14.9 -> 21.6 us 50.0 -> 41.0 us 13.5 -> 10.5 us
XML encode 16.5 -> 21.3 us 365.6 -> 219.6 us 172.1 -> 113.2 us
XML decode 29.5 -> 36.9 us 182.0 -> 159.2 us 78.6 -> 64.5 us

Small payloads pay 3-7 us once to compile their plans. Larger payloads are faster on the
first call too. First use is a single sample per run, so treat it as approximate.

Heap retained after the first call, per model graph:

Graph SmallNoList NestedLarge MapHeavy
JSON (both directions) 0.9 -> 6.0 KB 1.8 -> 11.7 KB 0.5 -> 3.9 KB
XML encode 0.4 -> 4.9 KB 1.9 -> 8.2 KB 0.5 -> 2.8 KB
XML decode 1.6 -> 5.8 KB 1.8 -> 9.7 KB 0.5 -> 3.0 KB

Note on the XML decode COERCE scalar tag

The XML decode plan precomputes a scalar coercion kind (COERCE_STRING/INT/FLOAT) rather
than re-reading $shape['type'] per leaf element. Without it, MapHeavy decode was
slower than the original parser on x86 (two runs, same baseline):

MapHeavy decode p50 vs original Without COERCE With COERCE
Run 1 +7.2% -18.0%
Run 2 +11.7% -17.3%

Design-doc conformance

Follows the JSON and XML serde-plan design. The XML decode descriptor carries a few
fields beyond the original proposal (M_ATTRKEY, M_TSFORMAT, M_COERCE) to remove
per-element model reads from the hot path. The COERCE ablation above shows the effect.

rohangavankar and others added 22 commits September 17, 2026 10:52
Add the shared cache foundation for serde plans (PR 0). Serializers and
parsers can compile protocol instructions once per shape or operation and
reuse them, replacing repeated model-array interpretation on every call.

- ShapePlanCache: central slot registry (protocol + direction).
- AbstractModel::getCachedPlan/setCachedPlan: per-slot plan storage.
- ShapeMap: monotonic graph generation for cross-object invalidation.
- clearResolvedModelCache hook on StructureShape, ListShape, MapShape,
  Operation, and Service; invoked on offsetSet/offsetUnset before the
  generation bump so plans derived from mutated shapes rebuild.
- Service::setDefinition clears the operation cache and getOperation now
  returns stable instances so operation-level plans can warm.

No protocol providers and no public API changes. Plan accessors are internal.
Add the json-serde-plans spec (requirements, design, tasks) grounded in the
committed plan-cache foundation and the docs/serde design docs. Add a focused
benchmark/json-serde.php that compares the legacy JsonBody::format() path
against the compiled-plan path in one process, reporting first-use, p50, and
p90 on small, nested-large, and map-heavy payloads.

Local benchmark numbers are dev-grade. Merge evidence requires the x86
m7i.xlarge runbook flow.
Replace JsonBody's per-request model reads with a compiled JsonEncodePlan
cached per shape via ShapePlanCache::JSON_ENCODE. JsonShapeType maps model
strings to integer tags, and JsonEncodePlanProvider compiles one shape level
lazily, fetching child plans when traversal reaches composite members. build()
runs the plan path; buildLegacy() retains the old format() path so the
benchmark can compare both in one process.

Wire output is byte-identical: all tests/Api/Serializer + ComplianceTest pass
(1252 tests). Local dev-grade encode numbers (200k iters, Xdebug off): warm p50
improves 27-47% (NestedLarge -26.8%, MapHeavy -36.8%, SmallNoList -46.7%).
Small scalar-only payloads pay a ~2us first-use plan-compile cost. The
PutItemRequest_Nested_L regression from the design doc did not reproduce here.

Numbers in benchmark/results/json-encode-2026-09-21.md. Merge evidence still
requires the x86 m7i.xlarge runbook flow.
Replace JsonParser's per-response model reads with a compiled JsonDecodePlan
cached per shape via ShapePlanCache::JSON_DECODE. Structure members are stored
as an ordered list to preserve V3 result order; the plan carries a union flag
for the Unknown fallback. JsonDecodePlanProvider compiles one shape level
lazily, fetching child plans when traversal reaches composite members. parse()
runs the plan path; parseLegacy() retains the old path for the benchmark.

Decode timestamp default stays null (DateTimeResult), null map values are
skipped, and blob base64_decode is preserved. Output is identical: all
tests/Api/Parser pass (677 tests). Extend benchmark/json-serde.php with
--direction=decode.

Local dev-grade decode numbers (200k iters, Xdebug off): warm p50 improves
21-51% (MapHeavy -50.7%, NestedLarge -31.5%, SmallNoList -21.0%), no
small-payload first-use regression. Numbers in
benchmark/results/json-decode-2026-09-21.md. Merge evidence still requires the
x86 m7i.xlarge runbook flow.
Add --items option to the focused harness to match the runbook JSON command
(--items=N sizes collections; 0 exercises the small scalar-only path). Capture
x86 before/after encode numbers on benchmark-v4 (m7i.xlarge, PHP 8.1.34, JIT
tracing, OPcache on).

Warm p50 improves 27-52% across SmallNoList, NestedLarge, and MapHeavy at both
items=50 and items=0. Small/empty payloads pay a few-microsecond first-use
plan-compile cost that repays within a few warm calls. Wire output identical.
The PutItemRequest_Nested_L regression did not reproduce on x86 (nested-large
-37.6% p50). Full data in benchmark/results/json-encode-x86-2026-09-21.md.
Capture x86 before/after decode numbers on benchmark-v4 (m7i.xlarge, PHP
8.1.34, JIT tracing, OPcache on), comparing JsonParser::parseLegacy against the
plan path.

Warm p50 improves 13-44% across SmallNoList, NestedLarge, and MapHeavy at both
items=50 and items=0. Empty-collection first-use pays a few-microsecond
plan-compile cost. A one-off MapHeavy first-use spike was confirmed a
measurement artifact across three isolated re-runs. Decoded output identical.
Full data in benchmark/results/json-decode-x86-2026-09-21.md.
Add --mode=memory to the focused harness (retained plan heap, both directions
on one graph) and capture the remaining evidence the design docs require.

Retained memory: bounded by the model, not payload size (identical at items=50
and items=2000). NestedLarge ~11.7 KB, MapHeavy ~3.9 KB per graph.

Large payload (items=2000, ~3 ms): encode warm p50 -37.6%, decode -40.7%,
output identical.

Full compliance corpus (scripts/benchmarks/serde_benchmark.php, 70 cases, all 5
protocols, foundation baseline vs serde-cache candidate): JSON RPC weighted p50
-35.9%; REST-JSON flat (HTTP-binding bound); REST-XML/Query/CBOR controls within
+-0.3% (no regression). Raw result JSON preserved under x86-compliance-corpus/.

Docs: json-memory-and-large-x86-2026-09-22.md,
compliance-corpus-x86-2026-09-22.md.
Remove the JSON serde benchmark result docs and raw result JSON from the SDK
tree and gitignore benchmark/results/. These are internal evidence and live in
php-fork-dev, not the public aws-sdk-php package. Only benchmark/json-serde.php
(the harness code, which needs the SDK autoloader) stays in the SDK.
Add the focused unit tests the design requires for the JSON serde plans:

- JsonEncodePlanProviderTest: structure/list/map/timestamp/blob/document
  descriptor variants, wire-name derivation, encode timestamp default
  (unixTimestamp), lazy child Shape retention, and per-shape caching in the
  JSON_ENCODE slot.
- JsonDecodePlanProviderTest: modeled member ordering, decode timestamp default
  (null), union flag, list/map value descriptors, and JSON_DECODE slot caching
  without occupying the encode slot.
- JsonPlanInvalidationTest: locationName mutation and members replacement
  rebuild both encode and decode plans; stable generation reuses the cached
  plan; mock-without-ShapeMap stays safe.

18 tests, 60 assertions, all green.
Add a decode test for a root-level timestamp shape (not a struct member),
covering the JsonShapeType::TIMESTAMP branch in JsonDecodePlanProvider. Brings
the JSON plan provider and shape-type classes to 100% line coverage.
Extend the invalidation tests to the remaining new foundation paths:

- offsetUnset on a shape trait bumps the generation and rebuilds the plan
  (the contract covers unset as well as set).
- ListShape and MapShape drop their resolved child (member / value) on
  mutation, so the rebuilt plan reflects the new child definition.

Covers AbstractModel::offsetUnset and the clearResolvedModelCache overrides on
StructureShape, ListShape, and MapShape. The base AbstractModel hook stays
uncovered by design (abstract; every concrete shape overrides it).
Replace XmlBody's per-request model reads with a compiled XmlEncodePlan cached
per shape via ShapePlanCache::XML_ENCODE (slot 4). XmlShapeType maps model
strings to integer tags; XmlEncodePlanProvider precomputes per shape the
namespace attribute, structure attribute-vs-element partition and element-name
resolution (honoring locationNameAtStructureLevel), flattened decisions, and
list/map element naming. XML timestamp default stays iso8601 (not JSON's
unixTimestamp). build() runs the plan path; buildLegacy() retains the old path
for the benchmark.

XMLWriter output is byte-identical: all tests/Api/Serializer pass (748 tests)
including RestXmlSerializerTest escaping and rest-xml ComplianceTest. 9 XML plan
unit tests; provider and shape-type at 100% line coverage.

Local x86 warm p50 improves 37-45% (SmallNoList -37.4%, NestedLarge -45.1%,
MapHeavy -42.8%), output identical. Add benchmark/xml-serde.php. Report in
php-fork-dev serde-plans results.
Replace XmlParser's per-response model reads with a compiled XmlDecodePlan
cached per shape via ShapePlanCache::XML_DECODE (slot 5). XmlDecodePlanProvider
precomputes per shape: element read-names (honoring the getOriginalDefinition
structure-level locationName special case), attribute key + namespace for
attribute members, flattened decisions, list/map element naming, union status,
timestamp decode default (null, not encode's iso8601), and a scalar coercion
tag. parse() runs the plan path; parseLegacy() retains the old dispatch for the
benchmark. Also benefits AWS Query response parsing (shared XmlParser).

A precomputed scalar coercion tag (COERCE_STRING/INT/FLOAT) removes a per-element
$shape['type'] read from the leaf path; without it the map-heavy case regressed
+6.8%, with it decode gains -10% to -23% warm p50 on x86.

Result arrays byte-identical: all 677 parser + 63 error-parser + 118 S3 parser
tests pass. 18 XML plan unit tests; XmlDecodePlanProvider at 100% coverage. Add
--direction=decode to benchmark/xml-serde.php.

Deviations from XML-Serde-Plans.md (extra descriptor fields, COERCE tag, doc not
yet updated) are documented in the XML report. Keep as a separate CR after XML
encode per the design.
Rename AbstractModel::getCachedPlan/setCachedPlan to getSerdePlan/cacheSerdePlan
to match the names specified in docs/serde/architecture.md. Updates all four
plan providers (JSON/XML encode/decode) and the serde tests. No behavior change;
all 1425 serializer + parser tests pass.
Restore two comments (XML node name extraction, locationName-from-definition
check) inadvertently dropped during the XmlParser plan-path rewrite. Comments
only; no behavior change. The legacy parse methods now match upstream.
Precompute the root element name (three-level precedence: ShapeMap original
locationName, resolved locationName, shape name) into XmlEncodePlan::rootName in
the provider, and read it in XmlBody::build(). This removes the per-request
determineRootElementName() metadata inspection from the plan path, satisfying
the XML design doc requirement that the runtime serializer not inspect shape
metadata to open the document root. determineRootElementName() is retained for
the legacy benchmark path.

Output byte-identical: all 748 serializer tests pass.
Add benchmark/ to .gitattributes export-ignore, matching how tests/, docs/, and
features/ are already handled. The serde benchmark harnesses live in the repo
for developers (and to sync to the x86 instance per the runbook) but should not
ship in the Composer-distributed package to customers.
@rohangavankar
rohangavankar requested a review from a team as a code owner September 29, 2026 19:52
Run phpcbf with phpcs.xml.dist on the files this change modifies.
Fixes control-structure spacing, missing instantiation parentheses,
foreach keyword spacing and one indentation error. Formatting only,
no behavior change.
XmlParser and XmlBody resolve their per-type methods at runtime by concatenating the shape type, so the snake_case names are part of the dispatch contract and cannot be renamed without changing that lookup. The names predate this change. PHPCS reports them only because the check scans whole touched files. Annotate each method with a targeted phpcs:ignore.

@stobrien89 stobrien89 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! left a few comments where some encoding/decoding results appeared to change and a minor style comment.

Also, are buildLegacy()/parseLegacy() still needed?

Comment thread src/Api/Parser/JsonParser.php
Comment thread src/Api/Serializer/XmlBody.php Outdated
Comment thread src/Api/Operation.php
Comment thread src/Api/Parser/XmlParser.php Outdated
The plan-based list loop passed each element to parseByType() without
the null check the legacy parser applied first, so a sparse list of
blobs decoded null as "" and a null timestamp as the epoch. Return null
at the top of parseByType() for every shape type, matching legacy.
The legacy XmlBody honors xmlAttribute only in add_string, so numeric,
boolean and timestamp members marked xmlAttribute were written as child
elements (still ordered first). The plan path wrote every scalar marked
xmlAttribute as an attribute, which changed the wire output.

Precompute the attribute name at plan-compile time and set it only for
string shapes, using the member locationName as legacy does. Apply the
same rule to list items and map keys and values. M_ATTRIBUTE keeps the
raw flag so member ordering is unchanged.
Plan invalidation clears the cached input, output and errors when the
operation definition changes, but context params were computed once in
the constructor. After replacing input, getInput() returned the new
shape while getContextParams() still described the old one, which can
break endpoint resolution for customized models.

Re-read staticContextParams and operationContextParams from the
definition on invalidation, and rebuild the dynamic context params
lazily on next access.
Match the SDK's existing multi-line condition style: the first condition
stays on the opening line, and each following condition starts its own
line.
The plan-based serializers and parsers are now the only code path.
parseLegacy() and buildLegacy() were kept solely so the benchmark
harness could compare old and new dispatch in one process; nothing in
the request or response pipeline calls them, and the harness is no
longer part of this change.

Drop those methods along with the private dispatch helpers that only
they used, and reword docblocks that referred to the removed methods.
No public or protected API from the base branch is removed.
@rohangavankar

Copy link
Copy Markdown
Contributor Author

@stobrien89 re buildLegacy()/parseLegacy(): not needed. Only the benchmark used them. Removed in ce903a9, along with the private helpers that only they called. No public or protected methods removed compared with master. I re-ran the benchmark with these fixes and saw no difference beyond noise.

@stobrien89 stobrien89 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just two more things, one minor formatting issue. Could we also rerun or confirm the benchmark results against the final implementation?

Comment thread src/Api/Operation.php Outdated
Comment thread src/Api/Parser/JsonParser.php Outdated
The XML encode plan compiler computed rootName for every shape it
reached, which called Shape::getName() on nested shapes. Inline shapes
have no 'name' key, so PHP 8 raised "Undefined array key" warnings for
hand-built models. Only the top-level plan's rootName is ever read.

XmlBody::build() now resolves the root name the first time a shape is
built as a document root and caches it on that plan, matching the old
serializer, which looked up the name once on the root shape only.

Add XmlWireCompatibilityTest, which pins encode and decode output for
root-name precedence, namespaces, multiple attributes, empty, flattened
and nested collections, blobs, timestamps, special floats and unions.
Expected values match the pre-plan serializer and parser. Add
XmlPlanInvalidationTest covering plan rebuilds after model mutation.
The benchmark harness is no longer part of this change, so the
benchmark/ export-ignore and results ignore entries have nothing to
match.
Context params were rebuilt only when the operation definition itself
changed. Mutating the resolved input shape, for example replacing its
members, bumped the shared ShapeMap generation but left
getContextParams() returning params built from the old members.

Record the ShapeMap generation the params were built against and
rebuild them when it moves, the same check serde plans use.
Remove blank lines before closing class braces and put short multi-line
conditions on one line, which satisfies PSR-12 without a line break
after the opening parenthesis.
@rohangavankar

Copy link
Copy Markdown
Contributor Author

@stobrien89 Reran the benchmarks against the final implementation (head 701c316) with the PR's base commit 4a8bc49 as the baseline, on the x86 m7i.xlarge, alternating baseline and candidate, two runs each, JIT on and off. Full tables are in the PR description.

  • Whole operation, JIT on: awsJson1 deserialize -49.7%, serialize -24.9%, awsQuery deserialize -4.4%, restXml deserialize -2.3%. REST serialize, Query serialize and CBOR are within ±1%.
  • Body codec only, warm p50: JSON -15% to -47%, XML encode -40% to -46%, XML decode -11% to -22%. Output hashes match the base commit in every case.
  • Small payloads pay 3-7 us once to compile plans, and plans retain 2-10 KB per model graph.

Also pushed: e12b87e resolves the XML root name only for the root shape, which removes an Undefined array key "name" warning for inline shapes, and adds XmlWireCompatibilityTest and XmlPlanInvalidationTest.

@stobrien89 stobrien89 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a couple more comments

Comment on lines +103 to +104
? str_replace($nsPrefix, '', $member['locationName'])
: null,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Passingnull to str_replace() emits a PHP 8.1+ deprecation. This is now evaluated eagerly while compiling the plan, whereas the previous parser only evaluated it when falling back to XML attribute parsing. Could this use the member name as the fallback?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in c1eee3e. Falls back to the member name (locationName ?: $name), same as XmlBody uses when encoding. Test: testDecodeAttributeWithoutLocationNameUsesMemberName, which fails on any PHP warning or deprecation. The old code also decoded the attribute as null, and it now returns the value.

Comment thread src/Api/Operation.php
// Context params derive from the input shape graph, so rebuild them
// when any shape in the shared ShapeMap has been mutated.
$generation = $this->shapeMap->getGeneration();
if ($this->contextParams === null || $this->contextParamsGeneration !== $generation) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new generation check rebuilds context parameters, but getContextParam() still returns the value copied during construction. Directly mutating or removing a member’s contextParam increments the generation but returns the old context parameter. Read from the mutable definition or refresh the cached property, and test direct mutation/removal.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in c1eee3e. getContextParam() now reads the live definition, so mutating or removing a member's contextParam is reflected in getContextParam() and getContextParams(). The protected $contextParam property is still set in the constructor for subclasses. Test: testContextParamsFollowMemberContextParamMutation covers mutation and removal.

Shape::getContextParam() returned the copy AbstractModel took at
construction, so mutating or removing a member's contextParam bumped
the ShapeMap generation and rebuilt the operation's context params from
the stale copy. Read the current definition instead. The protected
$contextParam property is still populated for subclasses that read it.

XmlDecodePlanProvider now compiles the attribute key eagerly, so an
xmlAttribute member without a locationName passed null to str_replace()
(a PHP 8.1+ deprecation) and the attribute decoded as null. Fall back to
the member name, matching how XmlBody names attributes on encode.

Add tests for contextParam mutation and removal, and for decoding an
attribute without a locationName under an error handler that fails on
any warning or deprecation.

@stobrien89 stobrien89 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just needs a changelog entry for the Api namespace. otherwise looks ready to go!

@rohangavankar

Copy link
Copy Markdown
Contributor Author

@stobrien89 Done. Added the Api changelog entry in e895ae7 (.changes/nextrelease/serde-plan-cache.json).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants