Conversation
|
| } | ||
|
|
||
| for (let i = 0; i < this.lookupStages.length; i++) { | ||
| // Within a stage, we can resolve lookups concurrently. |
There was a problem hiding this comment.
With the implementation from this PR, joins are no longer processed concurrently. It's possible to add that back with minor added complexity, but:
- this is only relevant for complex sync streams
- we already evaluate other users / queriers concurrently, to the point where bucket storage is likely the bottleneck and not JS
So I don't think this is necessarily something worth doing, but I can change this here / in a follow-up PR if needed.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 68a4c2dc8b
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: af866d1014
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
rkistner
left a comment
There was a problem hiding this comment.
Codex picked up a performance regression here with intersecting results. Example queries:
-- In this case we can intersect results
SELECT * FROM issues
WHERE id IN auth.parameter('allowed_issue_ids')
AND id IN subscription.parameter('visible_issue_ids')
AND id IN subscription.parameter('selected_issue_ids');
-- In this case we can compute the checks independently
SELECT * FROM issues
WHERE 'issues' IN auth.parameter('tables')
AND 'read' IN auth.parameter('permissions');With the old implementation, the intersection was computed first. The new implementation builds up the Cartesian product of the arrays before filtering them. In a cases like the above, the number of results can explode even with modestly-sized arrays, leading to slow performance and/or OOM-crashes.
dade362 to
849a419
Compare
|
That is a good point. Since the columns of intersections are known statically, we can evaluate them for each row added to the result set instead of removing rows after building the cartesian product. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 849a419d50
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
rkistner
left a comment
There was a problem hiding this comment.
I like the approach of using a ResultSet here, but Codex picked up another couple of performance regressions with the implementation (manually confirmed the findings).
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c9bf009cf7
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1183c12e16
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 07a4ef428c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
rkistner
left a comment
There was a problem hiding this comment.
The further optimizations did help, but there are still some further optimizations needed to avoid performance regressions. See the individual comments, analyzed using Codex.
| this.#checkInstantiable(); | ||
| } else { | ||
| // For this parameter to be part of the intersection optimization, it must return exactly one column. | ||
| for (const [value] of new Set(outputs)) { |
There was a problem hiding this comment.
Outputs is SqliteParameterValue[][], so this creates a Set<SqliteParameterValue[]>. That will compare arrays by identity, rather than value, so de-duplication doesn't work properly.
Since you're only using one array element, it should be sufficient to just change the order:
for (const value of new Set(outputs.map(([value]) => value))) {A regression test covering this case would also be good - see this suggestion from Codex:
For a IN subscription.parameter('x') AND a IN auth.parameter('y'):
- x = ['unauthorized', 'unauthorized'], y = ['allowed'] → no buckets.
- x = ['allowed', 'allowed'], y = ['allowed'] → exactly the allowed bucket.
| for (const [parameter, intersections] of parameterToIntersection.entries()) { | ||
| if (intersections.length == 1) { | ||
| this.resultSet.multiply(parameter.resultSetIndex, intersections[0].materializedRows); | ||
| } else { | ||
| const rows = intersections.reduce((acc, intersection) => | ||
| intersection.materializedRows.length < acc.materializedRows.length ? intersection : acc | ||
| ).materializedRows; | ||
|
|
||
| this.resultSet.multiply( | ||
| parameter.resultSetIndex, | ||
| rows.filter(([value]) => { | ||
| return intersections.every((intersection) => { | ||
| const matchingCount = intersection.rows.get(value); | ||
| return matchingCount === intersection.amountOfRequestParameters; | ||
| }); | ||
| }) | ||
| ); | ||
| } | ||
|
|
||
| this.#checkInstantiable(); | ||
| } |
There was a problem hiding this comment.
If I understand it correctly, the optimization above now removes values that don't intersect. But any remaining values are still multiplied, which can result in exponential computations and memory.
Maybe it can work to have a multiply function that takes into account the intersection, something like this?
resultSet.multiplyCorrelated(
parameters.map((p) => p.resultSetIndex),
intersection.materializedRows
);| using add = this.#prepareAddingResultSet(resultSetIndex); | ||
|
|
||
| const originalLength = this.#rows.length; | ||
| for (let i = 0; i < originalLength; i++) { | ||
| const lookup = lookupsByRow[i]; | ||
| if (lookup.foundRows.length === 0 || !this.#multiplyAtRow(resultSetIndex, add.filter, i, lookup.foundRows)) { | ||
| // The row has no matching join partner, so remove it. We can't split it immediately because #multiplyAtRow is | ||
| // still iterating through rows. | ||
| add.deletedRows.push(i); | ||
| } | ||
| } |
There was a problem hiding this comment.
This still does O(N*M) comparisons, even though the results are O(N+M). An in-memory index (Map) should be able to reduce the comparisons to O(N+M) as well.
Here it appears that the previous implementation did cover intersection on a single key efficiently. A stretch goal would be to cover multiple intersection parameters, such as x and y in this example:
SELECT items.*
FROM items
JOIN allowed a ON items.x = a.x AND items.y = a.y
JOIN enabled e ON items.x = e.x AND items.y = e.y
WHERE a.user_id = auth.user_id()I believe the previous implementation would efficiently intersect on x, then multiply on filter on y. We could make that even more efficient by intersecting on (x, y), but that's a further optimization rather than covering a performance regression.
When querying buckets for Sync Streams, we generally try to resolve parameters independently to form a cartesian product in the end. This is correct for most streams, but goes wrong when a parameter index has more than one column. For example, in
SELECT a.* FROM a, b WHERE a.c1 = b.c1 AND a.c2 = b.c2 AND b.u = auth.user_id(), the two parameters areb.c1andb.c2. If we encounter multiple rows ofbthrough a lookup result, we can't assume those to be independent parameters though! We can only pair parameters that originate from the same row.This is currently implemented by tracking provenance for each parameter value back to the lookup this originally came from. When we build the cartesian product in the end, we ignore values with incompatible provenance from different rows. Unfortunately, tracking provenance is both kind of expensive and very tricky to get right.
Semantically, evaluating bucket parameters involves:
This replaces the previous querier logic with an actual result set implementation: We start out with a unit set of one row without columns, then go through added lookups that are cross-joined (table-valued functions) or inner-joined (parameter lookups). If we end up with an empty intermediate result set at any point, we know there won't be any buckets and bail out early.
Intersection parameters require special consideration now, but can be implemented by going through the result set and deleting rows where the columns don't match.
AI use: The approach is manual, most tests and some implementation details are generated with Claude Code.