Repository navigation
docs: correct the SEARCH table function reference - #786
Conversation
search_columns is required (LanceSearchTableFunctions.search throws when it is empty) but has no positional slot, so the documented positional example always fails and SEARCH in fact requires named arguments. offset is documented but never read by search(); unknown named arguments are silently ignored, so it looks accepted. Execution described the removed bespoke single-partition path; the scan now runs server-side through queryTable only when the namespace supports it, and per-fragment otherwise. Validation claimed Docker coverage while the pytest case is xfail and the JVM case is @disabled.
| ## Execution | ||
|
|
||
| Spark plans `SEARCH` as a DataSource V2 batch read with one input partition. The partition reader calls the Lance namespace `queryTable` API. With a directory namespace the search runs in the Spark process executing that reader; with a REST namespace the REST server handles the namespace request. | ||
| Spark plans `SEARCH` as a batch read carrying the full-text query as a scan option, wrapped in an optional filter, a projection, `ORDER BY _score DESC`, and `LIMIT k`. The scan then runs one of two ways: a single-partition server-side read through the Lance namespace `queryTable` API when the namespace supports it, or a distributed per-fragment scan for catalog-only namespaces and for reads that target a branch or tag. |
There was a problem hiding this comment.
I haven't see how search can target a branch or tag is this true or should we drop?
There was a problem hiding this comment.
Tag is not reachable, you are right to push on it: version goes through optionalLong, so a tag name never parses, and there is no tag_ identifier suffix to carry one. Branch is reachable, since resolveLanceTable calls loadTable and that matches BRANCH_SUFFIX, so table => '...docs.branch_audit' lands on the branch and shouldNamespaceFtsScan falls back, but nothing tests that path. Dropped the clause in 52baafa and left only the catalog-only condition. Happy to put branch back if you want it documented.
There was a problem hiding this comment.
Two corrections to my last reply. I over-corrected by dropping branch as well as tag, and nothing tests that path was wrong — LanceScanBuilderTest#testBranchFullTextQueryDoesNotUseNamespaceScan (LanceScanBuilderTest.java:374) already asserts shouldNamespaceFtsScan() is false for LanceRef.ofBranch. 4ea4cb3 makes the per-fragment scan the general fallback and lists the three conditions the server-side route needs.
Review feedback: a tag cannot reach SEARCH at all. The version argument goes through optionalLong, so a tag name never parses, and there is no tag_ identifier suffix to carry one. A branch is reachable, since resolveLanceTable calls loadTable and that matches BRANCH_SUFFIX on the identifier, but nothing tests it. Leave the condition users can act on and drop the rest.
I over-corrected: only the tag half of the previous clause was wrong, and dropping branch too made the routing read as if catalog-only namespaces were the sole fallback trigger. State the fallback as the otherwise case and list the three conditions the server-side route needs, matching shouldNamespaceFtsScan: a namespace that implements queryTable, a ref that is not a branch or tag, and no pushed aggregation. Enumerating only the fallback triggers went stale twice.
|
cc @LuciferYang |
LuciferYang
left a comment
There was a problem hiding this comment.
Thanks for the cleanup — the corrections all match the code: search_columns is required with no positional slot, offset is never read on the SEARCH path, and the two-route Execution rewrite (server-side queryTable vs per-fragment scan) lines up with LanceScanBuilder. Two documentation notes below, both non-blocking.
| ## Validation | ||
|
|
||
| The Docker integration suite covers `SEARCH` against the directory namespace and a REST namespace backed by a directory namespace. The `Spark Search Docker` GitHub Actions workflow runs both backends for pull requests. | ||
| The `Spark Search Docker` workflow exercises `SEARCH` against directory and REST-directory namespaces. The Docker test still carries an `xfail` marker, but it can pass as `XPASS`; the JVM `SEARCH` cases remain `@Disabled`. [`VECTOR_SEARCH`](vector-search.md) and [`HYBRID_SEARCH`](hybrid-search.md) are also exercised by the workflow. |
There was a problem hiding this comment.
The Validation section leans on test-framework jargon (xfail, XPASS, @Disabled) that means little to someone reading the feature reference. More importantly, the reason those markers exist is that SEARCH's server-side queryTable route currently ignores the structured FTS query and cannot produce the _score column that the plan always projects and sorts by, so the query fails on the _score projection (the Python case is xfail with raises=Exception). That is exactly the path the Basic Usage example above takes on a queryTable-capable namespace such as a directory namespace, where it selects _score. So the section reads like "tested and passing" while the feature's headline example is actually broken.
Consider collapsing the CI mechanics into a one-line coverage statement and adding a short known-limitation note near the top: on a queryTable-capable namespace SEARCH currently fails on the _score projection. Leave the xfail/@Disabled details in the test code.
| !!! note "Named Arguments" | ||
| Named arguments require Spark 3.5 or later. On Spark 3.4, use the positional form. | ||
| !!! note "Named Arguments Required" | ||
| `search_columns` is required and has no positional slot, so `SEARCH` must be called with named arguments. Named arguments require Spark 3.5 or later, so `SEARCH` is not available on Spark 3.4. |
There was a problem hiding this comment.
SEARCH is not available on Spark 3.4 is slightly off: the 3.4 module registers SEARCH identically to 3.5/4.0 (no version gating), so the function is visible. What actually blocks it is that search_columns is required and cannot be passed positionally, only as a named argument, and named arguments need Spark 3.5+.
The conclusion (unusable on 3.4) is right, but "not available" reads as "not registered". Consider "cannot be used on Spark 3.4", or spell out that it is registered but uncallable because 3.4 lacks named-argument support.
The Validation section leaned on xfail/XPASS/@disabled and read as "tested and passing", while the Basic Usage example selects _score and therefore fails on a queryTable-capable namespace. Lead with that instead, in the words the repo's own xfail reason uses, and collapse the CI mechanics into one coverage line. Also correct "not available on Spark 3.4": the 3.4 module injects the search table function identically to 3.5 and 4.x, so it is registered but uncallable there.
I added a warning saying the server-side route cannot produce _score, taking the xfail reason and the review note as fact without reproducing it. Running the disabled JVM case on this head shows a different failure: it throws "SEARCH requires search_columns for full-text search" at analysis time, because the case uses the positional form. It never reaches execution, so nothing here demonstrates the _score claim. Remove the warning and the matching Validation claim, and state only the coverage the workflow provides. The Spark 3.4 rewording stays.
There was a problem hiding this comment.
✅ Gate recommendation: approve.
The latest revision removes the unproven server-side failure warning and matching validation claim. The SEARCH argument, Spark 3.4 availability, execution-route, and validation-coverage corrections now match the implementation and verified behavior.
LuciferYang
left a comment
There was a problem hiding this comment.
Docs review of the SEARCH reference. The parameter, plan-shape, and routing changes all check out against the code.
One thing worth resolving before merge: this PR drops the note that server-side SEARCH currently can't produce _score (dir.rs only handles string_query and ignores the structured_query the client sends), so the page now presents a route that still fails on a directory / REST-directory namespace as if it works, including the Basic Usage example. The repo's own tests still mark this xfail / @Disabled. Details inline, plus two smaller notes on the Validation coverage wording and the Spark 3.4 phrasing.
| ## Execution | ||
|
|
||
| Spark plans `SEARCH` as a DataSource V2 batch read with one input partition. The partition reader calls the Lance namespace `queryTable` API. With a directory namespace the search runs in the Spark process executing that reader; with a REST namespace the REST server handles the namespace request. | ||
| Spark plans `SEARCH` as a batch read carrying the full-text query as a scan option, wrapped in an optional filter, a projection, `ORDER BY _score DESC`, and `LIMIT k`. The scan then runs either as a single-partition server-side read through the Lance namespace `queryTable` API, or otherwise as a distributed per-fragment scan. The server-side route requires a namespace that implements `queryTable`, a read that does not target a branch or tag, and no pushed-down aggregation. |
There was a problem hiding this comment.
The Execution section presents the server-side queryTable route as a normal way SEARCH runs, but that route is currently broken for full-text search. The directory namespace's query_table (dir.rs) only handles string_query; its comment says "we only support string_query" and it ignores structured_query, while the Spark client only ever sends structured_query for SEARCH. So no FTS runs and _score is never produced, yet the SEARCH plan always projects and sorts by _score, so the query fails. Both a directory namespace and a REST-over-directory namespace implement queryTable and take this route, so the Basic Usage example above (which selects _score) fails on them.
The commit removed the warning calling the failure "unproven", but that was only because the @Disabled JVM case uses the positional form and throws "requires search_columns" at analysis, never reaching execution. The Python xfail case uses named arguments, does reach execution, and hits exactly this dir.rs gap.
On a queryTable-capable namespace SEARCH currently fails on _score. Keep a one-line known-limitation note (the top-of-page warning you removed is fine) until dir.rs adds structured FTS.
| ## Validation | ||
|
|
||
| The Docker integration suite covers `SEARCH` against the directory namespace and a REST namespace backed by a directory namespace. The `Spark Search Docker` GitHub Actions workflow runs both backends for pull requests. | ||
| The `Spark Search Docker` workflow covers `SEARCH`, [`VECTOR_SEARCH`](vector-search.md) and [`HYBRID_SEARCH`](hybrid-search.md) against directory and REST-directory namespaces. |
There was a problem hiding this comment.
"covers SEARCH, VECTOR_SEARCH and HYBRID_SEARCH" lists all three as equals, but in the workflow the SEARCH case is @pytest.mark.xfail (expected to fail) and the JVM cases are @Disabled; only VECTOR_SEARCH and HYBRID_SEARCH are real passing coverage. The previous revision still said "The SEARCH cases are currently expected to fail", and this PR dropped that clause too, so SEARCH now reads as tested and working.
Mark SEARCH as expected to fail, or pull it out of the "covers" list and note it is currently xfail. The root cause is that the server-side queryTable route can't produce _score; see my comment on the Execution section.
| !!! note "Named Arguments" | ||
| Named arguments require Spark 3.5 or later. On Spark 3.4, use the positional form. | ||
| !!! note "Named Arguments Required" | ||
| `search_columns` is required and has no positional slot, so `SEARCH` must be called with named arguments. Named arguments require Spark 3.5 or later. On Spark 3.4 the function is registered but cannot be called. |
There was a problem hiding this comment.
"registered but cannot be called" is right in direction (3.4 registers SEARCH the same as 3.5/4.x), but "cannot be called" is a touch strong. A positional call SEARCH('t','q',5) does parse and does reach the function body on 3.4; it just throws at analysis because search_columns is missing. A named call fails to parse because 3.4 has no named arguments (SPARK-44059 is 3.5+).
More precisely: it is registered but has no valid call on 3.4. Consider "registered but has no valid call on Spark 3.4", or spell out that a named call fails to parse and a positional call always throws for the missing search_columns. The conclusion is correct; this is precision only.
|
Thanks @LuciferYang — both addressed at 201b57c: the Spark 3.4 note now says "registered but cannot be called" (named-argument requirement), and the Validation section is collapsed to a one-line coverage statement, with the xfail/@disabled mechanics left in the test code. On the |
LuciferYang
left a comment
There was a problem hiding this comment.
The corrections (search_columns required, dropped positional example and offset, namespace-dependent Execution) are accurate, and dropping the unproven _score warning was the right call. One non-blocking note on the Validation wording inline. LGTM.
| ## Validation | ||
|
|
||
| The Docker integration suite covers `SEARCH` against the directory namespace and a REST namespace backed by a directory namespace. The `Spark Search Docker` GitHub Actions workflow runs both backends for pull requests. | ||
| The `Spark Search Docker` workflow covers `SEARCH`, [`VECTOR_SEARCH`](vector-search.md) and [`HYBRID_SEARCH`](hybrid-search.md) against directory and REST-directory namespaces. |
There was a problem hiding this comment.
Dropping the unproven _score warning was the right call: the mechanism was never reproduced, so it shouldn't be stated as fact. But that removal also took out the verifiable part: that the SEARCH cases currently fail. Validation now says the workflow "covers SEARCH, VECTOR_SEARCH and HYBRID_SEARCH", listing all three flatly, which reads as if SEARCH passes in CI too. In fact the Python test_search_table_function (test_lance_spark.py:1700, which uses named args and does reach execution) is xfail, the JVM SEARCH cases are @Disabled, and only VECTOR/HYBRID assert real results.
No need to reintroduce the unproven mechanism; just state the verifiable marker, e.g. split SEARCH out of "covers": "covers VECTOR_SEARCH and HYBRID_SEARCH …; SEARCH is run as a currently-xfail case." That keeps the _score claim out while stopping readers from following Basic Usage into a call that currently fails.
|
Merged into main. Thanks @jackylee-ch |
Summary
search_columnsis required with no positional slot, soSEARCHneeds named args (Spark 3.5+)offset, which is never readTest plan
mkdocs build --strict🤖 Generated with Claude Code