Skip to content

docs: correct the SEARCH table function reference - #786

Merged
LuciferYang merged 7 commits into
lance-format:mainfrom
jackylee-ch:docs/search-tvf-reference
Sep 28, 2026
Merged

LuciferYang merged 7 commits into
lance-format:mainfrom
jackylee-ch:docs/search-tvf-reference

Conversation

@jackylee-ch

@jackylee-ch jackylee-ch commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • search_columns is required with no positional slot, so SEARCH needs named args (Spark 3.5+)
  • Drop the positional example, which always throws, and offset, which is never read
  • Fix Execution (scan path is namespace-dependent) and clarify the current test markers

Test plan

  • mkdocs build --strict

🤖 Generated with Claude Code

search_columns is required (LanceSearchTableFunctions.search throws when
it is empty) but has no positional slot, so the documented positional
example always fails and SEARCH in fact requires named arguments.

offset is documented but never read by search(); unknown named arguments
are silently ignored, so it looks accepted.

Execution described the removed bespoke single-partition path; the scan
now runs server-side through queryTable only when the namespace supports
it, and per-fragment otherwise. Validation claimed Docker coverage while
the pytest case is xfail and the JVM case is @disabled.
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Aug 26, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Aug 26, 2026
Comment thread docs/src/operations/dql/search.md Outdated
## Execution

Spark plans `SEARCH` as a DataSource V2 batch read with one input partition. The partition reader calls the Lance namespace `queryTable` API. With a directory namespace the search runs in the Spark process executing that reader; with a REST namespace the REST server handles the namespace request.
Spark plans `SEARCH` as a batch read carrying the full-text query as a scan option, wrapped in an optional filter, a projection, `ORDER BY _score DESC`, and `LIMIT k`. The scan then runs one of two ways: a single-partition server-side read through the Lance namespace `queryTable` API when the namespace supports it, or a distributed per-fragment scan for catalog-only namespaces and for reads that target a branch or tag.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I haven't see how search can target a branch or tag is this true or should we drop?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tag is not reachable, you are right to push on it: version goes through optionalLong, so a tag name never parses, and there is no tag_ identifier suffix to carry one. Branch is reachable, since resolveLanceTable calls loadTable and that matches BRANCH_SUFFIX, so table => '...docs.branch_audit' lands on the branch and shouldNamespaceFtsScan falls back, but nothing tests that path. Dropped the clause in 52baafa and left only the catalog-only condition. Happy to put branch back if you want it documented.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two corrections to my last reply. I over-corrected by dropping branch as well as tag, and nothing tests that path was wrong — LanceScanBuilderTest#testBranchFullTextQueryDoesNotUseNamespaceScan (LanceScanBuilderTest.java:374) already asserts shouldNamespaceFtsScan() is false for LanceRef.ofBranch. 4ea4cb3 makes the per-fragment scan the general fallback and lists the three conditions the server-side route needs.

Review feedback: a tag cannot reach SEARCH at all. The version argument
goes through optionalLong, so a tag name never parses, and there is no
tag_ identifier suffix to carry one.

A branch is reachable, since resolveLanceTable calls loadTable and that
matches BRANCH_SUFFIX on the identifier, but nothing tests it. Leave the
condition users can act on and drop the rest.
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 3, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-changes Latest Gatekeeper recommendation requests changes. label Sep 3, 2026
I over-corrected: only the tag half of the previous clause was wrong, and
dropping branch too made the routing read as if catalog-only namespaces
were the sole fallback trigger.

State the fallback as the otherwise case and list the three conditions the
server-side route needs, matching shouldNamespaceFtsScan: a namespace that
implements queryTable, a ref that is not a branch or tag, and no pushed
aggregation. Enumerating only the fallback triggers went stale twice.
@lance-gatekeeper lance-gatekeeper Bot removed the K-changes Latest Gatekeeper recommendation requests changes. label Sep 3, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 3, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 15, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 15, 2026
@jackylee-ch

Copy link
Copy Markdown
Contributor Author

cc @LuciferYang

@LuciferYang LuciferYang left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the cleanup — the corrections all match the code: search_columns is required with no positional slot, offset is never read on the SEARCH path, and the two-route Execution rewrite (server-side queryTable vs per-fragment scan) lines up with LanceScanBuilder. Two documentation notes below, both non-blocking.

Comment thread docs/src/operations/dql/search.md Outdated
## Validation

The Docker integration suite covers `SEARCH` against the directory namespace and a REST namespace backed by a directory namespace. The `Spark Search Docker` GitHub Actions workflow runs both backends for pull requests.
The `Spark Search Docker` workflow exercises `SEARCH` against directory and REST-directory namespaces. The Docker test still carries an `xfail` marker, but it can pass as `XPASS`; the JVM `SEARCH` cases remain `@Disabled`. [`VECTOR_SEARCH`](vector-search.md) and [`HYBRID_SEARCH`](hybrid-search.md) are also exercised by the workflow.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Validation section leans on test-framework jargon (xfail, XPASS, @Disabled) that means little to someone reading the feature reference. More importantly, the reason those markers exist is that SEARCH's server-side queryTable route currently ignores the structured FTS query and cannot produce the _score column that the plan always projects and sorts by, so the query fails on the _score projection (the Python case is xfail with raises=Exception). That is exactly the path the Basic Usage example above takes on a queryTable-capable namespace such as a directory namespace, where it selects _score. So the section reads like "tested and passing" while the feature's headline example is actually broken.

Consider collapsing the CI mechanics into a one-line coverage statement and adding a short known-limitation note near the top: on a queryTable-capable namespace SEARCH currently fails on the _score projection. Leave the xfail/@Disabled details in the test code.

Comment thread docs/src/operations/dql/search.md Outdated
!!! note "Named Arguments"
Named arguments require Spark 3.5 or later. On Spark 3.4, use the positional form.
!!! note "Named Arguments Required"
`search_columns` is required and has no positional slot, so `SEARCH` must be called with named arguments. Named arguments require Spark 3.5 or later, so `SEARCH` is not available on Spark 3.4.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SEARCH is not available on Spark 3.4 is slightly off: the 3.4 module registers SEARCH identically to 3.5/4.0 (no version gating), so the function is visible. What actually blocks it is that search_columns is required and cannot be passed positionally, only as a named argument, and named arguments need Spark 3.5+.

The conclusion (unusable on 3.4) is right, but "not available" reads as "not registered". Consider "cannot be used on Spark 3.4", or spell out that it is registered but uncallable because 3.4 lacks named-argument support.

The Validation section leaned on xfail/XPASS/@disabled and read as "tested and
passing", while the Basic Usage example selects _score and therefore fails on a
queryTable-capable namespace. Lead with that instead, in the words the repo's own
xfail reason uses, and collapse the CI mechanics into one coverage line.

Also correct "not available on Spark 3.4": the 3.4 module injects the search table
function identically to 3.5 and 4.x, so it is registered but uncallable there.
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 16, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-changes Latest Gatekeeper recommendation requests changes. label Sep 16, 2026
I added a warning saying the server-side route cannot produce _score, taking the
xfail reason and the review note as fact without reproducing it. Running the
disabled JVM case on this head shows a different failure: it throws
"SEARCH requires search_columns for full-text search" at analysis time, because
the case uses the positional form. It never reaches execution, so nothing here
demonstrates the _score claim.

Remove the warning and the matching Validation claim, and state only the coverage
the workflow provides. The Spark 3.4 rewording stays.
@lance-gatekeeper lance-gatekeeper Bot removed the K-changes Latest Gatekeeper recommendation requests changes. label Sep 16, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

The latest revision removes the unproven server-side failure warning and matching validation claim. The SEARCH argument, Spark 3.4 availability, execution-route, and validation-coverage corrections now match the implementation and verified behavior.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 16, 2026

@LuciferYang LuciferYang left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs review of the SEARCH reference. The parameter, plan-shape, and routing changes all check out against the code.

One thing worth resolving before merge: this PR drops the note that server-side SEARCH currently can't produce _score (dir.rs only handles string_query and ignores the structured_query the client sends), so the page now presents a route that still fails on a directory / REST-directory namespace as if it works, including the Basic Usage example. The repo's own tests still mark this xfail / @Disabled. Details inline, plus two smaller notes on the Validation coverage wording and the Spark 3.4 phrasing.

## Execution

Spark plans `SEARCH` as a DataSource V2 batch read with one input partition. The partition reader calls the Lance namespace `queryTable` API. With a directory namespace the search runs in the Spark process executing that reader; with a REST namespace the REST server handles the namespace request.
Spark plans `SEARCH` as a batch read carrying the full-text query as a scan option, wrapped in an optional filter, a projection, `ORDER BY _score DESC`, and `LIMIT k`. The scan then runs either as a single-partition server-side read through the Lance namespace `queryTable` API, or otherwise as a distributed per-fragment scan. The server-side route requires a namespace that implements `queryTable`, a read that does not target a branch or tag, and no pushed-down aggregation.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Execution section presents the server-side queryTable route as a normal way SEARCH runs, but that route is currently broken for full-text search. The directory namespace's query_table (dir.rs) only handles string_query; its comment says "we only support string_query" and it ignores structured_query, while the Spark client only ever sends structured_query for SEARCH. So no FTS runs and _score is never produced, yet the SEARCH plan always projects and sorts by _score, so the query fails. Both a directory namespace and a REST-over-directory namespace implement queryTable and take this route, so the Basic Usage example above (which selects _score) fails on them.

The commit removed the warning calling the failure "unproven", but that was only because the @Disabled JVM case uses the positional form and throws "requires search_columns" at analysis, never reaching execution. The Python xfail case uses named arguments, does reach execution, and hits exactly this dir.rs gap.

On a queryTable-capable namespace SEARCH currently fails on _score. Keep a one-line known-limitation note (the top-of-page warning you removed is fine) until dir.rs adds structured FTS.

## Validation

The Docker integration suite covers `SEARCH` against the directory namespace and a REST namespace backed by a directory namespace. The `Spark Search Docker` GitHub Actions workflow runs both backends for pull requests.
The `Spark Search Docker` workflow covers `SEARCH`, [`VECTOR_SEARCH`](vector-search.md) and [`HYBRID_SEARCH`](hybrid-search.md) against directory and REST-directory namespaces.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"covers SEARCH, VECTOR_SEARCH and HYBRID_SEARCH" lists all three as equals, but in the workflow the SEARCH case is @pytest.mark.xfail (expected to fail) and the JVM cases are @Disabled; only VECTOR_SEARCH and HYBRID_SEARCH are real passing coverage. The previous revision still said "The SEARCH cases are currently expected to fail", and this PR dropped that clause too, so SEARCH now reads as tested and working.

Mark SEARCH as expected to fail, or pull it out of the "covers" list and note it is currently xfail. The root cause is that the server-side queryTable route can't produce _score; see my comment on the Execution section.

!!! note "Named Arguments"
Named arguments require Spark 3.5 or later. On Spark 3.4, use the positional form.
!!! note "Named Arguments Required"
`search_columns` is required and has no positional slot, so `SEARCH` must be called with named arguments. Named arguments require Spark 3.5 or later. On Spark 3.4 the function is registered but cannot be called.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"registered but cannot be called" is right in direction (3.4 registers SEARCH the same as 3.5/4.x), but "cannot be called" is a touch strong. A positional call SEARCH('t','q',5) does parse and does reach the function body on 3.4; it just throws at analysis because search_columns is missing. A named call fails to parse because 3.4 has no named arguments (SPARK-44059 is 3.5+).

More precisely: it is registered but has no valid call on 3.4. Consider "registered but has no valid call on Spark 3.4", or spell out that a named call fails to parse and a positional call always throws for the missing search_columns. The conclusion is correct; this is precision only.

@jackylee-ch

Copy link
Copy Markdown
Contributor Author

Thanks @LuciferYang — both addressed at 201b57c: the Spark 3.4 note now says "registered but cannot be called" (named-argument requirement), and the Validation section is collapsed to a one-line coverage statement, with the xfail/@disabled mechanics left in the test code. On the _score projection: I held off on the known-limitation note because on a queryTable-capable directory namespace the reproducer planned SEARCH as LanceSearchScan and returned a positive _score, so the headline example isn't broken on that route. Happy to add a narrower note if there's a namespace route where _score does fail.

@lance-gatekeeper lance-gatekeeper Bot added K-approved Latest Gatekeeper recommendation permits acceptance. and removed K-approved Latest Gatekeeper recommendation permits acceptance. labels Sep 26, 2026

@LuciferYang LuciferYang left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The corrections (search_columns required, dropped positional example and offset, namespace-dependent Execution) are accurate, and dropping the unproven _score warning was the right call. One non-blocking note on the Validation wording inline. LGTM.

## Validation

The Docker integration suite covers `SEARCH` against the directory namespace and a REST namespace backed by a directory namespace. The `Spark Search Docker` GitHub Actions workflow runs both backends for pull requests.
The `Spark Search Docker` workflow covers `SEARCH`, [`VECTOR_SEARCH`](vector-search.md) and [`HYBRID_SEARCH`](hybrid-search.md) against directory and REST-directory namespaces.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dropping the unproven _score warning was the right call: the mechanism was never reproduced, so it shouldn't be stated as fact. But that removal also took out the verifiable part: that the SEARCH cases currently fail. Validation now says the workflow "covers SEARCH, VECTOR_SEARCH and HYBRID_SEARCH", listing all three flatly, which reads as if SEARCH passes in CI too. In fact the Python test_search_table_function (test_lance_spark.py:1700, which uses named args and does reach execution) is xfail, the JVM SEARCH cases are @Disabled, and only VECTOR/HYBRID assert real results.

No need to reintroduce the unproven mechanism; just state the verifiable marker, e.g. split SEARCH out of "covers": "covers VECTOR_SEARCH and HYBRID_SEARCH …; SEARCH is run as a currently-xfail case." That keeps the _score claim out while stopping readers from following Basic Usage into a call that currently fails.

@LuciferYang
LuciferYang merged commit 710ee89 into lance-format:main Sep 28, 2026
3 checks passed
@LuciferYang

Copy link
Copy Markdown
Collaborator

Merged into main. Thanks @jackylee-ch

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation K-approved Latest Gatekeeper recommendation permits acceptance.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants