Ladybug version
0.18.3. Also 0.19.1, 0.20.0, 0.20.1 and 0.20.2.
What operating system are you using?
Linux x86-64, Python 3.10–3.14.
What happened?
CREATE_FTS_INDEX accepts an ignore_pattern option (default strips digits
and most punctuation). A custom pattern changes which tokens the index stores,
but QUERY_FTS_INDEX still normalises the query with the default pattern, so
any token the custom pattern was meant to preserve can never be matched
exactly.
CREATE (:T {id:1, content:'Audi A4 Avant 2022'}), (:T {id:2, content:'Volkswagen T-Roc Life'});
CALL CREATE_FTS_INDEX('T', 'idx', ['content'], stemmer := 'none', ignore_pattern := '[^[:alnum:]-]+');
CALL QUERY_FTS_INDEX('T', 'idx', 'A4') RETURN node.id; -- expected [1], actual []
CALL QUERY_FTS_INDEX('T', 'idx', 't-roc') RETURN node.id; -- expected [2], actual []
CALL QUERY_FTS_INDEX('T', 'idx', 'a?') RETURN node.id; -- [1] the index holds 'a4'
CALL QUERY_FTS_INDEX('T', 'idx', 't?roc') RETURN node.id; -- [2] the index holds 't-roc'
Expected: [1] / [2] for the exact queries. Actual: [] / [].
The wildcard forms return [1] / [2].
The wildcards prove the index side honoured the option: a4 and t-roc are
stored as single terms. The exact forms fail because the query is run through
the default pattern first — A4 → a, t-roc → t roc — and neither a
nor t/roc is a stored term.
| version |
exact A4 / t-roc |
wildcard a? / t?roc |
| 0.18.3 (ext 0.18.1) |
[] / [] |
[1] / [2] |
| 0.19.1 |
[] / [] |
[1] / [2] |
| 0.20.0, 0.20.1, 0.20.2 |
[] / [] |
[1] / [2] |
Identical on every version that loads the extension. (The reproducer uses
literal query strings on purpose: on 0.20.x a re-executed parameterized query
replays the first call's rows — #877 — which would mask this one.)
Not kuzu#5942. That issue was the default pattern stripping digits,
closed with "pass ignore_pattern". Here the option is passed and exact match
still fails.
Suspected origin, not bisected. CREATE_FTS_INDEX rewrites to
_CREATE_FTS_INDEX passing only stemmer and stopWords
(fts/src/function/create_fts_index.cpp). Index build uses the custom pattern
via the TOKENIZE macro; the catalog entry _CREATE_FTS_INDEX writes — and
that QUERY_FTS_INDEX reads — keeps the default. Same drop is on extensions
main.
Any corpus whose identifiers carry digits — product codes, car models (A4 vs
A6), vitamins (B12 vs B6), years — cannot be searched by those identifiers.
Wildcards retrieve the stored token (a? hits a4) but cannot tell A4 from
A6.
Workaround. None inside the engine. We index a letter-only alias next to
each digit-bearing token (A4 → xafour) and append the same alias to the
query text.
Are there known steps to reproduce?
#!/usr/bin/env -S uv run --script
# /// script
# requires-python = ">=3.10,<3.15"
# dependencies = ["ladybug==0.18.3"]
# ///
"""FTS `ignore_pattern` is applied when the index is built but not when a
query is normalised: QUERY_FTS_INDEX still splits on the default pattern.
Exits 1 when the bug reproduces.
uv run fts_ignore_pattern_query_side.py
"""
import os
import sys
import tempfile
import ladybug as kuzu
tmp = tempfile.mkdtemp(prefix="lbug_fts_ignore_")
def ids(result):
out = []
while result.has_next():
out.append(result.get_next()[0])
return sorted(out)
print(f"ladybug {kuzu.__version__}")
conn = kuzu.Connection(kuzu.Database(os.path.join(tmp, "db")))
conn.execute("LOAD EXTENSION fts;")
conn.execute("CREATE NODE TABLE T(id INT64, content STRING, PRIMARY KEY(id));")
conn.execute("CREATE (:T {id:1, content:'Audi A4 Avant 2022'}), (:T {id:2, content:'Volkswagen T-Roc Life'});")
# Keep digits and hyphens as token characters (default pattern strips both).
conn.execute("CALL CREATE_FTS_INDEX('T', 'idx', ['content'], stemmer := 'none', ignore_pattern := '[^[:alnum:]-]+');")
def q(term):
# Literal, not a parameter: on 0.20.x a re-executed parameterized query
# replays the first call's rows (separate bug) and would mask this one.
return ids(conn.execute(f"CALL QUERY_FTS_INDEX('T', 'idx', '{term}', top := 5) RETURN node.id;"))
exact_a4 = q("A4") # expected [1]
exact_troc = q("t-roc") # expected [2]
wild_a4 = q("a?") # [1] -> the index DID store 'a4'
wild_troc = q("t?roc") # [2] -> the index DID store 't-roc'
print(f"exact 'A4' -> {exact_a4} (expected [1])")
print(f"exact 't-roc' -> {exact_troc} (expected [2])")
print(f"wildcard 'a?' -> {wild_a4} (index stored the digit token)")
print(f"wildcard 't?roc' -> {wild_troc} (index stored the hyphenated token)")
bug = exact_a4 != [1] or exact_troc != [2]
print("BUG REPRODUCED" if bug else "clean on this version")
sys.exit(1 if bug else 0)
Ladybug version
0.18.3. Also 0.19.1, 0.20.0, 0.20.1 and 0.20.2.
What operating system are you using?
Linux x86-64, Python 3.10–3.14.
What happened?
CREATE_FTS_INDEXaccepts anignore_patternoption (default strips digitsand most punctuation). A custom pattern changes which tokens the index stores,
but
QUERY_FTS_INDEXstill normalises the query with the default pattern, soany token the custom pattern was meant to preserve can never be matched
exactly.
Expected:
[1]/[2]for the exact queries. Actual:[]/[].The wildcard forms return
[1]/[2].The wildcards prove the index side honoured the option:
a4andt-rocarestored as single terms. The exact forms fail because the query is run through
the default pattern first —
A4→a,t-roc→t roc— and neitheranor
t/rocis a stored term.A4/t-roca?/t?roc[]/[][1]/[2][]/[][1]/[2][]/[][1]/[2]Identical on every version that loads the extension. (The reproducer uses
literal query strings on purpose: on 0.20.x a re-executed parameterized query
replays the first call's rows — #877 — which would mask this one.)
Not kuzu#5942. That issue was the default pattern stripping digits,
closed with "pass
ignore_pattern". Here the option is passed and exact matchstill fails.
Suspected origin, not bisected.
CREATE_FTS_INDEXrewrites to_CREATE_FTS_INDEXpassing onlystemmerandstopWords(
fts/src/function/create_fts_index.cpp). Index build uses the custom patternvia the TOKENIZE macro; the catalog entry
_CREATE_FTS_INDEXwrites — andthat
QUERY_FTS_INDEXreads — keeps the default. Same drop is on extensionsmain.Any corpus whose identifiers carry digits — product codes, car models (A4 vs
A6), vitamins (B12 vs B6), years — cannot be searched by those identifiers.
Wildcards retrieve the stored token (
a?hitsa4) but cannot tell A4 fromA6.
Workaround. None inside the engine. We index a letter-only alias next to
each digit-bearing token (
A4→xafour) and append the same alias to thequery text.
Are there known steps to reproduce?