Back to releases
v1.1.1

Standalone and Chinese retrieval fixes (eval-driven) (v1.1.1)

Fixes three defects — standalone knowledge base creation returning 500, vector search unavailable in standalone, and FTS zero recall on Chinese questions — plus a disable-per-channel diagnostic semantic and an offline retrieval eval harness.

RetrievalBug FixesSQLiteChinese Tokenization

v1.1.1 is a fix release that the eval system dragged out of us. This cycle we built a quantitative retrieval evaluation (the full story is in this post); its very first smoke run caught two standalone defects, and the first real baseline exposed the Chinese-question full-text search problem. All three fixes landed in this release.

Fixed: standalone knowledge base creation always returned 500

Creating a knowledge base from scratch in standalone (single-binary SQLite) mode returned a 500 every time: a forward foreign key on one table was missing its deferred check declaration, so the creation transaction failed outright on SQLite. Anyone bringing up an instance from scratch hits it — the eval harness played exactly that “new user” and ran into it first.

Fixed: vector search unavailable in standalone

Vector search never worked in the standalone production binary: the vec extension wasn’t linked into the build, so every call failed. The problem was in build linking, not in logic — no existing test covered the “production binary + SQLite + vector search” combination, which is why it had gone unnoticed.

Fixed: zero recall on Chinese question sentences in FTS

An everyday question like “埃及有哪些民族?” returned literally nothing from full-text search: after tokenization, FTS matches with AND semantics, and question words and punctuation never co-occur in body text, so the query is always the empty set. Hybrid retrieval silently degraded to vector-only, with no errors and no alerts.

The fix filters punctuation and function words on the query side, effective for both SQLite FTS5 and PostgreSQL. After the fix, FTS recall recovered from zero, and hybrid came out strictly above vector-only for the first time. The mechanism and a self-check guide are in the postmortem.

Channel diagnostics and the eval harness

  • vector_top_k=0 / keyword_top_k=0 now disable the corresponding recall channel, enabling single-channel diagnosis without code changes;
  • The repo ships an offline retrieval eval (make eval): MIRACL-zh dataset, dual-track four-cell channel matrix, deterministic metrics with report fingerprints — full results in RETRIEVAL_BENCHMARK.md.

Verification

All three defects were found by the eval and verified after the fix with same-fingerprint reruns: FTS recall went from 0 to 0.13, hybrid first exceeded vector-only (0.9826 over 0.9799), and passage-level MRR reached 0.9975.