Back to releases
v1.2.0

Retrieval quality closure and MCP agent ergonomics (v1.2.0)

Four open retrieval directions closed by controlled experiments (semantic chunking, LLM context headers, chunk parameters, traditional/simplified normalization), then progressive disclosure for search results (detail full/lean) and discovery tools for agents (MCP tools 7→10); matched_children no longer carries child content.

MCPRetrievalAgentsEvaluation

v1.2.0 answers two questions: is the retrieval good enough? And is it ergonomic for agents? The answer to the first is “yes, with evidence” — the four most intuitive optimization paths were all closed by controlled experiments. The answer to the second is this release’s main feature: progressive disclosure for search results, plus environment discovery for agents.

Retrieval quality closure: four directions, judged by experiments

This cycle we extended the eval system to unstructured continuous text (Chinese meeting transcripts, the VCSUM dataset, make eval-vcsum), then ran a series of single-variable experiments. Every conclusion is backed by same-fingerprint reports (RETRIEVAL_BENCHMARK.md):

  • Semantic chunking is formally excluded. An oracle experiment using human-annotated topic boundaries — the theoretical ceiling of any semantic chunker — measured the upper bound at ±2 percentage points. No implementation can exceed it; the direction is closed.
  • LLM context-header generation: not pursued. Human topic titles in context headers reach the ceiling (+7.9pp vector / +36.7pp FTS), but both a local 7B model and a cloud flash-tier model captured only ~12–14% of the FTS gain. Crossing model tiers moved nothing — the bottleneck is the mismatch between summarization phrasing and question phrasing, not model capability.
  • Traditional/simplified normalization: closed. Measured zero benefit for hybrid/rerank (the vector channel was already immune); the remaining long-document gap is 100% intra-document section competition.
  • Net result: 98% recall at paragraph level, 94% on unstructured long documents with rerank enabled. The recommended configuration is hybrid + rerank.

Progressive disclosure: detail: full | lean

A single knowledge_search used to return top 10 × full parent chunks — roughly 45k characters per call, quickly polluting an agent’s context window. Search now supports two response levels (REST and MCP both default to full; lean is opt-in):

  • full: complete parent content plus matched-children metadata (a superset of today’s behavior);
  • lean: each hit returns only the best matched child’s content (evidence — the immediate “why it matched” proof), with the parent content dropped and its chunk_id kept as a drill-down handle — combined with the existing chunk_get, agents scan evidence first, then fetch full context on demand.

Measured: lean top 10 serializes to about one third of full (13.7k vs 41k+ characters). This is the same progressive-disclosure pattern as Claude Code’s Grep/Read: show “where and why it matched” first, let the agent decide what to read.

Breaking change: matched_children no longer carries content

matched_children[].content is removed from the REST/MCP contracts. It was structurally redundant — a parent chunk is assembled by concatenating its children, so child content is always a substring of parent content. Position markers, chunk IDs, and per-channel scores are all retained. Callers that need child content should use detail=lean’s evidence or chunk_get.

MCP discovery tools: 7 → 10

Agents dropped into a workspace can now figure out the environment on their own:

  • knowledge_base_list: list accessible knowledge bases;
  • document_list: page through a base’s documents (title/kind/status);
  • document_get: read a document’s normalized full text and section outline (with max_chars truncation).

All read-only, reusing existing services, zero schema migrations. The agent’s full work loop is now: discover (list) → search (lean) → read (document_get / chunk_get) → search again.

Verification

go test ./... (50 packages), make test-integration (51 packages, ephemeral docker PostgreSQL), and the web console’s pnpm check / test / build (333 tests) all pass; ordering consistency between detail levels, the lean payload budget, and the contract field removal each have dedicated assertions.