Skip to content

feat(embed): Implements multimodal image-vector retrieval - #3909

Open
chengjoey wants to merge 1 commit into
Tencent:mainfrom
chengjoey:feat/multi-embedding
Open

chengjoey wants to merge 1 commit into
Tencent:mainfrom
chengjoey:feat/multi-embedding

Conversation

@chengjoey

@chengjoey chengjoey commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Implements multimodal image-vector retrieval

  • Add a MultimodalEmbedder interface and a multimodalEmbedder decorator (chat / sglang envelopes, one embedding request per image). Text calls still go through the catalog protocol, so only the multimodal call shape changes.
  • New chunk type image_vector plus an indexing_strategy.image_vector_enabled switch. It is decoupled from VLM: images embedded inside docs (PDF/Word/Markdown) are indexed without turning on image processing.
  • Indexing and query sides share a single usesMultimodalEnvelope predicate. A text query is encoded through the same envelope and recalled via ANN in the same collection; image_vector chunks are exempted from reranking (their content is only a markdown image link).
  • Flipping the switch on an existing KB enqueues kb:reindex_vectors to backfill image vectors.

Bug fix

writeImageVectorChunk omitted IsEnabled when constructing IndexInfo, so the image vectors were persisted with is_enabled = false and silently filtered out by every retrieval engine — written but never recallable. Fixed and covered by a regression assertion (with mutation verification).

Verification

hybrid-search with query_text="图片里有什么" successfully recalled all 4 images cross-modally.

Description

Type of Change

  • 🐛 Bug fix
  • ✨ New feature
  • 💥 Breaking change
  • 📚 Documentation update
  • 🎨 Refactor
  • ⚡ Performance improvement
  • 🧪 Test
  • 🔧 Configuration / Build / CI

Related Issue

feat: #2875

Checklist

  • git diff --check origin/main...HEAD passes
  • Changed source files are formatted
  • Targeted tests for the changed packages/components pass
  • Diff-scoped lint passes where applicable (for Go: golangci-lint run --new-from-rev=origin/main ./...)
  • Full-repository checks were run, or any unrelated/environment-dependent failures are documented above
  • Self-reviewed the code
  • Added/updated tests covering the change
  • Updated related documentation (README, website-docs/, Swagger annotations, etc.)
  • Breaking changes are clearly called out in the description above

Screenshots / Recordings

image image image

@chengjoey
chengjoey force-pushed the feat/multi-embedding branch 2 times, most recently from e190989 to 5f24b28 Compare September 30, 2026 09:45
Implements multimodal image-vector retrieval (issue Tencent#2875):

- Add a `MultimodalEmbedder` interface and a `multimodalEmbedder` decorator
  (chat / sglang envelopes, one embedding request per image). Text calls still
  go through the catalog protocol, so only the multimodal call shape changes.
- New chunk type `image_vector` plus an `indexing_strategy.image_vector_enabled`
  switch. It is decoupled from VLM: images embedded inside docs (PDF/Word/Markdown)
  are indexed without turning on image processing.
- Indexing and query sides share a single `usesMultimodalEnvelope` predicate.
  A text query is encoded through the same envelope and recalled via ANN in the
  same collection; `image_vector` chunks are exempted from reranking (their
  content is only a markdown image link).
- Flipping the switch on an existing KB enqueues `kb:reindex_vectors` to backfill
  image vectors.

## Bug fix

`writeImageVectorChunk` omitted `IsEnabled` when constructing `IndexInfo`, so the
image vectors were persisted with `is_enabled = false` and silently filtered out
by every retrieval engine — written but never recallable. Fixed and covered by a
regression assertion (with mutation verification).

Signed-off-by: joeyczheng <joeyczheng@tencent.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant