Skip to content
CAS
Multimodal VisionOngoingCAS Research · Internal Engineering Demonstration

Multimodal Search

Cross-modal search over documents containing layout, text and figures — measuring where aligned encoders break on real page structures.

Findings so far

  • Layout-heavy pages degrade naive chunking; structure-aware parsing recovered most of the loss.
  • Joint embedding space search underperformed late-fusion on table-heavy documents.

Stack

Vision encoderText encoderAlignment projectionHybrid search

Architecture in ModLens

Multimodal Architecture
Ask CAS