Multimodal VisionOngoingCAS Research · Internal Engineering Demonstration
Multimodal Search
Cross-modal search over documents containing layout, text and figures — measuring where aligned encoders break on real page structures.
Findings so far
- Layout-heavy pages degrade naive chunking; structure-aware parsing recovered most of the loss.
- Joint embedding space search underperformed late-fusion on table-heavy documents.
Stack
Vision encoderText encoderAlignment projectionHybrid search
