Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture explores the quest for a good image-text embedding for remote sensing visual question answering, discussing various methods such as element-wise multiplication, Multimodal Compact Bilinear pooling, and Multimodal Tucker Fusion. The presentation delves into the baseline system, related works, and the results obtained from low and very high-resolution image sets.