Fields · Physical Sciences · Computer Science · Computer Vision and Pattern Recognition

Multimodal Machine Learning Applications

This cluster of papers focuses on the development and improvement of visual question answering systems, image captioning techniques, and neural networks for understanding and generating descriptions of images and videos. The research involves semantic reasoning, multimodal fusion, scene graph generation, attention mechanisms, and deep learning approaches to bridge the gap between vision and language.

67,398 works

Papers listed on taxonomy pages are the top few works per node from the OpenAlex snapshot. That list is not exhaustive and is not an endorsement. The topic map and the journal registry remain separate: there is still no authoritative topic-to-venue or topic-to-organization edge. Search is a lexical lookup, not a claim that a venue publishes a topic.

Search papers

Keywords