Disaster Recognition Through Image Captioning Features and Shifted Attention
Abstract
The author proposes a novel methodology that integrates image captioning with a custom-developed attention-shifting algorithm designed to dynamically refocus the model on less conspicuous yet essential elements within images. By leveraging the inherent strengths of Vision Encoder-Decoder (VED) models, along with innovative optimal masking strategies, we enable the system to discern and articulate the specifics of disaster impacts in diverse imaging conditions, from satellite to ground-level perspectives. The empirical results underscore the superiority of our approach over conventional image captioning models, exhibiting enhanced detection capabilities with accuracies exceeding 91% for landslide detection from side-view image captions and 87.5% for shipborne view detection. These figures not only reflect the technical prowess of the system but also its practical applicability in real-world disaster assessment scenarios. This work carries profound implications for the field of disaster management. By augmenting the quality and reliability of disaster region identification, our framework facilitates more informed decision-making in allocating resources for relief efforts. Additionally, the adaptive nature of the model paves the way for its application across a spectrum of environmental monitoring and emergency response tasks, heralding a new era of AI-enabled disaster management tools. Future research avenues include scaling the model to encompass a broader range of disaster types and integrating real-time data for swift, actionable insights during crisis events.