MARATTO

article

MiT-RelTR: An Advanced Scene Graph-Based Cross-Modal Retrieval Model for Real-Time Capabilities

Abstract

In the current technological landscape, cross-modal retrieval systems have become essential, bridging the gap between diverse data types to boost accessibility and interaction across digital platforms. Our research enhances these systems by aiming for the efficient handling of low-resolution inputs, a common challenge in various real-life fields. This was conducted while ensuring robust performance even when high-resolution data is unavailable. The paper introduces advancement to the Local-Global Scene Graph Matching (LGSGM) architecture for cross-modal image/text retrieval, by incorporating a lightweight replacement of the scene graph generation module. The novel MiT-RelTR scene graph generation model is used to optimize the retrieval process. Our contribution improved caption retrieval by achieving a 0.4% increase in Recall@10, which signifies boosted accuracy in processing textual data. Conversely, it resulted in a decline in the image retrieval Recall@10 by 0.9%. Nonetheless, the system's inference speed improved notably, with a 38% increase in frames per second (FPS), bolstering its fitness for real-time applications. These findings illustrate the trade-offs and benefits of refining system components and suggest a need for balanced optimization strategies that equally benefit all modalities.

Research topics

  • Advanced Image and Video Retrieval Techniques
  • Multimodal Machine Learning Applications
  • Image Retrieval and Classification Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/miucc62295.2024.10783495

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.