article
Visual impairment significantly challenges an individual's ability to navigate and interact with their environment independently. While retinal prostheses have provided partial vision restoration, they offer limited support for complex visual tasks such as object recognition, text reading, and real-time navigation. Recent advancements in technology have introduced Large Language Models (LLMs) as promising tools to enhance these assistive technologies, potentially enhancing the capabilities of existing devices. This study explores the integration of LLMs into assistive devices for visually impaired individuals, evaluating several models, including MiniCPM-Llama3-V-2_5, PaliGemma, Blip2, Microsoft Git Large COCO, and LLaVA-Next. The models are assessed in tasks like image captioning, object detection, and optical character recognition (OCR), focusing on their effectiveness in enhancing user independence and quality of life. Our findings indicate that MiniCPM-Llama3-V-2_5 and PaliGemma excel in providing accurate, detailed, and practical outputs, crucial for real-world applications. These models demonstrate significant potential in transforming retinal prostheses from passive aids into interactive tools that enhance user engagement with their surroundings. The study concludes with discussions on the implications of these findings and suggests future research directions to further improve the integration of LLMs in assistive technologies for visually impaired individuals.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/niles63360.2024.10753262
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.