Skip to content
AI.info

Research

AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired

Overview Research area: Assistive technology and human–computer interaction, specifically AI-based mobile assistance for people with visual impairments, built on computer vision (object detection, vis

AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired
arXiv
2511.06080
Published
2025-11-08
Authors
Luis Marquez-Carpintero, Francisco Gomez-Donoso, Zuria Bauer, Bessie Dominguez-Dager, Alvaro Belmonte-Baeza, Mónica Pina-Navarro, Francisco Morillas-Espejo, Felix Escalona, Miguel Cazorla

AI summary

Overview

  • Research area: Assistive technology and human–computer interaction, specifically AI-based mobile assistance for people with visual impairments, built on computer vision (object detection, vision-language models) and accessibility engineering.
  • Technical level: Intermediate. The components (YOLOv8 detection, LLaVA-based visual question answering and OCR, a client–server mobile app, haptic feedback design) are standard and clearly described; no new model architecture is proposed.
  • Scope: One sentence — the paper describes the design, integration, and a 28-participant pilot evaluation of AIDEN, a smartphone-based assistant that routes spatial guidance to haptic feedback while reserving audio for semantic content such as OCR and visual question answering.

What This Paper Is About

Visually impaired users need help with object identification, text reading, and locating things in unfamiliar places, but most existing assistive apps push nearly all guidance through audio, which can overload the user, mask critical environmental sounds, and require them to switch between fragmented single-purpose tools. AIDEN is a smartphone-first assistant that unifies optical character recognition, object description/visual question answering, and real-time object finding into one workflow, deliberately moving spatial guidance onto continuous haptic (vibration) feedback and keeping audio for semantic content, while storing no personal data.

Key Contributions

  1. Accessibility-first multimodal interaction: A practical interaction paradigm that combines screen-reader-compatible navigation with continuous haptic guidance for object retrieval, intended to reduce auditory overload and preserve situational awareness.
  2. End-to-end system integration: A cross-platform mobile application that unifies OCR, scene understanding (Object Description/VQA), and object finding inside a smartphone-first client–server pipeline optimized for resource-constrained devices and guided by data minimization.
  3. Empirical evidence with visually impaired users: Results from a pilot evaluation with 28 visually impaired participants, including Technology Acceptance Model (TAM) usability outcomes and technical runtime performance measurements.
  4. A comparative analysis against leading commercial systems (Seeing AI, Be My Eyes, Envision AI, Meta Ray-Ban) across criteria derived from identified research gaps: haptic and audio guidance, object search, privacy and data handling, VQA capability, and smartphone-only deployment.

Main Findings

  • High TAM ratings overall: Participants rated AIDEN highly, with average ratings falling between the "Excellent" and "Best" categories for most of the 16 questionnaire items, using the mapping of Very Poor (1.0–1.25), Poor (1.25–2.5), Acceptable (2.5–3.5), Good (3.5–4.0), Excellent (4.0–4.5), and Best (4.5–5.0).
  • Top-rated item: The highest mean score was for willingness to incorporate the final version of the system into daily routines (Q14), followed by the overall positive evaluation of the system (Q16).
  • Favorable adoption intention: High ratings were also obtained for intention to use the system in real-world settings and to recommend it to others (Q13 and Q15). The exact mean and standard deviation for individual questions are not given numerically in the available text; they are presented only in a figure.
  • Object finding was by far the fastest function: "Find an Object" (one request) took 0.23 ± 0.18 s server-side and 0.51 ± 0.22 s on the smartphone.
  • Semantic tasks were slower and similar to each other: Object Description & QA took 7.34 ± 1.10 s on the server and 10.10 ± 1.29 s on the smartphone; OCR took 7.09 ± 2.57 s on the server and 9.54 ± 2.94 s on the smartphone.
  • Real-time framing rate: The "Find an Object" module achieved an average processing speed of 1.96 frames per second on the mobile device under the tested conditions, described as low-latency performance.
  • A short adaptation period is needed for haptics: Users initially struggled to interpret the varying vibration frequencies, but after the third attempt in the training phase they displayed a smoother scanning motion.
  • Distance estimation was a recurring practical problem: In the OCR task, some participants held the document too close to the lens, while others held it too far away or captured a broad area of unnecessary text.
  • Competitive positioning: In the authors' hands-on comparison, AIDEN was the only system marked as offering full integration across haptic guidance, audio guidance, object search, privacy/data, VQA capability, and smartphone-only deployment. Seeing AI and Envision AI were marked partial on object search (post-photo/finger interaction and non-orientation guidance, respectively) and partial on privacy/data (subscription-dependent); Be My Eyes and Meta Ray-Ban were marked as lacking haptic guidance, object search, and VQA.
  • Sample and study size: 28 participants (15 male, 13 female); 14 with total blindness, 10 with profound visual impairment, and 4 with severe visual impairment; 20 on iOS and 8 on Android.
  • Explicit limitation: The authors state the sample size is insufficient for strong inferential or broadly generalizable claims, and that the statistical analyses should be interpreted as exploratory rather than confirmatory.
  • Baseline comparisons: No comparisons against other systems' measured runtime or usability scores are reported; the comparison is feature-based, not performance-based.

Methodology in Plain English

The team built a cross-platform mobile app (Ionic v6.20.1, Capacitor Core v4.7.0, Vue.js v3.3.7) that runs on both Android and iOS, with accessibility handled through the Capacitor ScreenReader plugin and the Vue i18n library for multilingual support. The interface was designed screen-reader-first — interactive elements carry semantic labels so users navigate by standard gestures rather than by visually locating buttons — and the design targets WCAG 2.1 Level AA, with high-contrast colors and scalable typography for users with residual vision.

Because running deep learning models on a phone would demand high-end hardware and hurt accuracy, the app uses a distributed client–server architecture: the phone only captures images and presents results, while heavy models run on a server (Ubuntu 18.04 LTS, Intel i7-8700 CPU, 32 GB RAM) with two GPUs — an NVIDIA GTX 1080Ti dedicated to YOLOv8 for the low-latency object-finding path, and an NVIDIA A40 reserved for LLaVA-based VQA and OCR. Requests are handled first-in-first-out, with "Find an Object" prioritized to minimize round-trip time. The app was deployed on an LG G6 smartphone (Qualcomm Snapdragon 821, 4 GB RAM, 5.7-inch QHD+ display) over 4G LTE.

Three functions were built. OCR sends an image to LLaVA with the prompt "Transcribe the text present in this image." Object Description/VQA uses the fixed prompt "What is this? Provide the answer as summarized as possible," with optional follow-up natural-language questions. "Find an Object" uses YOLOv8 (80 standard categories, extended with an extra "door" class) and guides the user in two stages: coarse directional speech ("Move camera to the left," "Tilt up") when the object is in the periphery, then continuous audio–haptic feedback once the object is near center. That feedback follows a Geiger-counter metaphor: pulse frequency is inversely proportional to the Euclidean distance between the object centroid and the camera's optical center. Peripheral detection uses intermittent pulses (200 ms ON, 300 ms OFF), approaching uses rapid pulsing (200 ms ON, 100 ms OFF), and a centered object produces continuous vibration plus a confirmation tone.

Evaluation had two parts. The runtime study executed each of the three core functionalities 20 times, recording server processing time and total smartphone-observed response time (including transmission and retrieval). The user study ran in a controlled laboratory with constant lighting and minimized background noise; each of the 28 participants (recruited through ONCE and Retina Comunidad Valenciana, none involved in development, all with at least intermediate smartphone and screen-reader experience) spent roughly 25 minutes across three phases: training and familiarization, task execution (object retrieval, information access via OCR, and scene understanding via VQA, in counterbalanced order), and data collection. A Think-Aloud protocol captured verbal feedback, and a task counted as successful if the goal was achieved without physical intervention. Subjective perception was measured with a 16-item TAM questionnaire on a 5-point Likert scale via Google Forms.

Why This Matters

Impact on research. The paper argues that the field forces a trade-off between functional assistance and sensory safety, and it demonstrates a design response: offload spatial guidance to haptics so the auditory channel stays free for semantic content and environmental cues. It also contributes a concrete, feature-by-feature comparison of commercial state-of-the-art systems against a research prototype, and pilot TAM evidence from a low-incidence population — a group the authors note is hard to recruit and underrepresented in evidence that is limited in quantity and quality, unevenly distributed, and largely observational.

Real-world applications:

  • Everyday object retrieval — locating keys, a cup, a backpack, or a door in cluttered or unfamiliar spaces, with vibration rather than speech guiding the camera.
  • Reading printed information — transcribing labels, documents, medicine packaging, and street signage, with example outputs including product labels, a medication box, a fabric softener label, and a street name.
  • Scene understanding — asking what is in front of the camera or asking follow-up questions such as whether a door is open.
  • Low-vision support — high-contrast schemes and scalable typography for users with residual vision who can perceive forms and shapes.

Industry relevance. The work matters for developers of commercial assistive apps (Seeing AI, Be My Eyes, Envision AI) and for wearable players (Meta Ray-Ban, HoloLens-based prototypes such as MR.NAVI): it argues that a smartphone-first architecture is more affordable, more socially acceptable, and avoids the power consumption and battery-life limits of head-mounted devices. It also argues for a privacy posture — no personal data stored, data minimization, no cloud-dependent capture the user cannot visually verify — as an adoption factor rather than a compliance afterthought, and shows a client–server split as a way to deliver heavy models on cheap phones.

Future Directions

  • Larger-scale validation. The authors explicitly motivate larger-scale validations; the current 28-participant pilot is described as insufficient for inferential or generalizable claims, so confirmatory studies are the natural next step.
  • Improving distance estimation and capture framing. The observed tendency to hold documents too close or too far in the OCR task, or to capture excess text, points to guidance or automatic framing support during capture.
  • Shortening the haptic learning curve and smoothing the handoff. Users needed roughly three attempts to interpret varying vibration frequencies, and the system switches between verbal coarse navigation and continuous fine feedback — refining that transition and the training phase is an open design question.
  • Addressing the semantic-task latency gap. OCR and Object Description took on the order of 7 s server-side and 9.5–10 s on the phone, while Find an Object took fractions of a second; narrowing that gap is an obvious engineering target.
  • Robustness in uncontrolled conditions. The pilot ran in a laboratory with constant lighting and minimized background noise, so performance in noisy, variable real-world environments remains untested here.

Target Audience

Researchers and practitioners in assistive technology, accessibility, and human–computer interaction; mobile and computer vision engineers building on-device or client–server AI assistants; usability researchers who study small, specialized, low-incidence user populations; and product teams at companies developing commercial assistive apps or smart eyewear who want a concrete comparison of feature trade-offs and pilot acceptance data.

Authors’ abstract

This paper presents AIDEN, an artificial intelligence-based assistant designed to enhance the autonomy and daily quality of life of visually impaired individuals, who often struggle with object identification, text reading, and navigation in unfamiliar environments. Existing solutions such as screen readers or audio-based assistants facilitate access to information but frequently lead to auditory overload and raise privacy concerns in open environments. AIDEN addresses these limitations with a hybrid architecture that integrates You Only Look Once (YOLO) for real-time object detection and a Large Language and Vision Assistant (LLaVA) for scene description and Optical Character Recognition (OCR). A key novelty of the system is a continuous haptic guidance mechanism based on a Geiger-counter metaphor, which supports object centering without occupying the auditory channel, while privacy is preserved by ensuring that no personal data are stored. Empirical evaluations with visually impaired participants assessed perceived ease of use and acceptance using the Technology Acceptance Model (TAM). Results indicate high user satisfaction, particularly regarding intuitiveness and perceived autonomy. Moreover, the ``Find an Object'' achieved effective real-time performance. These findings provide promising evidence that multimodal haptic-visual feedback can improve daily usability and independence compared to traditional audio-centric methods, motivating larger-scale clinical validations.

Read the original paper