Edge AI-Based Object Detection via Voice Recognition with an LLM-Based Emotional Assistant for Elderly Care Robots

PDF (1429KB), PP.50-67

Views: 0 Downloads: 0

Author(s)

Sarra Ben Halima 1,* Faten Ben Abdallah 1 Joseph Haggege 1

1. Automation Research Laboratory (LARA), LR11ES18, National Engineering School of Tunis, University of Tunis El Manar, 1002 Tunis, Tunisia

* Corresponding author.

DOI: https://doi.org/10.5815/ijitcs.2026.04.04

Received: 16 Feb. 2026 / Revised: 29 Apr. 2026 / Accepted: 27 May 2026 / Published: 8 Aug. 2026

Index Terms

Elderly Care, Edge AI, NVIDIA Jetson Nano, Speech Recognition, Generative AI, Object Detection, Embedded Systems, Assistive Robotics, LLM-based Emotional Assistant

Abstract

This paper presents a fully integrated, real-time assistive system that combines voice-based object recognition with a generative conversational interface, specifically designed to enhance elderly care through edge AI deployment. The proposed framework enables intuitive human–robot interaction in domestic environments by fusing natural language understanding, optimized visual detection, and local generative response. Voice commands are processed through a speech-to-text pipeline using the Google Web Speech API, with keyword extraction triggering object detection via a quantized YOLOv8n model accelerated through TensorRT with FP16 inference on an NVIDIA Jetson Nano. In parallel, a locally deployed generative AI assistant, executed entirely on-device, provides empathetic dialogue to support social engagement and emotional well-being. The proposed system adopts a hybrid edge architecture in which object detection, robot control, and LLM-based dialogue generation are executed on-device, while speech-to-text transcription relies on a cloud-based service. This generative interface is implemented as an LLM-based Emotional Assistant. The system achieves 13 FPS with an inference latency of 70 ms for object detection, 94.3% speech recognition accuracy, and an F1-score of 0.69 at a 0.5 confidence threshold. All AI components are executed on-board, preserving privacy for on- device processing while maintaining real-time responsiveness. Experimental validation confirms the effectiveness of deploying multimodal AI, including generative models, on resource-constrained hardware. This work lays the foundation for autonomous, voice-guided care robots that not only assist in locating objects but also engage users socially, promoting greater autonomy and quality of life for older adults.

Cite This Paper

Sarra Ben Halima, Faten Ben Abdallah, Joseph Haggege, "Edge AI-Based Object Detection via Voice Recognition with an LLM-Based Emotional Assistant for Elderly Care Robots", International Journal of Information Technology and Computer Science(IJITCS), Vol.18, No.4, pp.50-67, 2026. DOI:10.5815/ijitcs.2026.04.04

Reference

[1]World Health Organization. Mental Health of Older Adults. https://www.who.int/news-room/ fact-sheets/detail/mental-health-of-older-adults, 2021.
[2]National Academies of Sciences, Engineering, and Medicine. Social Isolation and Loneliness in Older Adults: Opportunities for the Health Care System. The National Academies Press, Washington, DC, 2020.
[3]G. Peleka et al. RAMCIP—A service robot for MCI patients at home. In Proc. IEEE/RSJ IROS, Madrid, Spain, 2018.
[4]P. Asgharian, A. M. Panchea, and F. Ferland. A Review on the Use of Mobile Service Robots in Elderly Care. Robotics, 11(6):127, Nov. 2022.
[5]S. Ben Halima, F. Ben Abdallah, and J. Haggege. Edge AI-based fall detection for an elderly care robot. In Proc. IEEE 22nd Int. Conf. on Sciences and Techniques of Automatic Control and Computer Engineering (STA). IEEE, 2025.
[6]P. H. Varshini, V. Swati, D. Navya, and A. M. K. Aiswariya. EchoVision: RealTime Object Detection and Voice Assistance for the Visually Impaired. In Proc. 3rd Int. Conf. on Intelligent Systems, Advanced Computing and Communication (ISACC), pages 1–7, 2025.
[7]S. Cos, ar et al. ENRICHME: Perception and Interaction of an Assistive Robot for the Elderly at Home. Int. J. Soc. Robot., 12:779–805, 2020.
[8]A. K. Pandey and R. Gelin. A mass-produced sociable humanoid robot: Pepper: The first machine of its kind. IEEE Robot. Autom. Mag., 25:40–48, 2018.
[9]F. Fracasso et al. Social Robots Acceptance and Marketability in Italy and Germany: A Cross-National Study Focusing on Assisted Living. Int. J. Soc. Robot., 14:1463–1480, 2022.
[10]F. Cavallo et al. Robotic services acceptance in smart environments with older adults: User satisfaction and accept- ability study. J. Med. Internet Res., 20(9):e264, 2018.
[11]G. A. Sai et al. Visio-Voice Transforming Images into Sound for the Visually Impaired. In Proc. IEEE Int. Conf. on Information Technology, Electronics and Intelligent Communication Systems (ICITEICS), pages 1–7, 2024.
[12]A. K. Luke et al. Envision: Assistance System for the Visually Impaired. In Proc. 14th Int. Conf. on Computing Communication and Networking Technologies (ICCCNT), pages 1–8, 2023.
[13]R. K. Megalingam, P. T. K. Sai, A. Ashvin, P. N. Reddy, and B. R. Gamini. Trinetra App: A Companion for the Blind. In Proc. IEEE Bombay Section Signature Conf. (IBSSC), pages 1–5, 2021.
[14]P. J. Duh et al. V-Eye: A Vision-Based Navigation System for the Visually Impaired. IEEE Trans. Multimedia, 23:1567–1580, 2020.
[15]G. Yang and J. Saniie. Sight-to-Sound Human-Machine Interface for Guiding and Navigating Visually Impaired People. IEEE Access, 8:185416–185428, 2020.
[16]R. C. Joshi, N. Singh, A. K. Sharma, R. Burget, and M. K. Dutta. AI-SenseVision: A Low-Cost Artificial- Intelligence-Based Robust and Real-Time Assistance for Visually Impaired People. IEEE Transactions on Human- Machine Systems, 2024.
[17]R. C. Joshi et al. AI-SenseVision: A Low-Cost Artificial-Intelligence-Based Robust and Real-Time Assistance for Visually Impaired People. IEEE Trans. Human-Machine Systems, 2024.
[18]G. Agarwal, K. Jindal, A. Chowdhury, V. K. Singh, and A. Pal. Image and Video Captioning for Apparels Using Deep Learning. IEEE Access, 2024.
[19]J. P. Docto, A. I. Labininay, and J. F. Villaverde. Third Eye Hand Glove Object Detection for Visually Impaired Using YOLO v4-Tiny Algorithm. In Proc. IEEE Int. Conf. on Artificial Intelligence in Engineering and Technology (IICAIET), pages 1–6, 2022.
[20]D. Kleinberg, R. Yozevitch, I. Abekasis, Y. Israel, and E. Holdengreber. A Haptic Feedback System for Spatial Orientation in the Visually Impaired: A Comprehensive Approach. IEEE Sensors Letters, 7(9):1–4, 2023.
[21]M. Osama, A. Yehia, S. Mohamed, R. Sherief, N. Elmasry, V. Adel, and A. Hamdy. Design and Implementation of Visually Impaired Assistant System. In Proc. Int. Mobile, Intelligent, and Ubiquitous Computing Conf. (MIUCC), pages 303–310, 2021.
[22]I. A. R. Djinko and T. Kacem. Video-based Object Detection Using Voice Recognition and YOLOv7. In Proc. Conf. on Artificial Intelligence and Computer Vision. University of the District of Columbia, 2025.
[23]M. Gygli and V. Ferrari. Fast Object Class Labelling via Speech. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 15365–15374. IEEE, 2019.
[24]L. Ramesh, P. M. S. Pruthiga, M. Muthulakshmi, P. S. N. Sri, and S. Navaneetha. Voice Assistance for Visually Impaired Persons Using AI and IoT. In Proc. 11th Int. Conf. on Bio Signals, Images, and Instrumentation (ICBSII). IEEE, 2025.
[25]R. Miyata, N. Yamaguchi, O. Fukuda, and H. Okumura. Object Search Using Edge-AI Based Mobile Robot. In Proc. 6th Int. Conf. on Intelligent Informatics and Biomedical Sciences (ICIIBMS). IEEE, 2021.
[26]A. Gordeev, V. Klyachin, E. Kurbanov, and A. Driaba. Autonomous mobile robot with AI based on jetson nano. In Proc. Future Technologies Conf. (FTC) 2020, Vol. 1, pages 190–204. Springer, 2021.
[27]Anthony Zhang (Uberi). Speechrecognition: Speech recognition module for python. https://pypi.org/ project/SpeechRecognition/, 2023. Accessed: 2025-07-17.
[28]Google Cloud. Cloud speech-to-text documentation. https://cloud.google.com/speech-to-text/ docs, 2025. Accessed: 2025-07-17.
[29]Hubert Pham and PyAudio Contributors. Pyaudio: Python bindings for portaudio. https://pypi.org/ project/PyAudio/, 2023. Accessed: 2025-07-17.
[30]P. Warden. Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition. arXiv preprint arXiv:1804.03209, 2018.
[31]T. Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dolla´r, and C. L. Zitnick. Microsoft COCO: Common Objects in Context. In Proc. European Conf. on Computer Vision (ECCV), pages 740–755. Springer, 2014.
[32]S. Han, H. Mao, and W. J. Dally. Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. arXiv preprint arXiv:1510.00149, 2015.
[33]G. Hinton, O. Vinyals, and J. Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
[34]B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 2704–2713, 2018.
[35]Y. Zhou and K. Yang. Exploring TensorRT to Improve Real-Time Inference for Deep Learning. In Proc. IEEE 24th Int. Conf. on High Performance Computing & Communications (HPCC), 8th Int. Conf. on Data Science & Systems, 20th Int. Conf. on Smart City, and 8th Int. Conf. on Dependability in Sensor, Cloud & Big Data Systems & Application (HPCC/DSS/SmartCity/DependSys), pages 2011–2018, 2022.
[36]I. Noreen, A. Akbar, and U. Siddiqui. AI-Enabled Elderly Care Robot. J. Inf. Commun. Technol. Robot. Appl., 11(2):22–29, 2020.
[37]R. K. Megalingam et al. Trinetra App: A Companion for the Blind. In Proc. IEEE Bombay Section Signature Conf. (IBSSC), pages 1–5, 2021.
[38]K. Jivrajani et al. AIoT-Based Smart Stick for Visually Impaired Person. IEEE Trans. Instrum. Meas., 72:1–11, 2022.