Sign language recognition and reasoning pipeline for human-robot interaction

Laen...
Pisipilt

Kuupäev

Ajakirja pealkiri

Ajakirja ISSN

Köite pealkiri

Kirjastaja

Tartu Ülikool

Abstrakt

Modern technologies enable seamless interactions between humans and machines through voice recognition, as seen in in-car multimedia systems and smart speakers. Such benefits are not accessible to deaf and hard-of-hearing people. Sign Language Recognition (SLR) translates sign language into text, widening the means of communication. However, many studies primarily focus on improving recognition accuracy, and the solutions are usually not deployed in real life to help people. Additionally, deaf and hard-of-hearing people are usually not well involved in their development or performance evaluation, and this gap has repeatedly raised questions about the real-world applicability of such studies. This thesis aims to develop a sign-language-controllable robot. First, signers sign in front of the robot’s camera, and MediaPipe, a machine learning framework for real-time hand and pose landmark detection, is used to estimate participants’ skeletal models during sign language performance. Then, the SLR model, which recognises American Sign Language (ASL) letters, is utilised to construct sentences letter by letter. Minor distortions are corrected by Gemma 3n, a locally deployed lightweight Large Language Model (LLM). Lastly, Gemma 3n also performs context extraction, and the robot moves accordingly. The entire pipeline was implemented on ROS 2. 15 videos were collected from 3 native signers performing sentences in everyday places such as kitchens and living rooms. The average Character Error Rate (CER) was 0.079, and the plan success rate and execution success rate both reached 1.00. In addition, the in-person user study with one of the native signers in a domestic environment showed that the average CER was 0.254. Despite such circumstances, the plan and execution success rates achieved 0.70. This thesis can be further developed for real-world robot applications to make technologies more accessible to deaf and hard-of-hearing people.

Kirjeldus

Märksõnad

Sign Language Recognition (SLR), Large Language Model (LLM), Robotics, Robot Operating System (ROS)

Viide