NED University Journal of Research
ISSN 2304-716X
E-ISSN 2706-5758




MULTI-FRAME TRANSFORMER-BASED COMMUNICATION SIGNAL MODELING FOR WORD PREDICTION IN APHASIA

Author(s): Nur Syahmina Ahmad Azhar1, Nik Mohd Zarifie Hashim2 *, Mohd Norzali Hj Mohd3 *, David Lim4
1PhD Candidate, Fakulti Teknologi dan Kejuruteraan Elektronik dan Komputer, Universiti Teknikal Malaysia Melaka, Jalan Hang Tuah Jaya, 76100 Durian Tunggal, Melaka, Malaysia, Ph. +601127853983, Email: p122410008@student.utem.edu.my

2Senior Lecturer, Fakulti Teknologi dan Kejuruteraan Elektronik dan Komputer, Universiti Teknikal Malaysia Melaka, Jalan Hang Tuah Jaya, 76100 Durian Tunggal, Melaka, Malaysia, Ph. +60176129997, Email: nikzarifie@utem.edu.my

3Associate Professor, Faculty of Electrical and Electronic Engineering, Universiti Tun Hussein Onn, Malaysia, Ph. +60137475869, Email: norzali@uthm.edu.my

4Key Technical Figure, Sena Traffic Systems Sdn Bhd, Malaysia, Email: david.lim@senatraffic.com.my

https://doi.org/10.35453/NEDJR-Icon3E2025-014-R1

Volume: 23

No. Special issue on Icon3E'25

Pages: 288-307

Date: August 2026

Publication Type: Open-Access Publication

Abstract:
Aphasia affects a person's ability to retrieve words and coordinate the movements of the speech articulators, resulting in impaired speech production and pronunciation. Current word prediction tools do not use lip movement patterns, which limits their usefulness for people with speech disorders. This study aims to predict the word a person intends to say by analyzing both lexical activation and articulatory stability. A Multi-Frame Transformer (MFT) model that processes five consecutive frames of lip movements is developed, along with a proposed Single-Frame Transformer (SFT) and three pre-trained comparative models including BERT, RoBERTa, and a 3D CNN. The proposed MFT model achieved an 84.50% accuracy, outperforming BERT with 63.33%, the 3D CNN with 51.11%, the Single-Frame Transformer with 26.67%, and RoBERTa with11.11%. The proposed model maintained 74% accuracy when word-finding information was missing and 58% accuracy when lip movement information was missing. These findings show that motor signals are more critical for correct prediction. The Multi-Frame Transformer (MFT) model can effectively predict intended words from lip movement patterns, offering a foundation for future communication aids for people with aphasia.

Keywords:
Aphasia, Articulatory Stability, Lexical Activation, Multi-Frame Processing, Perceptron, Single-Frame Processing, Transformer

Full Paper | Close Window      X |