WWW.JAMRIS.ORG
pISSN 1897-8649 (PRINT)/eISSN 2080-2145 (ONLINE)
VOLUME 20, N° 3, 2026
Indexed in SCOPUS
AI-generated illustration
Journal of Automation, Mobile Robotics and Intelligent Systems A peer-reviewed quarterly focusing on new achievements in the following fields: • automation • systems and control • autonomous systems • multiagent systems • decision-making and decision support • • robotics • mechatronics • data sciences • new computing paradigms • Editor-in-Chief
Typesetting
Janusz Kacprzyk (Polish Academy of Sciences, Łukasiewicz-PIAP, Poland)
Paradigm Publishing Services, reference-global.com
Advisory Board
Webmaster
Dimitar Filev (Research & Advenced Engineering, Ford Motor Company, USA) Kaoru Hirota (Beijing Institute of Technology, China) Witold Pedrycz (ECERF, University of Alberta, Canada)
TOMP, www.tomp.pl
Editorial Office
Co-Editors Roman Szewczyk (Warsaw University of Technology, Systems Research Institute Polish Academy of Sciences, Poland) Oscar Castillo (Tijuana Institute of Technology, Mexico) Marek Zaremba (University of Quebec, Canada)
ŁUKASIEWICZ Research Network – Industrial Research Institute for Automation and Measurements PIAP Al. Jerozolimskie 202, 02-486 Warsaw, Poland (www.jamris.org) tel. +48-22-8740109, e-mail: office@jamris.org The reference version of the journal is e-version. Printed in 100 copies. Articles are reviewed, excluding advertisements and descriptions of products.
Executive Editor Katarzyna Rzeplinska-Rykała, e-mail: office@jamris.org (Łukasiewicz-PIAP, Poland)
Associate Editor Piotr Skrzypczyński (Poznań University of Technology, Poland)
Statistical Editor
Papers published currently are available for non-commercial use under the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 (CC BY-NC-ND 4.0) license. Details are available at: https://www.jamris.org/index.php/JAMRIS/ LicenseToPublish Open Access.
Małgorzata Kaliczyńska (Łukasiewicz-PIAP, Poland)
Editorial Board: Chairman – Janusz Kacprzyk (Polish Academy of Sciences, Łukasiewicz-PIAP, Poland) Plamen Angelov (Lancaster University, UK) Adam Borkowski (Polish Academy of Sciences, Poland) Wolfgang Borutzky (Fachhochschule Bonn-Rhein-Sieg, Germany) Bice Cavallo (University of Naples Federico II, Italy) Chin Chen Chang (Feng Chia University, Taiwan) Jorge Manuel Miranda Dias (University of Coimbra, Portugal) Andries Engelbrecht ( University of Stellenbosch, Republic of South Africa) Pablo Estévez (University of Chile) Bogdan Gabrys (Bournemouth University, UK) Fernando Gomide (University of Campinas, Brazil) Aboul Ella Hassanien (Cairo University, Egypt) Joachim Hertzberg (Osnabrück University, Germany) Tadeusz Kaczorek (Białystok University of Technology, Poland) Nikola Kasabov (Auckland University of Technology, New Zealand) Marian P. Kaźmierkowski (Warsaw University of Technology, Poland) Laszlo T. Kóczy (Szechenyi Istvan University, Gyor and Budapest University of Technology and Economics, Hungary) Józef Korbicz (University of Zielona Góra, Poland) Eckart Kramer (Fachhochschule Eberswalde, Germany) Rudolf Kruse (Otto-von-Guericke-Universität, Germany) Ching-Teng Lin (National Chiao-Tung University, Taiwan) Piotr Kulczycki (AGH University of Science and Technology, Poland) Andrew Kusiak (University of Iowa, USA) Mark Last (Ben-Gurion University, Israel) Anthony Maciejewski (Colorado State University, USA)
Krzysztof Malinowski (Warsaw University of Technology, Poland) Andrzej Masłowski (Warsaw University of Technology, Poland) Patricia Melin (Tijuana Institute of Technology, Mexico) Fazel Naghdy (University of Wollongong, Australia) Zbigniew Nahorski (Polish Academy of Sciences, Poland) Nadia Nedjah (State University of Rio de Janeiro, Brazil) Dmitry A. Novikov (Institute of Control Sciences, Russian Academy of Sciences, Russia) Duc Truong Pham (Birmingham University, UK) Lech Polkowski (University of Warmia and Mazury, Poland) Alain Pruski (University of Metz, France) Rita Ribeiro (UNINOVA, Instituto de Desenvolvimento de Novas Tecnologias, Portugal) Imre Rudas (Óbuda University, Hungary) Leszek Rutkowski (Czestochowa University of Technology, Poland) Alessandro Saffiotti (Örebro University, Sweden) Klaus Schilling (Julius-Maximilians-University Wuerzburg, Germany) Vassil Sgurev (Bulgarian Academy of Sciences, Department of Intelligent Systems, Bulgaria) Helena Szczerbicka (Leibniz Universität, Germany) Ryszard Tadeusiewicz (AGH University of Science and Technology, Poland) Stanisław Tarasiewicz (University of Laval, Canada) Piotr Tatjewski (Warsaw University of Technology, Poland) Rene Wamkeue (University of Quebec, Canada) Sławomir Wierzchon (Polish Academy of Sciences, Poland) Janusz Zalewski (Florida Gulf Coast University, USA) Teresa Zielińska (Warsaw University of Technology, Poland)
Publisher:
Copyright © by Łukasiewicz Research Network - Industrial Research Institute for Automation and Measurements PIAP All rights reserved
i
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20, N˚3, 2026
Contents 1
45
Some Extensions of the Controllability and Observability Tests to Linear Systems Tadeusz Kaczorek DOI: 10.14313/jamris‐2026‐032
A Hybrid LSTM‐DNN Model with Fuzzy Inference for Adaptive Dispatcher Control of Industrial Information‐Control Systems Barnokhon Temerbekova, Gulnora Bekimbetova, Ulugbek Mamanazarov, Bakhodir Bekimbetov DOI: 10.14313/jamris‐2026‐037
7
Validation of Eddy-Current Energy Losses in the Gyrator-Capacitor Model of Magnetic Hysteresis in a Nanocrystalline Soft Magnetic Core Roman Szewczyk, Piotr Gazda, Michał Nowicki, Paweł Nowak, Na Peng, Adam Bieńkowski, Tomasz Charubin DOI: 10.14313/jamris-2026-033
Micro‐ROS System Integration onto the Turtlebot3 Mobile Robot Platform Bartłomiej Stadnik, Artur Wymysłowski DOI: 10.14313/jamris‐2026‐038
15
63
Modified Hybrid Gray Wolf‐Cuckoo Search Algorithm for Optimum Tuning of the PI Controller for Controlling the Transients of the Doubly‐Fed Induction Generator Ashutosh Kashiv, H. K. Verma, Nagendra Singh DOI: 10.14313/jamris‐2026‐034
Research on Fuzzy-PID Control System for Mecanum Robot Trong-Tai Nguyen, Quang-Tho Le, Quang-Phuoc Pham, Tran-Long Le DOI: 10.14313/jamris-2026-039
26
86
Optimal Pid Level Control of a Nonlinear Spherical Tank Using a Hybrid WOA–GA Optimization Approach Abhay Pakhare, Sharad Jadhav, Mukesh Patil DOI: 10.14313/jamris‐2026‐035
Vision‐Based Gesture‐Controlled Mobile Robot for Human–Robot Interaction Using Ros 2 Kaveh Hooshmandi, Mahdi Molavi DOI: 10.14313/jamris‐2026‐040
Comparative Study of LQR and Neural Network Controllers for Quadcopter Roll Angle Stabilization Hiba Arif, Marouane Kadi, Aymane Bidah, Omar Zakary DOI: 10.14313/jamris‐2026‐036
Efficient Coverage Path Planning Via Gradient‐ Based Rectangular Segmentation Hubert Baraniak, Konrad Cop, Morteza Haghbeigi DOI: 10.14313/jamris‐2026‐041
36
ii
55
94
105
155
An Efficient Swarm Control Algorithm for Covering an Area with Communication Network Dariusz Miedziński, Mariusz Jacewicz, Kacper Kaczmarek, Sebastian Topczewski, Robert Głębocki, Antoni Kopyt DOI: 10.14313/jamris‐2026‐042
Revolutionizing Big Data Assessment for Human Activity Recognition with Metapath Context and Bi‐Directional Cascade Networks Praveen S. Banasode, Sunita Padmannavar DOI: 10.14313/jamris‐2026‐047
118
170
Fast Aerial Manipulation of a Varying Payload Tariel Simonyan, Oleg Gasparyan DOI: 10.14313/jamris‐2026‐043
126
Automated Detection and Severity Grading of Knee Osteoarthritis from X-Ray Images Using Machine Learning with Detailed Literature Review R.Gokulapriya, M.Balamurugan, Rakoth Kandan Sam‐ bandam, Divya Vetriveeran DOI: 10.14313/jamris-2026-048
Design and Development of Vacuum-Actuated Soft Gripper for Pick and Place Different Irregular Shaped Objects S.Senthil Raja, R.Gangadevi, M. Dharmaraj DOI: 10.14313/jamris-2026-044
Swinnetplus: A Novel Deep Learning Approach Using Swin Transformer For Enhancing Brain Tumor Segmentation Vikash Verma, Pritaj Yadav DOI: 10.14313/jamris-2026-049
Energy-Efficient Control of Collaborative Robots through Intelligent Optimization Algorithms Nandkishor Marotrao Sawai, Satpalsing Devising Ra‐ jput, Minal Vilas Gade, Dipak D. Bage, Aniruddha S. Rumale DOI: 10.14313/jamris-2026-045
Application of Both Static and Dynamic Analysis Methodologies for Malware Detection Andrzej Mycek, Mirosław Roszkowski DOI: 10.14313/jamris‐2026‐050
132
183
195
143
A General Form of Cautious Approximate Reasoning for Symbolic and Complex Data Saoussen Bel Hadj Kacem DOI: 10.14313/jamris‐2026‐046
iii
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
SOME EXTENSIONS OF THE CONTROLLABILITY AND OBSERVABILITY TESTS TO LINEAR SYSTEMS Submitted: 8th June 2024; accepted: 12th August 2025
Tadeusz Kaczorek DOI: 10.14313/jamris-2026-032 Abstract: The new controllability tests of the pairs (f(A), MB), where f(A) is a function well-defined on the spectrum of the matrix A and M is a nonsingular matrix defined by (3.7). They are proposed and illustrated by numerical examples. In the dual versions the tests can be used to test the observability of the pairs (f(A), CM). Keywords: Controllability, observability, linear, system, test, example
1. Introduction The notions of controllability and observability introduced by Kalman [9] are the basic concepts of the modern control theory. Modified tests for the controllability and the observability have been proposed in [2]. The tests have been extended to other classes of linear continuous-time and discrete-time systems [3–5, 7, 8, 10–12, 14]. Global stability of discrete-time nonlinear systems with descriptor and fractional linear parts and scalar feedbacks has been analyzed in [6]. The stabilization of positive descriptor fractional discrete-time linear systems with two different fractional order by a decentralized controller has been investigated in [13]. In this paper the tests will be extended to the pair (f (A), MB), where f (A) is a function welldefined on the spectrum of the matrix A and M is a nonsingular matrix defined by (3.7). The paper is organized as follows: Section 2 offers basic definitions and theorems concerning controllability, stability and Frobenius canonical forms of the continuoustime linear systems that have been recalled. The controllability of linear systems with different single inputs have been considered in Section 3. The controllability of linear systems with functions of the state matrices has been investigated in Section 4. The general case has been analyzed in Section 5. Concluding remarks are given in Section 6. The following notation will be used: ℜ - the set of real numbers, ℜ𝑛×𝑚 - the set of 𝑛 × 𝑚 real matrices, ℜ𝑛×𝑚 - the set of 𝑛 ×𝑚 real matrices with nonnegative + entries and ℜ𝑛+ = ℜ𝑛×1 + , 𝐼𝑛 - the 𝑛 × 𝑛 identity matrix.
2. Preliminaries Consider the continuous-time linear system 𝑥̇ = 𝐴𝑥 + 𝐵𝑢,
(2.1a)
𝑦 = 𝐶𝑥,
(2.1b)
where 𝑥 = 𝑥(𝑡) ∈ ℜ𝑛 , 𝑢 = 𝑢(𝑡) ∈ ℜ𝑚 , 𝑦 = 𝑦(𝑡) ∈ ℜ𝑝 are the state, input and output vectors and 𝐴 ∈ ℜ𝑛×𝑛 , 𝐵 ∈ ℜ𝑛×𝑚 , 𝐶 ∈ ℜ𝑝×𝑛 . Definition 2.1. [2–4, 6, 10, 11] The continuous-time linear system (2.1) is called controllable if for given initial state 𝑥(0) ∈ ℜ𝑛 and a given final state 𝑥𝑓 ∈ ℜ𝑛 there exists an input 𝑢(𝑡) ∈ ℜ𝑚 for 𝑡 ∈ [0, 𝑡𝑓 ] which steers the system from 𝑥(0) to 𝑥𝑓 = 𝑥(𝑡𝑓 ). Theorem 2.1. [2–4, 6, 10, 11] The linear system (2.1) is controllable if and only if one of the following conditions is satisfied: 1. (Kalman condition) rank[𝐵
𝐴𝐵
...
𝐴𝑛−1 𝐵] = 𝑛
2. (Hautus condition) rank[𝐼𝑛 𝑠 − 𝐴
𝐵]
= 𝑛 for s ∈ C (the field of complex numbers). (2.3) Definition 2.2. [2–4, 6, 10, 11] The linear system (2.1) is called observable if knowing its input 𝑢(𝑡) ∈ ℜ𝑚 and its output 𝑦(𝑡) ∈ ℜ𝑝 for 𝑡 ∈ [0 𝑡𝑓 ] it is possible find its unique initial condition 𝑥(0) ∈ ℜ𝑛 . Theorem 2.2. [2–4, 6, 10, 11] The linear system (2.1) is observable if and only if one of the following conditions is satisfied: 1. (Kalman condition) 𝐶 ⎡ ⎤ 𝐶𝐴 ⎥=𝑛 rank ⎢ ⋮ ⎢ ⎥ ⎣ 𝐶𝐴𝑛−1 ⎦
(2.4)
2. (Hautus condition) rank �
𝐼𝑛 𝑠 − 𝐴 � 𝐶
= 𝑛 for s ∈ C (the field of complex numbers). (2.5) Definition 2.3. [2-4, 6, 10, 11] The system (2.1) for 𝑢(𝑡) = 0 is called asymptotically stable if lim 𝑥(𝑡) = 0 for any x(0) ∈ ℜn+ .
𝑡→∞
Open Access. © 2026 Tadeusz Kaczorek, published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
(2.2)
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
(2.6)
1
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
Theorem 2.3. [2–4, 6, 10, 11] The linear system (2.1) for 𝑢(𝑡) = 0 is asymptotically stable if and only if the eigenvalues of the matrix A satisfies the condition Re𝜆i < 0 for i = 1, …, n.
(2.7)
Definition 2.4. The matrix A has the Frobenius canonical form if it has one of the following forms: 0 ⎡ ⎢ 0 𝐴1 = ⎢ ... ⎢ 0 ⎣ −𝑎0
1 0 ... 0 −𝑎1
0 0 ⎡ ⎢ 1 0 𝐴2 = 𝐴1𝑇 = ⎢ 0 1 ⎢ ... ... ⎣ 0 0 ⎡ ⎢ 𝐴3 = ⎢ ⎢ ⎣
−𝑎𝑛−1 1 0 ... 0
0 1 ... 0 −𝑎2
1 0 ... 0 0
... ... ... ... ...
−𝑎1 0 0 ... 1
0 1 ... 0 0
... ... ... ... ...
0 0 ⎡ ⎢ 1 0 𝐴2 = 𝑃2−1 𝐴𝑃2 = ⎢ 0 1 ⎢ ... ... ⎣ 0 0 1 ⎡ ⎤ 0 ⎥ 𝐵2 = 𝑃2−1 𝐵 = ⎢ , ⎢ ⋮ ⎥ ⎣ 0 ⎦
Controllability of a class of multi-inputs linear systems
Theorem 3.1. The above pair (A,B) is controllable if the matrix A is nonsingular (det 𝐴 ≠ 0).
−𝑎0 ⎤ 0 ⎥ 0 ⎥, .. ⎥ 0 ⎦ 0 ⎤ 0 ⎥ ... ⎥ . 1 ⎥ 0 ⎦
Proof. Let the matrix A has the form
(2.8)
... 0 ⎤ ... 0 ⎥ ... ... ⎥, 0 1 ⎥ ... −𝑎𝑛−1 ⎦
0 0 ⎡ ⎢ 1 0 𝐴 =⎢ 0 1 ⎢ ... ... ⎣ 0 0
(2.9b)
1) Multiplication of any i-th row (column) by the number 𝑎. This operation will be denoted by
... 0 −𝑎0 ⎤ ... 0 −𝑎1 ⎥ ... 0 −𝑎2 ⎥ ... ... ... ⎥ ... 1 −𝑎𝑛−1 ⎦
(3.1a)
and the matrix B one of the following forms
1 ⎤ ⎡ 0 ⎥ , 𝐵1 = ⎢ ⎢ ⋮ ⎥ ⎣ 0 ⎦
0 ⎤ ⎡ 1 ⎥ , 𝐵2 = ⎢ ⎢ ⋮ ⎥ ⎣ 0 ⎦
0 ⎤ ⎡ ⋮ ⎥ . ..., 𝐵𝑛 = ⎢ ⎢ 0 ⎥ ⎣ 1 ⎦
(3.1b)
The pair (𝐴, 𝐵1 ) is controllable since by Hautus condition (2.3) rank�𝐼𝑛 𝑠 − 𝐴
(2.9a)
In a similar way we may obtain the remaining canonical forms of 𝐴3 and 𝐴4 given by (2.8). The following elementary operations on real matrices will be used [3, 4, 7]:
2
2) Addition to any i-th row (column) of the j-th row (column) multiplied by any number 𝑏. This operation will be denoted by 𝐿[𝑖 + 𝑗 × 𝑏] for row operation and by 𝑅[𝑖+𝑗×𝑏] for column operation.
Consider the matrix A in the canonical form given by (2.8) and the matrix 𝐵 ∈ ℜ𝑛×1 with only one nonzero entry equal 1.
... 0 −𝑎0 ⎤ ... 0 −𝑎1 ⎥ ... 0 −𝑎2 ⎥ , ... ... ... ⎥ ... 1 −𝑎𝑛−1 ⎦
det 𝑃2 ≠ 0
𝐿[𝑖 × 𝑎] for row operation and by 𝑅[𝑖 × 𝑎] for column operation.
3.
The inverse matrices of the matrices (2.8) have also the Frobenius canonical forms [6]. It is well-known that [4, 9, 10, 12] the controllable pair (𝐴 ∈ ℜ𝑛×𝑛 , 𝐵 ∈ ℜ𝑛×1 ) by the similarity transformation 𝑥̄ = 𝑃𝑥 can be transform to its canonical form 0 1 0 ⎡ 0 0 1 ⎢ ... ... 𝐴1 = 𝑃1−1 𝐴𝑃1 = ⎢ ... 0 0 ⎢ 0 ⎣ −𝑎0 −𝑎1 −𝑎2 0 ⎡ ⎤ ⋮ ⎥ −1 𝐵1 = 𝑃1 𝐵 = ⎢ , det 𝑃1 ≠ 0 ⎢ 0 ⎥ ⎣ 1 ⎦
2026
3) for interchange of rows and 𝑅[𝑖, 𝑗] for interchange of columns.
... 0 −𝑎0 ⎤ ... 0 −𝑎1 ⎥ ... 0 −𝑎2 ⎥ , ... ... ... ⎥ ... 1 −𝑎𝑛−1 ⎦
−𝑎𝑛−2 0 1 ... 0
−𝑎𝑛−1 ⎡ ⎢ −𝑎𝑛−2 𝑇 ... 𝐴4 = 𝐴3 = ⎢ −𝑎 1 ⎢ ⎣ −𝑎0
... 0 ⎤ ... 0 ⎥ ... ... ⎥, 0 1 ⎥ ... −𝑎𝑛−1 ⎦
N∘ 3
𝐵1 �
𝑠 0 ... 0 ⎡ −1 𝑠 ... 0 ⎢ = rank⎢ 0 −1 ... 0 ... ... ... ⎢ ... 0 ... −1 ⎣0
𝑎0 𝑎1 𝑎2 ... 𝑠 + 𝑎𝑛−1
= 𝑛 for ∀s ∈ C.
1 ⎤ 0 ⎥ 0 ⎥ … ⎥ 0 ⎦ (3.2a)
The pair (𝐴, 𝐵2 ) is also controllable since rank�𝐼𝑛 𝑠 − 𝐴
𝐵2 �
𝑠 0 ... 0 ⎡ −1 𝑠 ... 0 ⎢ = rank⎢ 0 −1 ... 0 ... ... ... ⎢ ... 0 ... −1 ⎣0 = 𝑛 for ∀s ∈ C and 𝑎0 ≠ 0.
𝑎0 𝑎1 𝑎2 ... 𝑠 + 𝑎𝑛−1
0 ⎤ 1 ⎥ 0 ⎥ … ⎥ 0 ⎦ (3.2b)
Journal of Automation, Mobile Robotics and Intelligent Systems
Continuing this procedure after n steps we obtain rank�𝐼𝑛 𝑠 − 𝐴
𝐵𝑛 �
𝑠 0 ⎡ −1 𝑠 ⎢ = rank⎢ ... ... ⎢0 0 ⎣0 0
... ... ... ... ...
0 0 ... 𝑠 −1
𝑎0 𝑎1 ... 𝑎𝑛−2 𝑠 + 𝑎𝑛−1
0 ⎤ 0 ⎥ ... ⎥ 0 ⎥ 1 ⎦ (3.2c)
= 𝑛 for ∀s ∈ C
Therefore, the pairs (𝐴, 𝐵𝑖 ), i = 1,2,…,n is controllable if 𝑎0 = det 𝐴 ≠ 0. Example 3.1. Consider the pair (A,B), for 0 1 0 0 𝐴1 = � −1 −2
0 1 �, −3
0 𝐴 2 = �1 0
0 0 1
−1 −2� −3
(3.3a)
VOLUME 20,
N∘ 3
This confirm the Theorem 3.1. Given the controllable pair 0 ⎡ ⎢ 0 𝐴̄ = 𝑃−1 𝐴𝑃 = ⎢ ... ⎢ 0 ⎣−𝑎0 0 ⎡ ⎤ ⋮ ⎥ 𝐵̄ = 𝑃−1 𝐵 = ⎢ , ⎢ 0 ⎥ ⎣ 1 ⎦
1 0 ... 0 −𝑎1
0 1 ... 0 −𝑎2
det 𝐴 ≠ 0,
... 0 ⎤ ... 0 ⎥ ... ... ⎥ , 0 1 ⎥ ... −𝑎𝑛−1 ⎦ det 𝑃 ≠ 0 (3.6)
Find a nonsingular matrix 𝑀 ∈ ℜ𝑛×𝑛 such that the pair ̄ is also controllable. (𝐴,̄ 𝑀𝐵) Let ̄ −1 𝑀 = 𝑃𝑀𝑃 (3.7) then
and
𝑃−1 𝑀𝑃 = 𝑃−1 𝑀𝑃𝑃−1 𝐵 = 𝑀̄ 𝐵̄ = 𝐵̂
1 𝐵1 = � 0 � , 0
0 𝐵2 = � 1 � , 0
0 𝐵3 = � 0 � . 1
𝐴1 𝐵1
1 = rank�0 0 rank�𝐵2
𝐴1 𝐵2
0 = rank�1 0 rank�𝐵3
𝐴12 𝐵1 �
0 0 0 −1� = 3, −1 3
𝑀̄ = 𝑃−1 𝑀𝑃,
rank�𝐵2
(3.9)
Let 𝑀̄ ∈ ℜ𝑛×𝑛 be a monomial matrix (in each row and in each column only one element is nonzero and the remaining elements are zero). In this case the matrix 𝑀̄ 𝐵̄ ∈ ℜ𝑛×1 has only one nonzero element and the remaining elements are zero and by Theorem 3.1 the ̂ is controllable. pair (𝐴,̄ 𝐵) Therefore, the following theorem has been proved: Theorem 3.2. If the pair (A,B) is controllable and ̂ is also controllable. det 𝐴 ≠ 0 then the pair (𝐴,̄ 𝐵)
(3.4b)
Step 1. For the controllable pair (A,B), det 𝐴 ≠ 0, find ̄ defined by (3.6). the matrix P and the pair (𝐴,̄ 𝐵) Step 2. Choose the monomial matrix 𝑀.̄ Step 3. Compute the desired matrix (3.7).
𝐴12 𝐵3 �
0 1 1 −3� = 3 −3 7
Procedure 3.1.
(3.4c)
Example 3.2. For the given controllable pair −0.5 𝐴 = �−0.5 1
and rank�𝐵1
𝐵̄ = 𝑀−1 𝐵
To find the matrix M the following procedure can be used.
1 0 0 −2� = 3, −2 5
0 = rank�0 1
(3.4a)
𝐴12 𝐵2 �
𝐴1 𝐵3
(3.8)
where (3.3b)
Using the Kalman condition (2.2) for (3.3) we obtain rank�𝐵1
2026
𝐴2 𝐵1
1 𝐴22 𝐵1 � = rank�0 0
0 1 0
0 0� = 3, 1 (3.5a)
𝐴2 𝐵2
0 𝐴22 𝐵2 � = rank�1 0
0 0 1
−1 −2� = 3, −3 (3.5b)
−1 −0.75 1 0.75 � , −6 −2.5
1 𝐵 =� 0 � 2
(3.10)
find the matrix M such that the pair (A,MB) is also controllable. Using Procedure 3.1 we obtain Step 1. In this case the matrix P has the form 2 𝑃 = �1 0
0 1 0
1 0� 2
(3.11)
and rank�𝐵3
𝐴2 𝐵3
0 𝐴22 𝐵3 � = rank�0 1
−1 −2 −3
3 5� = 3. 7 (3.5c)
0 1 0 𝐴̄ = 𝑃−1 𝐴𝑃 = � 0 0 1 � , −2−3−2
0 𝐵̄ = 𝑃−1 𝐵 = � 0 � . 1 (3.12) 3
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
Step 2. We choose the monomial matrix 𝑀̄ in the form 0 𝑀̄ = �1 0
1 0 0
0 0� . 1
Ψ(𝜆) = (𝜆 − 𝜆1 )(𝜆 − 𝜆2 )...(𝜆 − 𝜆𝑛 )
(3.13)
𝑛
𝑓(𝐴) = � 𝑍𝑘 𝑓(𝜆𝑘 ) (3.14) where
Remark 3.1. The above considerations for single-input systems can be easily extended to multi-input case in the following way. Let 𝐵 ∈ ℜ𝑛×𝑚 , 𝑚 > 1 and 𝑘 ∈ ℜ𝑚 .
𝑛
𝑍𝑘 = � 𝑖=1 𝑖≠𝑘
...
𝐴
� 𝑍𝑖1 = 𝐼𝑛
(3.15)
𝐵𝑘� ≠ 0.
Ψ(𝜆) = (𝜆 − 𝜆1 )𝑚1 (𝜆 − 𝜆2 )𝑚2 ...(𝜆 − 𝜆𝑟 )𝑚𝑟
𝑍𝑖𝑗 𝑍𝑘𝑗 = 0 for 𝑖 ≠ 𝑘
(4.6b)
𝑍𝑖1 𝑍𝑘𝑗 = 𝑍𝑖𝑗 for 𝑖 = 1, ..., 𝑟; 𝑗 = 1, ..., 𝑚
(4.6c)
𝑘 𝑍𝑖1 = 𝑍𝑖1 for 𝑘 = 1, 2, ...; 𝑖 = 1, ..., 𝑟
(4.6d)
1 (𝐴 − 𝐼𝑛 𝜆𝑖 )𝑗−1 𝑍𝑖1 for (𝑗 − 1)! 𝑖 = 1, ...𝑟; 𝑗 = 1, ..., 𝑚
...,
(4.6e)
In particular case the matrices (4.5b) satisfy the equalities 𝑛
(4.7a)
� 𝑍𝑘 = 𝐼𝑛
(4.1)
𝑘=1
𝑍𝑖 𝑍𝑗 = 0 for 𝑖 ≠ 𝑗, 𝑖, 𝑗 = 1, ..., 𝑛
(4.7b)
𝑍𝑖𝑘 = 𝑍𝑖 for, 𝑘 = 1, 2, ... 𝑖 = 1, ..., 𝑛
(4.7c)
To simplify the notation, we shall consider the details for the case when the matrix A has only distinct eigenvalues 𝜆𝑖 ≠ 𝜆𝑗 , 𝑖 ≠ 𝑗. Theorem 4.1. The pair (f (A),B) is controllable if, and only if, the pair (A,B) is controllable and the function 𝑓(𝜆) is well defined on the spectrum of the matrix A.
𝑘
𝑑 𝑚𝑘 −1 𝑓(𝜆) 𝑓 (𝑚𝑘 −1) (𝜆𝑘 ) = � , 𝑑𝜆𝑚𝑘 −1 𝜆=𝜆
(4.6a)
𝑖=1
𝑍𝑖𝑗 =
where 𝜆1 , ..., 𝜆𝑟 are the eigenvalues of the matrix A 𝑟 and ∑𝑖=1 𝑚𝑖 = 𝑚 ≤ 𝑛. It is assumed that the function 𝑓(𝜆) is well defined on the spectrum 𝜎𝐴 = {𝜆1 , ..., 𝜆𝑟 } of the matrix A, i.e. 𝑑𝑓(𝜆) � , 𝑑𝜆 𝜆=𝜆
(4.5b)
(3.16)
Consider the controllable linear system (2.1a) and the minimal characteristic polynomial of the matrix A in the form
𝑓 (1) (𝜆𝑘 ) =
𝑘 = 1, … , 𝑛
𝑟
4. Controllability of systems with functions of state matrices
𝑓(𝜆),
𝐴 − 𝜆𝑖 𝐼𝑛 , 𝜆𝑘 − 𝜆𝑖
It is well-known [1] that the matrices 𝑍𝑖𝑗 satisfy the equalities
The vector k is chosen so that the equivalent single input system with input matrix b is controllable and satisfies the condition 𝐴𝐵𝑘
(4.5a)
𝑘=1
It is easy to check that the pair (A,MB) for A and B given by (3.10) and M by (3.14) is also controllable.
det�𝐵𝑘
(4.4)
then the formula (4.3a) has the form
0 −1 −0.25 ̄ −1 = �−0.25 0 0 � 𝑀 = 𝑃𝑀𝑃 0.5 −3 −1.25
𝑛−1
2026
In particular case when the eigenvalues 𝜆1 , ..., 𝜆𝑛 of the matrix A are distinct (𝜆𝑖 ≠ 𝜆𝑗 , 𝑖 ≠ 𝑗)
Step 3. The desired matrix M has the form
𝐵𝑘 = 𝑏 ∈ ℜ𝑛 ,
N∘ 3
𝑘 = 1, ..., 𝑟
𝑘
(4.2) are finite [1]. In this case the matrix f (A) is well-defined and it is given by the Lagrange-Sylvester formula [1] 𝑟
𝑓(𝐴) = � 𝑍𝑖1 𝑓(𝜆𝑖 )+𝑍𝑖2 𝑓 (1) (𝜆𝑖 )+...+𝑍𝑖𝑚𝑖 𝑓 (𝑚𝑖 −1) (𝜆𝑖 ) 𝑖=1
(4.3a)
where
Proof. Using (2.2) and (4.5) we obtain rank�𝐵
𝑚𝑖 −1
𝑘=𝑗−1
Ψ𝑖 (𝐴)(𝐴 − 𝜆𝑖 𝐼𝑛 )𝑘 𝑑𝑘−𝑗+1 1 � �� 𝑘−𝑗+1 (𝑘 − 𝑗 + 1)!(𝑗 − 1)! 𝑑𝜆 Ψ𝑖 (𝜆) 𝜆=𝜆
...
𝑛
�∑𝑖=1 𝑍𝑖 𝑓(𝜆𝑖 )�
𝑛−1
𝐵� (4.8)
Taking into account that 𝑍𝑖 𝑍𝑗 = 0 for 𝑖 ≠ 𝑗 and 𝑍𝑖𝑘 = 𝑍𝑖 for 𝑘 = 1, 2, ... 𝑖 = 1, ..., 𝑛 (relations (4.7)) we obtain 𝑓(𝐴)𝐵
...
∑𝑛𝑖=1 𝑍𝑖 𝑓(𝜆𝑖 )𝐵
�𝐵 𝑖
(𝑓(𝐴))𝑛−1 𝐵� = rank
...
∑𝑛𝑖=1 𝑍𝑖 𝑓(𝜆𝑖 )𝐵
�𝐵
rank�𝐵
𝑍𝑖𝑗 = �
𝑓(𝐴)𝐵
(𝑓(𝐴))𝑛−1 𝐵� = rank ...
∑𝑛𝑖=1 𝑍𝑖 (𝑓(𝜆𝑖 ))𝑛−1 𝐵� (4.9)
since
(4.3b) and
𝑛
Ψ(𝜆) Ψ𝑖 (𝜆) = , (𝜆 − 𝜆𝑖 )𝑚𝑖 4
𝑖 = 1, ..., 𝑟.
(4.3c)
𝑘
𝑛
�� 𝑎𝑖 𝑍𝑖 � = � 𝑎𝑘𝑖 𝑍𝑖 for 𝑘 = 1, 2, ... 𝑖=1
𝑖=1
(4.10)
Journal of Automation, Mobile Robotics and Intelligent Systems
After some algebraic manipulations we obtain
VOLUME 20,
N∘ 3
2026
Step 3. Using (3.14) and (4.15a) we obtain
𝑓(𝐴)𝐵
...
(𝑓(𝐴))𝑛−1 𝐵�
rank�𝐵
= rank�𝐵
𝐴𝐵
...
𝐴𝑛−1 𝐵� 𝐻
= rank�B
(𝑎0 𝐼3 + 𝑎1 𝐴 + 𝑎2 𝐴2 )𝐵
(𝑎0 𝐼3 + 𝑎1 𝐴 + 𝑎2 𝐴2 )2 𝐵�
= rank�𝐵
𝐴𝐵
...
𝐴𝑛−1 𝐵�
= rank�B
𝑎0 𝐵 + 𝑎1 𝐴𝐵 + 𝑎2 𝐴2 𝐵
𝑎02 𝐵 + 𝑎12 𝐴𝐵 + 𝑎22 𝐴2 𝐵�
= rank�B
AB
rank�𝐵
(4.11a)
where the matrix 1 𝑎0 ⎡ 0 𝑎1 𝐻 =⎢ ... ... ⎢ 0 𝑎 ⎣ 𝑛−1
𝑖 ≠ 𝑗,
(𝑒 𝐴 )2 𝐵�
A2 𝐵� 𝐻 = rank�B
AB
A2 𝐵� (4.16a)
where the matrix
... 𝑎0𝑛−1 ⎤ ... 𝑎1𝑛−1 ⎥ ... ... ⎥ 𝑛−1 ... 𝑎𝑛−1 ⎦
is nonsingular for 𝑠𝑖 ≠ 𝑠𝑗 , This completes the proof.
𝑒 𝐴𝐵
𝑖, 𝑗 = 1, ..., 𝑛.
𝑎02 𝑎12 � 𝑎22
1 𝑎0 𝐻 = �0 𝑎1 0 𝑎2
(4.11b)
(4.16b)
is nonsingular. Therefore, the pair (𝑒𝐴 , 𝐵) is controllable.
Remark 4.1. In general case the proof of Theorem 4.1 is similar. In this case the relations (4.6) should be used instead of (4.7). The controllability of the pair (f (A), B) can be checked by the use of the following procedure.
5.
Procedure 4.1.
Theorem 5.1.If the function 𝑓(𝜆) is well-defined on the spectrum of the matrix A, det 𝐴 ≠ 0 and the pair (A,B) is controllable (A and B are defined by (3.6) and (3.8)) then the pair (f (A),MB) is controllable for matrix M defined by (3.7). Proof follows immediately from the proofs of Theorem 3.1 and 4.1. The controllability of the pair (f (A),MB) can be checked by the use of Procedures 3.1 and 4.1. Example 5.1. Check the controllability of the pair (𝐴−1 , 𝑀𝐵) for the matrices A, B given by (3.10) and the matrix M by (3.14). The pair (3.10) is controllable and the matrix M has been obtained from the matrix (3.13) by similarity transformation (3.14). The eigenvalues of the matrix A are 𝑠1 = −1, 𝑠2 = −2, 𝑠3 = −3 and the there exists the its inverse matrix
Step 1. Compute the eigenvalues 𝜆1 , ..., 𝜆𝑛 of the matrix A. Step 2. Check if the function 𝑓(𝜆) satisfies the conditions (4.2). Step 3. Check the controllability of the pair (f (A),B). Example 4.1. Consider the controllable pair 0 𝐴 =� 0 −6
1 0 0 1 �, −11 −6
0 𝐵 =� 0 � . 1
(4.12)
Using Procedure 4.1 we obtain the following. Step 1. The characteristic polynomial of the matrix A given by (4.12) has the form 𝑠 −1 0 −1 � = 𝑠 3 + 6𝑠 2 + 11𝑠 + 6 det[𝐼3 𝑠 − 𝐴] = �0 𝑠 6 11 𝑠 + 6 (4.13) and the eigenvalues of the matrix are 𝑠1 = −1, 𝑠2 = −2, 𝑠3 = −3. Step 2. In this case the conditions (4.2) are satisfied and using (4.5b) we obtain (𝐴 − 𝐼3 𝑠2 )(𝐴 − 𝐼3 𝑠3 ) 5 1 = 3𝐼3 + 𝐴 + 𝐴2 , (𝑠1 − 𝑠2 )(𝑠1 − 𝑠3 ) 2 2 (𝐴 − 𝐼3 𝑠1 )(𝐴 − 𝐼3 𝑠3 ) 𝑍2 = = −3𝐼3 − 4𝐴 − 𝐴2 , (𝑠2 − 𝑠1 )(𝑠2 − 𝑠3 ) (𝐴 − 𝐼3 𝑠1 )(𝐴 − 𝐼3 𝑠2 ) 3 1 𝑍3 = = 𝐼3 + 𝐴 + 𝐴2 (4.14) (𝑠3 − 𝑠2 )(𝑠3 − 𝑠1 ) 2 2
General case and the extension to the observability of the systems
In this section the general case will be considered using the results of sections 3 and 4.
−1
𝐴
0 =� 0 −6 =�
1 0 −11
−1.833 1 0
0 0� −6 −1 0 1
−1
−0.167 0 � 0
(5.1)
𝑍1 =
The matrices (4.14) satisfy the relations (4.7) and 𝑒𝐴 = 𝑎0 𝐼3 + 𝑎1 𝐴 + 𝑎2 𝐴2
𝑎2 =
1 −1 1 𝑒 − 𝑒 −2 + 𝑒 −3 2 2
0 −1 −0.25 1 −0.5 0 �� 0 � = � −0.25 � 𝑀𝐵 = �−0.25 0 −0.5 −3 −1.25 2 −2 (5.2) The pair (𝐴−1 , 𝑀𝐵) is controllable since
(4.15a) det�𝑀𝐵
where 𝑎0 = 3(𝑒 −1 − 𝑒 −2 ) + 𝑒 −3 ,
Taking into account (3.14) and the matrix B we obtain
𝑎1 =
5 −1 3 𝑒 − 4𝑒 −2 + 𝑒 −3 , 2 2 (4.15b)
𝐴−1 𝑀𝐵
(𝐴−1 )2 𝑀𝐵�
−0.5 1.5 −2.208 1.5 � = −2.93 = �−0.25 −0.5 −2 −0.25 −0.5
(5.3)
This simple example confirms the Theorem 5.1. 5
Journal of Automation, Mobile Robotics and Intelligent Systems
Using the well-known duality between the controllability and the observability of linear systems the above presented results for the controllability can be extended to the observability of the above classes of linear systems. For example, the Theorem 5.1 can be extended to the observability case as follows. Theorem 5.2. If the function 𝑓(𝜆) is well-defined on the spectrum of the matrix A, det 𝐴 ≠ 0 and the pair (A,C) is observable then the pair (f (A),CM) is observable for the nonsingular matrix M defined by (3.7).
6. Concluding remarks The controllability and observability tests have been extended to the pair (f (A), MB), where f (A) is a function well-defined on the spectrum of the matrix A and M is a nonsingular matrix defined by (3.7). The controllability tests of linear systems with different single inputs have been investigated (Theorem 3.1 and 3.2). The controllability tests of linear systems with functions of the state matrices has been established (Theorem 4.1). Procedures for checking the controllability tests have been given and illustrated by numerical examples. The tests can be easily extended to check the observability of the systems. Using the approach given in [5] the results of this paper can be extended to convex linear combination of the controllability (observability) pairs. Open problems are extensions of these results to discrete-time linear systems and to fractional linear systems.
AUTHOR Tadeusz Kaczorek∗ – Bialystok University of Technology, Poland, e-mail: kaczorek@ee.pw.edu.pl. ∗ Corresponding author
ACKNOWLEDGEMENTS The studies have been carried out in the framework of work No. WZ/WE-IA/5/2023 and financed from the funds for science by the Polish Ministry of Science and Higher Education.
References [1] F.R. Gantmacher, Theory of Matrices, Chelsea Publishing Company, 1977. [2] M.L.J. Hautus and M. Heymann, “Linear Feedback-an Algebraic Approach,” SIAM j. Contr. and Optim, vol. 16, no. 1, 1978, pp. 83–105.
6
VOLUME 20,
N∘ 3
2026
[3] T. Kaczorek, Linear Control Systems, vol. 1 and 2,Reaserch Studies Press LTD, J. Wiley, 1992. [4] T. Kaczorek and K. Borawski, Descriptor Systems of Integer and Fractional Orders, Studies in Systems, Decision and Control, Springer, 2021. [5] T. Kaczorek and J. Klamka. “Convex Linear Combination of the Controllability Pairs for Linear Systems,” Control and Cybernetics, vol. 50, no. 4., 2021. [6] T. Kaczorek and A. Ruszewski, “Global Stability of Discrete-Time Nonlinear Systems with Descriptor and Fractional Positive Linear Parts and Scalar Feedbacks,” Archives of Control Sciences, vol. 30, no. 4, 2020, pp. 667–681. [7] T. Kaczorek and K. Rogowski, Fractional Linear Systems and Electrical Circuits, Springer Cham Springer, 2014. [8] T. Kailath, Linear Systems, Prentice-Hall, 1980. [9] Kalman, “Mathematical Description of Linear Dynamical Systems,” SIAM Journal of Control, Series A, 1963, pp. 152-192. [10] J. Klamka, Controllability and Minimum Energy Control, Studies in Systems, Decision and Control, Springer Verlag, vol. 162, 2018. [11] J. Klamka, Controllability of Dynamical Systems, Kluwer Academin. Publ., 1991. [12] W. Mitkowski, “Outline of Control Theory,” Publishing House AGH, 2019. [13] L. Sajewski, “Stabilization of Positive Descriptor Fractional Discrete-Time Linear Systems with Two Different Fractional Orders by Decentralized Controller,” Bull. Pol. Acad. Sci. Techn., vol. 65, no. 5, 2017, pp. 709–714 [14] S. Zak, Systems and Control, Oxford University Press, 2003.
VOLUME 20, N° 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
Validation of Eddy-Current Energy Losses in the Gyrator-Capacitor Model of Magnetic Hysteresis in a Nanocrystalline Soft Magnetic Core Submitted: 25th August 2025; accepted: 20th November 2025
Roman Szewczyk, Piotr Gazda, Michał Nowicki, Paweł Nowak, Na Peng, Adam Bieńkowski, Tomasz Charubin DOI: 10.14313/jamris-2026-033 Abstract: This article presents the results of validating the gyrator-capacitor model of the frequency dependence characteristic of an inductive core made of Fe73.5Cu1Nb3Si15.5B7 nanocrystalline alloy. The model, focused on eddy current losses, was implemented using the LTspice software. Experimental results confirmed that a relatively simple description of eddy current losses enables effective modeling of the magnetic hysteresis loop for driving frequencies up to 36 kHz. A very good agreement between the modeling and experimental results was confirmed quantitatively by the R² coefficient exceeding 0.995. Results indicate that further research should consider more detailed functional models of the frequency dependence of the coercive field. Keywords: eddy current energy losses, gyrator-capacitor model, nanocrystalline materials
1. Introduction Inductive components play a crucial role in power conversion devices [1], telecommunications [2], and the automotive industry [3]. However, despite the fact that the total yearly turnover of the inductive components market exceeds 23 billion US$ [4], its operational parameters selection process is often based on a set of rules of thumb, common practices, and rough estimations. The primary barrier to the efficient optimization process of inductive components is the limited applicability of the components commonly used physical models. These physical models primarily rely on the linearization of characteristics or are restricted to a narrow range of operating parameters. The gyrator-capacitor model presents an opportunity to overcome this limited applicability problem, enabling the efficient modeling of inductive component characteristics in the SPICE (Simulation Program with Integrated Circuit Emphasis) environment [5]. However, from the perspective of the power conversion device and automotive industries, the most important thing is accurately modeling the power loss characteristics of the inductive component. From the physical point of view, inductive component power losses consist of hysteresis losses, eddy current losses, and excess (also known as anomalous) losses [6, 7]. In the case of power conversion devices
operating at driving frequencies up to 50 kHz, eddy current losses play a key role in efficient modeling. While there is a large variety of eddy current loss models [8–12], the commonly used gyrator-capacitor model utilizes the approach presented by Bertotti et al. [13]. This approach simplifies the model through a constant-value resistor that describes eddy current losses. However, the proposed approach of eddy current losses in the gyrator-capacitor model of the inductive component for SPICE modeling has not yet been directly validated. This article fills this gap. The eddy current losses of a ring-shaped core made of nanocrystalline alloy were measured directly using an advanced hysteresis graph in the frequency range direct current to 36 kHz. In this frequency range, eddy current losses are the primary power losses in the inductive component. Next, the gyrator-capacitor model was implemented in the LTspice environment [14] offered by Linear Technology. Finally, the modeling results were compared with experimental measurements, enabling validation of the eddy current-based loss model for inductive components.
2. The material and methods of experimental investigation
A commercially available Vacuumschmeltze W915 core was used for the experimental investigation. It is a core made of nanocrystalline material, Vitroperm 500F, with the following chemical composition: Fe73.5Cu1Nb3Si15.5B7. The core was a ribbon-wound ring core in the form of a toroid, with magnetic material dimensions of 6.5 mm x 9.8 mm x 4.5 mm (inner diameter, outer diameter, height). The sample was deliberately selected to be as small as possible, so that the induced voltage in the flux density (B) induction measurements was as low as possible and within the measuring range of the hysteresis graph. Hysteresis loop measurements were performed using an AMH-200K-S device (Laboratorio Elettrofisico, Milan, Italy). The samples were magnetized with sinusoidal excitation, with an amplitude of 350 A/m, and the magnetization frequency (f) ranged from 1 to 36 kHz. Due to its high permeability, the sample was magnetized using only two turns of the coil, wound symmetrically on the sample. Figure 1 presents the results of the hysteresis loop measurements as a function of frequency.
Open Access. © 2026 Roman Szewczyk et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP. This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 License.
7
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Se is the core’s cross-section, le is the length of the magnetic circuit of the core, and μ0 is the magnetic constant. In the LTspice software, the nonlinear capacitor is modeled considering its charge Q [18] in the following form: Q C1 · x V r · o ·
Figure 1. Results of experimental measurements of the B(H) hysteresis loop as a function of magnetizing field frequency (f) for the Fe73.5Cu1Nb3Si15.5B7 (Vitroperm 500F) nanocrystalline alloy. Consecutive values of frequency follow the direction indicated by the arrow
Figure 2. The gyrator-capacitor circuit for modeling an inductor with the magnetic core exhibits nonlinear magnetization characteristics
3. Method of modeling The proposed model of a nonlinear magnetic core was implemented in LTspice [14]. The gyrator circuit [15] with a gyrator resistance of 1/N [16] was enhanced by incorporating two ultra-high-value resistors to improve the circuit’s numerical stability. From a technical point of view, in this case, N should be considered the number of winding turns on the modeled magnetic core. In the gyrator-capacitor method of modeling [17], the anhysteretic magnetization curve Bah(H) is modeled by the nonlinear variable capacitor C1. Next, the hysteretic losses (caused by the eddy currents) are added by the resistor (R3). The simplified schematic diagram of the gyrator-capacitor circuit for modeling an inductor with the nonlinear magnetic core is presented in Figure 2. The nonlinear relative magnetic characteristic can be modeled by the nonlinear capacitor C1. In such a case [18, 19]: S C1 r H · 0 · e (1) le
8
Se · x (2) le
In Equation 2, x is a voltage at capacitor, V(μr ∙ μ0) is the tabularized form of μr(H)∙μ0 provided to the simulation as a separate nonlinear voltage source, and x is the voltage at the nonlinear capacitor C1 [18]. Presenting nonlinear capacitors in this form in the LTspice model enables stable and efficient solutions to differential equations describing the state of the nonlinear system. The proposed model neglects two additional sources of power losses in the core: static hysteresis losses and excess losses [18]. Although such losses can be replicated in the gyrator-capacitor model [18, 20], in the case of the nanocrystalline ribbon ring core (such as the Vacuumschmeltze W915 made of Fe73.5Cu1Nb3Si15.5B7 alloy), the static hysteresis loop B(H) has an area of about two orders of magnitude smaller than the area of the loop measured at a frequency of 36 kHz. As a result, the static hysteresis is neglected without significantly influencing the model’s accuracy, and an anhysteretic loop represents the nonlinear part of the magnetization curve. The value of excess losses is directly connected to the value of saturation magnetostriction (λs) of anisotropic alloy [21]. The Fe73.5Cu1Nb3Si15.5B7 alloy exhibits nearly zero saturation magnetostriction, which is smaller than 0.5 µm/m. Additionally, measuring excess losses is a complex task with high uncertainty expected [22]. For this reason, the excess losses were neglected in the chosen model for the driving frequency of up to 36 kHz, which is sufficient for most practical cases in high-power switching mode power supplies. Although a nonlinear description of this part of losses is omitted [23], the participation of excess losses in total losses of nearly zero saturation magnetostriction nanocrystalline alloy should be the subject of further research. The anhysteretic magnetization curve, Mah(H), of the Fe73.5Cu1Nb3Si15.5B7 alloy can be modeled using the Jiles-Atherton model [24, 25] with anisotropic extension [26], which is validated for materials with axial anisotropy [27] (parallel or transverse in ribbon ring cores). Such anisotropy occurs in the Fe73.5Cu1Nb3Si15.5B7 alloy due to stresses introduced during rapid quenching, a part of the production process for nanocrystalline materials [28]. The anhysteretic magnetization curve is given by the following set of equations with energies E(1) and E(2) determining the axial anisotropy [27]: e E (1) E ( 2) sin cos ·d M ah ( H ) M s o E (1) E ( 2) (3) sin ·d o e E 1
He K an cos sin 2 (4) o M s a a
Journal of Automation, Mobile Robotics and Intelligent Systems
E 2
VOLUME 20,
N° 3
2026
He K an cos sin 2 (5) o M s a a
Ms is the saturation magnetization, He is H+αM, H is a magnetizing field, α is quantifying the interdomain coupling, M is the total magnetization of the material, and a quantifies the domain wall density in the material. In the above equations, Kan is the average axial magnetic anisotropy energy density, and ψ is the angle between the direction of the magnetizing field H and the material’s magnetization easy axis. Finally, the anhysteretic magnetic relative permeability μah(H), required in Equation 2, and the anhysteretic magnetization curve Bah(H) can be calculated as:
ah H
M ah H H
(a)
(6)
Bah ( H ) 0 ( M ah ( H ) H ) (7)
where μ0 is the magnetic constant. The Jiles-Atherton model parameters of the anisotropic anhysteretic magnetization curve were identified using a differential evolution-based minimization process [29, 30]. The target function F for optimization was given as: F in 1 ( Banh model ( H i ) Bmeas ( H i )) 2 (8)
Banh model(Hi) was the result of modeling based on Equations 1–8, and Bmeas(Hi) was the result of quasistatic measurements carried out for the driving frequency (f) of 1 Hz. Considering the relatively narrow quasistatic hysteresis loop of the Fe73.5Cu1Nb3Si15.5B7 alloy, the anhysteretic curve was approximated by the average value from the hysteretic curve for up and down magnetization. The set of parameters identified during the minimization process is given in Table 1. Figures 3a and 3b present the results of an anhysteretic curve Bah(H) fitting to the quasistatic hysteresis loop B(H) of the Fe73.5Cu1Nb3Si15.5B7 alloy and the results of modeling the relative permeability dependence μah(H), respectively. Notably, the presented experimental results of the B(H) hysteresis loop measurements don’t reach physical magnetic saturation. However, observed saturation is valid from the technical point of view. For this reason, the presented values of saturation magnetization Ms and saturation flux density Bs should be considered technical saturation values, which are slightly lower than the declared values of physical saturation. Table 1. The parameters of the anhysteretic magnetization curve of the Fe73.5Cu1Nb3Si15.5B7 nanocrystalline alloy determined during the optimization process Parameter
Units
Value
Ms
A/m
7.455⋅105
a
A/m
Kan ψ
-
J/m
3
deg
8.587⋅10-7 0.489 4.054 90
(b) Figure 3. The results of an anhysteretic curve modeling for the nanocrystalline Fe73.5Cu1Nb3Si15.5B7 alloy: (a) Bah(H) fitting to the quasistatic hysteresis loop B(H) (black – model, red – results of experimental measurements), (b) the results of modeling of its relative permeability dependence μah(H) The presented results confirm that the inductive core Vacuumschmeltze W915, made of the Fe73.5Cu1Nb3Si15.5B7 nanocrystalline alloy, exhibits transverse anisotropy introduced during the production process [31]. Moreover, the anhysteretic curve properly reproduces the nonlinear shape of the magnetization curve of the alloy. The complete gyrator-capacitor circuit used for the modeling is presented in Figure 4. The coil’s driving current waveform, acquired from measurements using the hysteresis graph setup, is provided as a text file Icurr.txt describing the current source I1. The modeled core’s nonlinear anhysteretic magnetization characteristic is tabulated in the description of voltage sources B6 as a function of magnetic permeability μah versus magnetizing field H, according to Equation 6. LTspice enables the even use of long tables, simplifying the precise description of nonlinear anhysteretic magnetization characteristics. 9
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 4. The complete gyrator-capacitor circuit used for the modeling B(H) hysteresis loop as a function of magnetizing field frequency f for the Fe73.5Cu1Nb3Si15.5B7 (Vitroperm 500F) nanocrystalline alloy
In the gyrator-capacitor model presented in Figure 4, the eddy current losses are modeled by the resistor R2. As previously presented by Chikazumi, the instantaneous power losses per unit volume can be calculated as [32]: d 02 dB 2 dt 2
ped (t )
(9)
ρ is the material’s resistivity, d0 is the lamination thickness, and β is connected with the core’s geometry. The shape factor β is 6 for laminations, 16 for cylinders, and 20 for spheres [6]. The ked value of resistor R2 can be calculated as [19]: ked
10
Vc d 02
2 Ac2
(10)
Vc is the volume of the core, and Ac is the core’s effective cross-section. It should be noted that, since amorphous alloys are typically ribbon-wound cores, the effective cross-sectional area Ac should be estimated based on the core mass and material density. As a result, the value Ac is the value provided by the core’s producer. In addition, considering the physical aspects of the gyrator-capacitor [19], the parameter ked is given in 1/Ω. The inductor’s output signal is collected by the instrumental amplifier and integrated by the operational amplifier U2, which serves as the integrator. Finally, the values of the magnetic field strength H and flux density B in the nonlinear core of the inductor are calculated using the voltage sources B1 and B4, respectively [20]. The proposed model operates in the transient mode. Time plots of the magnetic field strength H and flux density B in the nonlinear core of the inductor are written to output files, enabling further processing. Identification of the parameter determining the eddy current loss For the determination of the optimal value ked of resistor R2, the target function G(ked) was determined as:
Figure 5. The results of modeling the target function G dependence on the value of ked for the resistor R2
Figure 6. The dependence of quality Q(i) of regression versus the order of the polynomial used for the approximation n
G ked Bg cmodel Hi i 1
2
ked
Bmeas Hi (11)
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
(a)
(b)
(c)
(d)
(e)
(f)
N° 3
2026
Figure 7. The comparison of the experimental results and the results of modeling using the gyrator-capacitor model (black – model, red – results of experiments) for the frequency f of the magnetizing field equal: (a) 1 Hz, (b) 2 kHz, (c) 6 kHz, (d) 15 kHz, (e) 26kHz, (f) 36kHz
11
Journal of Automation, Mobile Robotics and Intelligent Systems
Bg–c model(Hi) are the results of modeling the flux density B in the core, using the model presented in Figure 4, for the determined value of resistor R2 equal to ked. The results of the modeling are presented in Figure 5. It can be observed that the value of the function G(ked) exhibits numerical uncertainty. This uncertainty is caused by the measurement uncertainty described in Section 2 and the numerical uncertainty associated with solving ordinary differential equations [5], as stated in the numerical model presented in Figure 4. As a result, it would be difficult to identify the ked value in the optimization process. To overcome this problem, the polynomial response curve method [33] was applied. In this method, the obtained G(ked) dependence results are modeled by a polynomial curve, using the least squares method [34]. The order of the polynomial curve, i equals 3, was determined experimentally, considering the quality Q(i) of regression (least squares value of the differences between the model and results) presented in Figure 6. Finally, the optimal value of parameter ked was estimated as 0.0856 1/Ω.
4. Results of modeling
Figure 7 compares the experimental and modeling results using the gyrator-capacitor model. The gyrator-capacitor model represents the frequency dependence of the shape of the magnetic B(H) hysteresis loop of Vacumschmeltze W915 core made of nanocrystalline Vitroperm 500F very well. The goodness of fit is confirmed by the value of the R2 determination coefficient, which exceeds 0.995 for all loops measured at the magnetizing field frequency, ranging from quasistatic to 36 kHz. the parameters of the Vacuumschmeltze W915 core provided by the manufacturer were as follows: The effective cross-section Ac is 5.94⋅10-6 m2; the magnetic path length Lc is 2.56⋅10-3 m, and the material’s generalized resistivity ρ is 1.15⋅10-6 Ωm, because the
Figure 8. The frequency f dependence of the coercive field Hc value (black – model, red – results of experimental measurements) 12
VOLUME 20,
N° 3
2026
nanocrystalline ribbon thickness is about 30 μm. As a result, the value of the core’s shape parameter β (given by Equation 10) is 19.5. This value indicates that the layers of the nanocrystalline ribbon in the Vacuumschmeltze W915 core are only partially isolated [35]. This observation aligns with previous experimental measurements of nanocrystalline soft magnetic cores produced by annealing amorphous alloys [36]. Although the described gyrator-capacitor model very accurately represents the frequency dependence of the magnetic B(H) hysteresis loop shape of the Vacuumschmeltze W915 core, the more detailed analysis of the frequency dependence of the coercive field Hc value (presented in Figure 8) clearly shows that the root-squareshaped curve of the modeled Hc(f) curve is not a good fit to previous experimental results. This is connected to the fact that Chikazumi’s model (given by Equation 9) is simplified [32] and doesn’t describe all effects related to eddy current power losses in the laminated materials [37]. A more detailed functional model of the Hc(f) curve and its implementation in the gyrator-capacitor model should be considered as a goal of further research.
5. Conclusions
The presented results confirm that the modeling results based on the gyrator-capacitor model with eddy current losses, as described by a resistor, are in agreement with other experimental measurements. In the case of the Vacuumschmeltze W915 core made of nanocrystalline Fe73.5Cu1Nb3Si15.5B7 alloy (Vitroperm 500F), this agreement is confirmed by the value of the R2 determination coefficient, which exceeds 0.995 for all loops measured for the magnetizing field frequency, from quasistatic up to 36 kHz. In addition, considering the core’s parameters provided by the manufacturer, the value of the core’s shape parameter β was estimated as 19.5, indicating that the layers of the nanocrystalline ribbon in the Vacuumschmeltze W915 core are only partially isolated. On the other hand, the frequency dependence of the coercive field Hc value indicates that the rootsquare-shaped curve Hc(f) exhibits limited fit to experimental results. As a result, further research of more detailed functional models of the Hc(f) curve and their implementation in the gyrator-capacitor model should be conducted.
AUTHORS Roman Szewczyk* – Warsaw University of Technology, Faculty of Mechatronics, ul. ś� w. Andrzeja Boboli 8, 02-525 Warsaw, Poland, email: roman.szewczyk@ pw.edu.pl. Piotr Gazda – Warsaw University of Technology, Faculty of Mechatronics, ul. ś� w. Andrzeja Boboli 8, 02-525 Warsaw, Poland , email piotr.gazda@pw.edu.pl. Michał Nowicki – Gediminas Technical University (VILNIUS TECH), Plytinė� s St. 25, LT-10105 Vilnius, Lithuania, email: michal.nowicki@pw.edu.pl Paweł Nowak – Warsaw University of Technology, Faculty of Mechatronics, ul. ś� w. Andrzeja Boboli 8, 02-525
Journal of Automation, Mobile Robotics and Intelligent Systems
Warsaw, Poland, email: pawel.nowak2@pw.edu.pl. Na Peng – Key Laboratory of Coal Conversion and New Carbon Materials of Hubei Province & Institute of Advanced Materials and Nanotechnology, College of Chemistry and Chemical Engineering, Wuhan University of Science and Technology, Wuhan 430081, P.R. China and Belt and Road Joint Laboratory on Measurement and Control Technology, Huazhong University of Science and Technology, Wuhan, Hubei 430074, China, email: pengna@wust.edu.cn. Adam Bieńkowski – Warsaw University of Technology, Faculty of Mechatronics, ul. ś� w. Andrzeja Boboli 8, 02-525 Warsaw, Poland, email : adam.bienkowski@ pw.edu.pl. Tomasz Charubin – Warsaw University of Technology, Faculty of Mechatronics, ul. ś� w. Andrzeja Boboli 8, 02-525 Warsaw, Poland, email : tomasz.charubin@ pw.edu.pl. ∗ Corresponding author
References
[1] M.R. Kiran et al., “A comprehensive review of advanced core materials-based high-frequency magnetic links used in emerging power converter applications,” IEEE Access, vol. 12, 2024, pp. 107769–107799. [2] C. Enkrich et al., “Magnetic metamaterials at telecommunication and visible frequencies,” Physical Review Letters, vol. 95, no. 20, 2005,; doi:10.1103/PhysRevLett.95.203901 [3] P.K. Mallick, “Advanced Materials for Automotive Applications: An overview,” Advanced Materials in Automotive Engineering, 2012, pp. 5–272012; doi:10.1533/9780857095466.5 [4] Pandey V., R. Sharma, “Inductive Components Market Research Report 2032,” Research Report 2032, https://dataintelo.com/report/globalinductive-components-market (accessed Nov. 15, 2024). [5] R. Pratap, V. Agarwal, and R.K. Singh, “Review of various available spice simulators,” 2014 International Conference on Power, Control and Embedded Systems (ICPCES), 2014, pp. 1–6; doi: 10.1109/icpces.2014.7062809. [6] D. Jiles, Introduction to Magnetism and Magnetic Materials, CRC Press, 2017. [7] M. Hasiak et al., “Microstructure and magnetic properties of NANOPERM-type soft magnetic material,” Acta Physica Polonica A, vol. 135, no. 2, 2019, pp. 284–287. [8] V. Bertolini et al., “Eddy current losses model and physical parameters evaluation for ferrite magnetic cores in frequency domain,” Journal of Magnetism and Magnetic Materials, vol. 594, 2024; doi:10.1016/j.jmmm.2024.171905. [9] H. Fujimori et al., “Anomalous eddy current loss and amorphous magnetic materials with low core loss (invited),” Journal of Applied Physics, vol. 52, no. 3, 1981, pp. 1893–1898.
VOLUME 20,
N° 3
2026
[10] P. Jabłoń� ski, M. Najgebauer, and M. Bereź� nicki, “An improved approach to calculate eddy current loss in soft magnetic materials based on measured hysteresis loops,” Energies, vol. 15, no. 8, 2022; doi:10.3390/en15082869. [11] E. Dlala et al., “Interdependence of hysteresis and eddy-current losses in laminated magnetic cores of electrical machines,” IEEE Transactions on Magnetics, vol. 46, no. 2, 2010, pp. 306–309. [12] B. Koprivica, and K. Chwastek, “Verification of Bertotti’s loss model for non-standard excitation,” Acta Physica Polonica A, vol. 136, no. 5, 2019, pp. 709–712. [13] G. Bertotti, F. Fiorillo, and G.P. Soardo, “The prediction of power losses in soft magnetic materials,” Le Journal de Physique Colloques, vol. 49, no. C8, 1988; doi:10.1051/jphyscol:19888867. [14] F. Asadi, “Simulation of Electric Circuits with LTspice®,” Essential Circuit Analysis using LTspice®, 2022, pp. 1–175; doi:10.1007/978-3031-09853-6_1. [15] B.H.D Tellegen, “The gyrator, a new electric network element,” Philips Research Reports, vol. 3, pp. 81–101, 1948. [16] D.C. Hamill, “Lumped equivalent circuits of magnetic components: The gyrator-capacitor approach,” IEEE Transactions on Power Electronics, vol. 8, no. 2, 1993, pp. 97–103. [17] M. Lambert et al., “Magnetic circuits within electric circuits: Critical Review of existing methods and new mutator implementations,” IEEE Transactions on Power Delivery, vol. 30, no. 6, 2015, pp. 2427–2434. [18] Q. Chen et al., “Gyrator-capacitor simulation model of Nonlinear Magnetic Core,” 2009 Twenty-Fourth Annual IEEE Applied Power Electronics Conference and Exposition, 2013, pp. 1740– 1746; doi:10.1109/APEC.2009.4802905. [19] H. Zhang et al., “Improved gyrator–capacitor model considering eddy current and excess losses based on loss separation method,” AIP Advances, vol. 10, no. 3, 2020; doi:10.1063/1.5143172. [20] R. Szewczyk et al., “Improved gyrator-capacitor modeling of inductive components with a FINEMET-type nanocrystalline alloy core using SPICE,” Journal of Magnetism and Magnetic Materials, vol. 555, 2022; doi:10.1016/j. jmmm.2022.169376. [21] K. Suzuki, “Recent Advances in Nanocrystalline Soft Magnetic Materials: A Critical Review for Way Forward,” SSRN Electronic Journal, 2023;doi: 10.2139/ssrn.4621685. [22] B. Jez et al., “Share of Additional Losses in Total Core Losses in the Remagnetization Process of Amorphous FeB-Based Alloys,” Acta Physica Polonica A, vol. 147, no. 3, Apr. 2025, p. 159. [23] D.-X. Chen et al., “Anomalous loss factor of annealed nearly non-magnetostrictive
13
Journal of Automation, Mobile Robotics and Intelligent Systems
amorphous wire,” Journal of Magnetism and Magnetic Materials, vol. 221, no. 3, 2000, pp. 317–326. [24] D.C. Jiles and D.L. Atherton, “Theory of ferromagnetic hysteresis (invited),” Journal of Applied Physics, vol. 55, no. 6, 1984, pp. 2115–2120. [25] D.C. Jiles and D.L. Atherton, “Theory of ferromagnetic hysteresis,” Journal of Magnetism and Magnetic Materials, vol. 61, no. 1–2, 1986, pp. 48–60. [26] A. Ramesh, D.C. Jiles, Y. Bi, “Generalization of hysteresis modeling to anisotropic materials,” Journal of Applied Physics, vol. 81, no. 8, 1997, pp. 5585–5587. [27] R. Szewczyk, “Validation of the anhysteretic magnetization model for soft magnetic materials with perpendicular anisotropy,” Materials, vol. 7, no. 7, 2014, pp. 5109–5116. [28] M.A. Willard, M. Daniil, “Nanocrystalline soft magnetic alloys two decades of progress,” Handbook of Magnetic Materials, vol. 21, 2013, pp. 173–342; doi:10.1016/B978-0-444-595935.00004-0. [29] R. Storn and K. Price, “Differential Evolution - A Simple and Efficient Heuristic for Global Optimization over Continuous Spaces,” Journal of Global Optimization, vol. 11, no. 4, 1997, pp. 341–359. [30] R. Szewczyk, “Two step, differential evolution-based identification of parameters of Jiles-Atherton model of magnetic hysteresis loops,” Advances in Intelligent Systems and Computing, 2018, pp. 635–641; doi:10.1007/978-3319-77179-3_60.
14
VOLUME 20,
N° 3
2026
[31] F. Mazaleyrat et al., “A novel method determining longitudinally induced magnetic anisotropy in amorphous and Nanocrystalline Soft Materials,” Journal of Magnetism and Magnetic Materials, vol. 280, no. 2–3, 2004, pp. 391–394. [32] S. Chikazumi and S.H. Charap, Physics of Magnetism, R.E. Krieger, 1986. [33] R.H. Hardin and N.J.A. Sloane, “A new approach to the construction of optimal designs,” Journal of Statistical Planning and Inference, vol. 37, no. 3, 1993, pp. 339–369. [34] A.C. Atkinson, A.N. Donev, and R. Tobias, Optimum Experimental Designs, with SAS, Oxford University Press, 2023. [35] A. Kolano-Burian et al., “High-frequency soft magnetic properties of Finemet modified with Co,” Journal of Magnetism and Magnetic Materials, vol. 316, no. 2, 2007; doi:10.1016/j. jmmm.2007.03.116. [36] C. Beatrice et al., “Broadband magnetic losses of nanocrystalline ribbons and powder cores,” Journal of Magnetism and Magnetic Materials, vol. 420, 2016, pp. 317–323; doi:10.1016/j. jmmm.2016.07.045. [37] S. Singh et al., “Homogenization and eddy current loss approximation of soft magnetic composite material for electrical machines,” Proc. 10th International Conference on Power Electronics, Machines and Drives (PEMD 2020), vol. 2020, no. 7, pp. 1024–1029.
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
MODIFIED HYBRID GRAY WOLF-CUCKOO SEARCH ALGORITHM FOR OPTIMUM TUNING OF THE PI CONTROLLER FOR CONTROLLING THE TRANSIENTS OF THE DOUBLY-FED INDUCTION GENERATOR Submitted: 6th August 2024; accepted: 2nd October 2024
Ashutosh Kashiv, H. K. Verma, Nagendra Singh DOI: 10.14313/jamris-2026-034 Abstract: Double-fed induction generators play an important role in the wind energy sector due to their ability to operate at variable speeds, superior efficiency, and compatibility with the grid. The optimal controller design for a doubly-fed induction generator ensures efficient power conversion, stable operation, and fault tolerance in wind turbines. To maintain stability and control active and reactive power, this article suggests the design of a PI controller. This article suggests the design of a new PI controller to control the stability and active and reactive power of a doubly-fed induction generator. The PI controller gains are tuned in such a way that they control the operation of the doubly-fed induction generator. PI controller parameter tuning is done by a novel modified hybrid Gray Wolf Cuckoo search algorithm. Furthermore, this work applied a modified Gray Wolf Cuckoo search algorithm for the tuning of PI controller parameters. In this work, the proposed controller has been tested for a 2 MW DFIG wind power system. The results obtained from the proposed controller are compared with the results of the Jaya optimization and Whale optimization algorithms. The results show that the proposed method tuned PI controller parameters accurately, hence improving the stability and efficiency of the DFIG model. Keywords: Doubly Fed Induction Generator (DFIG), Modified Hybrid Gray Wolf Cuckoo Search Algorithm (MGWOCS), Wind Energy Conversion System (WECS), Proportional Integral controller (PI)
1. Introduction Electricity is important to improve people’s living standards. Without electricity we cannot imagine life. It plays an important role in enhancing the economy of any country [1]. Fossil fuels are available all the time and most generation plants use it for the generation of electrical power. Since the sources of fossil fuels are limited and emit toxic gases in the environment, it is required to use substitutes [2]. Now we can focus on renewable energy sources that are environmentally friendly and fulfill the demands of the people. A wind energy conversion system is the best substitute for classical fossil fuel generation plants.
Figure 1. Schematic diagram of the DFIG-WECS
Wind power plants have one major problem: The airflow is not constant. Due to such changes in airflow, transients arise in wind generators. These transients induced many power quality issues in the generated power. Transients can affect the reactive power, frequency, and voltage profiles. Control of the transient operation of a wind power plant with a doubly-fed induction generator using PI controller is implemented in this work. For accurate control of the DFIG generator gain, the PI controller is tuned using a hybrid optimization technique. The DFIG is a type of electrical generator used in wind turbines to generate renewable electrical energy [6]. DFIG wind energy conversion system is shown in Figure 1. For stability of operation and control of active and reactive power, the stator and rotor currents of DFIG must be controlled. The main objective of this work is to improve the transient response of the DFIG-WECS. For this purpose, a PI controller is mathematically modeled. After that, the goal is to find the optimized tuning of PI controller parameters used to control the operation of DFIG-WECS. A novel modified hybrid (Gray Wolf and Cuckoo Search) algorithm is used to find out optimal values of the PI controller gain parameters (Kp and Ki). The output results of the DFIG are verified by comparing the results in terms of THD values. The unique approach of this suggested work is the use of the modified hybrid algorithm which has never been suggested by the any other researcher for DFIG operation and control.
Open Access. © 2026 Ashutosh Kashiv et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
15
Journal of Automation, Mobile Robotics and Intelligent Systems
In the first section of the paper, an introduction to the current scenario of renewable and wind power generation is given, and it is explained that DFIG-WECS are a very important part of the wind power system. Section 2 presents the literature review and existing research in this field and discusses the research gap. The mathematical modeling of DFIG and its controller design are explained in Section 3 of the paper. In the subsection of this section, the vectororiented control strategy is applied to the design of a rotor-side controller. The MGWOCS optimization method was developed by proposing small modifications to the hybrid methods of the Cuckoo search algorithm and the Gray Wolf algorithm describe in Section 4. In the Section 5, the implementation of MGWOCS in an experimental model is explained. In the same section, we implemented simulation of Jaya optimization, MGWOCS method, and the Whale optimization method for tuning of PI controller parameters. The results of the proposed experimental model are obtained in terms of graph and numeric values and listed comparatively for results of the Jaya optimization and Whale optimization algorithm. The final section of the paper draws conclusions from the research work.
2. Related Work Vector and direct control methods are suggested for the stable operation of the DFIG system. Such control schemes are good but fail when loads are increased [1]. Wind conversion system with maximum power point tracking to enhance the harvest of wind generation is suggested in article [2]. MPPT moves the position of the head of the generator in the direction of the flow of air, and hence the rotor of the wind generator maintains good speed and generates high power [2]. The slip power recovery control technique used in article [3]. They used a power converter attached to the rotor windings to recover part of the slip power The productivity of the entire system can be improved by feeding this recovered energy back into the grid. Pitch control in turbines refers to changing the angle of the turbine blades to regulate the aerodynamic load on the rotor. This helps to prevent over speeding and overloading of the generator in high wind flow conditions [4]. The DFIG’s power converter enables precise control of active and reactive power. This is crucial for grid integration and stability, as the generator can respond to grid demands and support voltage regulation [5]. The power converter of the DFIG can regulate the terminal voltage of the generator and thus ensure that it remains within the permissible limits even under changing operating conditions. In the event of shortterm faults, DFIG systems must maintain the grid connection and continue operation [6]. In many cases, DFIG systems must remain connected to the grid in order to operate during short-term grid faults [7]. Modern control techniques allow the generator to withstand these faults without shutting down. By 16
VOLUME 20,
N∘ 3
2026
limiting the maximum permissible rotor current [8], the control system can prevent overloading and overheating of the rotor and associated power electronics. These techniques are often implemented with advanced control algorithms and real-time monitoring systems to ensure optimal performance, grid stability, and efficient energy conversion in DFIG-WECS [9]. The optimal design of proportional-integralderivative (PID) controllers for DFIG wind generators is suggested in article [10]. They control the DFIGbased wind generators with a grid interconnection system. The fundamental principles of PID control, which include proportional, integral, and derivative components, form the basis for developing effective controllers for DFIG wind generators used by [11]. Researchers have explored various tuning methods to optimize PID parameters. Classical approaches such as Ziegler-Nichols [12] and Internal Model Control (IMC) [13] have been adapted and extended for DFIG applications. More recent studies have explored advanced optimization techniques and adaptive strategies for tuning the PI controller in DFIG wind generators. Metaheuristic algorithms, such as particle swarm optimization [14], genetic algorithms [15], and artificial bee colony [16], have been employed to search for optimal PI parameter sets that balance stability, performance, and energy capture. In addition, adaptive PI control strategies that adjust controller gains in real time based on system dynamics have shown better transient response and higher robustness. The integration of advanced control methods with PI controllers has shown promise in improving the overall performance of DFIG wind generators [17]. Fuzzy logic, artificial neural networks, and model predictive controllers (MPC) have been combined with PI strategies to improve fault tolerance, power tracking, and grid synchronization [18]. These hybrid approaches leverage the strengths of multiple techniques to overcome complex challenges. The study of the transient behavior of DFIG, which includes faults in DC links, converters, rotor terminals, and other forms of stator terminal faults, such as phase-to-phase and phase-to-ground, is still lacking in the literature. However, these studies are only able to investigate and analyze the response of DFIG in threephase short circuits. There is still a lack of research on the optimal design of regulators. To fill this gap, swarm intelligence is used in this proposed study, namely the modified Gray Wolf Optimization (GWO) method. The aim of this work is to design an optimal controller for DFIG-WECS. A novel modified hybrid Gray Wolf Cuckoo search algorithm is proposed and applied for obtaining the optimum parameters of the PI controller. To the best of the author’s knowledge, this optimization method has never been applied to DFIG-WECS in the past.
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
3 3 Re{Vr . ir ∗ } = �vqr idr − vdr iqr � 2 2 3 d 3 L m Te = {λs ⋅I∗r } = �λ i − λds iqr �� � 2 2 dt Ls qs dr Qr =
2026
(12) (13)
In above equations “*” denotes the complex conjugate of quantity. 3.1.
Controller designing using Vector Control Strategy
The vector-oriented control strategy was used for the design of a rotor-side controller [21]. The basic equations of current, flux and Emf of DFIG is given by assuming a reference frame connected to the stator flux (i.e. 𝜆𝑑𝑠 = 𝜆𝑠 , 𝜆𝑞𝑠 = 0). Figure 2. Three-dimensional positioning of the vectors
1 �λ − Lm idr � Ls ds −Lm iqs = �iqr � Ls Lm λdr = λ + σLr idr Ls ds λqr = σLr iqr
ids =
3. Mathematical Modeling Mathematical modeling of the DFIG using the direct-quadrature (d-q) transform is a fundamental step in the development of the controller to control the transient behavior of DFIG systems [19]. The d-q transform, often referred to as the Park transform, provides a powerful mathematical tool to simplify the analysis of three-phase system as shown in figure 2. The d-q transformation comprises two orthogonal components: The direct (d-axis), which is connected to the rotor flow, and the quadrature (q-axis), which is orthogonal to the d-axis. The three-dimensional distribution of the dq-transformation is explained in mathematical equations as follows [20] Vqs = rs iqs + ωs λds + λ′qs Vds = rs ids − ωs λqs + λ′ds Vqr = rr iqr + σLr i′qr + (ωs − ωr )λdr + λ′qr Vdr = rr idr − (ωs − ωr )λqr + λ′dr
(1)
(4)
Flux linkages for stator and rotor in d-q frame are given as: λqs = Ls iqs + Lm iqr
(5)
λds = Ls ids + Lm idr
(6)
λqr = Lm iqs + Lr iqr
(7)
λdr = Lm ids + Lr idr
(8)
In the above equations, Vqs , Vds , Vqr , Vdr , iqs , ids , iqr , idr , λqs , λqr , λqr , λdr , λds are the stator and rotor emf, currents, and the flux-linkages respectively, 𝑟𝑠 and 𝑟𝑟 denote resistance, 𝐿𝑠, 𝐿𝑟, self-inductance of stator and rotor windings, 𝐿𝑚, is the mutual-inductance and 𝜔𝑠, 𝜔𝑟 are the synchronous and rotor angular speed. The active, reactive power, and torque in the d-q frame in mathematical form can be written as: 3 3 Re{Vs . is ∗ } = �vds ids + vqs iqs � 2 2 3 3 ∗ Pr = Re{Vr . ir } = �vdr idr + vqr iqr � 2 2 3 3 ∗ Qs = Im{Vs . is } = �vqs ids − vds iqs � 2 2 Ps =
(9) (10) (11)
(15) (16) (17)
Equations (1) to (4) can be modified as follows (assuming stator resistance is negligible, and stator flux linkage is constant (𝜆′𝑑𝑠 = 0),) vds = 0
(18)
vqs = ωs λds Vdr = rr idr + σLr i′dr − (ωs − ωr )σLr iqr Vqr = rr iqr + σLr i′qr + (ωs − ωr ) �σLr idr +
(2) (3)
(14)
(19) (20) Lm λ � Ls ds (21)
Real and reactive powers and electromagnetic torque is given as: Ps =
−3Lm �vqs iqr � 2Ls
3vqs (λ − Lm idr ) 2Ls ds −3 d Lm Te = �λ i �� � 2 dt Ls ds qr Qs =
(22) (23) (24)
L 2
Where, 𝑠= 1 − m Ls Lr In Equations (20) and (21), by putting −(ωs − ωr )Lr σiqr = error1 Lm λ ) = error2 1 (ωs − ωr ) (σLr idr + Ls ds
(25) (26)
Then we get, Vdr = rr idr + σLr i′ dr + error1 ′
Vqr = rr iqr + σLr i qr + error2
(27) (28)
by solving equations (27) and (28), we get: (vdr − error1 ) rr + s.σLr (vqr − error2 ) iqr = rr + s.σLr
idr =
(29) (30)
17
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 5. PI current control system using close loop feedback system
Figure 3. Depiction of the relationship between voltage and rotor current in d-q frame
Figure 6. Block diagram of RSC controller
4.
Figure 4. Modified system after adding PI block in the d-q System Gain for the current control is given by G (s) =
idr iqr 1 = = vdr vqr rr + s.σLr
(31)
Equation (27) and (28) is shown in Figure 3. The modified system is given in Figure 4 for the d-q axis. After adding PI controller, the governing equations are as follows. vqr = (rr + Lr σs) iqr = �kp +
ki � �iqrref − iqr � (32) s
vdr = (rr + Lr σs) idr = �kp +
ki � (idrref − idr ) (33) s
So, the transfer function of current controller is, ir irref
=
1 �skp + ki � σLr
s2 +
s�rr +kp � σLr
+
ki σLr
This article implemented a new technique, which is a modification of hybrid Gray Wolf and Cuckoo search algorithm (MGWOCS). The MGWOCS method was developed by changes to the hybrid methods of the Cuckoo search and the Gray Wolf algorithm (GWOCS). Advantages of using MGWOCS and GWOCS over other methods is as follows: 1. These techniques have a better balance between exploration and exploitation. 2. These techniques convergence rate is very fast and gives the globally best solution. 3. These techniques handle complex and multimodal problems easily. 4.1.
Cuckoo Search Algorithm (CSA)
The CSA method is based on the behavior of young cuckoos, which fly in the surrounding area to search for new nests efficiently. The random walk model known as Lé vy’s flight, whose name goes back to the mathematician Paul Lé vy, is characterized by the step length and obeys the power law [19].
(34)
A closed-loop feedback PI controller with the current control system is shown in Figure 5 and the proposed control system is sustainable energy grids and networks shown in Figure 6. 18
Optimization Techniques
N (s) = s−1
(35)
The probability distribution length of run or jump steps is given by P (i) = i−x at 1 < x = 3
(36)
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Via Lé vy’s flight Cuckoo search a new position as given by equation (37): Gα = �C1 Kα − AK(𝑡)�
(43)
(37)
Gβ = �C2 Kβ − AK(𝑡)�
(44)
where, α is the length of the step taken after the Lé vy flight distribution, which has variance and infinite mean value [20] and, s is the step range is given as; Lé vy (s, y) ∼s−y , (1 < y = 3) (38)
Gδ = �C3 Kδ − AK(𝑡)� Xg1 = Xα − A1 ∗ (Gα )
(45)
Xg2 = Xβ − A2 ∗ �Gβ �
(47)
Xg3 = Xδ − A2 ∗ (Gδ ) Xg1 + Xg2 + Xg3 X′(t + 1) = 3
(48)
(t+1)
si
4.2.
(t)
= si + α⊕Lé vy(s, y)
Gray Wolf Optimization (GWO) Method
The social structure and hunting strategy of gray wolves serve as inspiration for the metaheuristic method of GWO [21]. This algorithm mimics the hierarchy and hunting patterns found in wolf packs. In GWO, a population of predicted solutions (wolves) is developed and updated in iterations to search for optimal solutions to complex optimization problems. Alpha is the first optimum solution of the GFO’s algorithm; similarly, beta, gamma, and omega are the second, third, and next optimum solutions, respectively. It is assumed that the remaining possible solutions are omega [21]. The hunt in the GWO method is led by these three wolves (α, β, and δ) and followed by the ω-wolves. The alpha wolf, representing the best solution candidate, surrounds the prey (optimal solution) to reduce the search space. The wolves coordinate their hunting efforts to approach and catch the prey together. This hunting behavior is transferred to the optimization process, in which the candidate solutions work together to refine the search space. This hunting process involves (i) chasing, running, and moving towards the prey; (ii) following and circling until the prey keeps moving; and (iii) capturing the prey [21, 22]. The mathematical equations to represent the encircling behavior are given below [22]. G = |CKP − AK (t)|
(39)
K (t+1) = Kp (t) − A ∗ D
(40)
Where t is the initial iteration, A,C and, D are constants, Kp(t) is the gray wolf position vector, K(t) is the gray wolf position vector at initial iteration. Constant A and C are given as: ��⃗ A = 2 �⃗ q �����⃗ n1 − q
(41)
�⃗ C = 2 �����⃗ n2
(42)
Where n1 , and n2 are random vectors between [0, 1] and q decrease from 2 to 0 as the iteration increases. Gray wolf hunting’s best three solutions are saved. The remaining other agents update their potions. It is mathematically expressed by the following equations [23, 24]:
(46)
(49)
Gray wolves attack when their prey pauses movement. The value of A lies randomly between -2n and 2n, and a declines from 2 to 0 with the iterations. Wolves attack the prey if |A| is less than one, which is called exploitation activity. If |A| is greater than one, it indicates the wolves are searching for others, so they diverge from the present prey [25]. This process aims to make a tradeoff between exploration and exploitation, helping the algorithm efficiently navigate the search space and converge towardsoptimal solutions. 4.3.
Modified Hybrid Gray Wolf Optimization and Cuckoo-Search Method (MGWOCS)
To improve the exploration process in the GWO method, a new modified GWOCS technique is used in this work. A new factor γ is applied to amend the encircling and position update calculations [22]. The next phase of the MGWOCS is encircling, which comes after population initialization, where gray wolves work together to encircle the prey before attack. A vector sum approach is used for writing the equation of encircling. Factor γ is applied to measure the distance appropriately and to have the wolves in an improved loop. The improved encircling process of GWO is shown by the following equations [22], �⃗ �⃗ ��⃗ D = 1/γ| G𝑝(i) + G(i)|
(50)
�G(i ⃗ + 1) = 1/γ {�G𝑝(i) ⃗ + ��⃗ A��⃗ D}
(51)
where D is the position from the prey, ����⃗ Gp and ��������⃗ G (i) are position of prey and the current wolf, γ is the scaling factor (γ=1.4 is the optimum value to get effective results) [22]. The MGWO-CS hybrid method is developed by suggesting small changes in the GWO and CSA algorithms. GWO is a metaheuristic technique created based on how gray wolves hunt the prey. In the hybrid MGWOCS algorithm, location updates for GWO are added to the updated position equation of CSA to achieve faster convergence. The equation is changed to include a fourth term in the numerator to generate the position update in the MGWO-CS algorithm, as indicated in the following equation [23]. �G⃗ (t+1) =
�G⃗1 + �G⃗2 + �G⃗3 + �G⃗4 4
(52)
19
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
�⃗1 , G �⃗2 , and G �⃗3 are the positions of hunting Above, G agents. After the best hunting agent’s position Gα , the following and next best hunting position are Gβ and �⃗4 Position vector is estimated using the CS Gδ [23]. G update rule [19]. Every egg in the nest is a solution and is replaced by more effective ones. Such a process is used to update the locations of �G⃗4 terms in the hybrid MGWOCS algorithm. It is given by equations as given below [24]. ���⃗ �⃗t + α⊕Lé vy (λ′ ) G′ 4 = G
(53)
Here, �G⃗t is the candidate’s position in the current iteration, the step dimension α, is multiplied by the input and has a 0 to 1 range. The proposed algorithm gets more effective as �G⃗4 is introduced since it uses Lé vy flights for exploring the search space. Figure 7 shows the flow chart for the implementation of the MGWOCS algorithm. The following steps are adopted for the implementation of the MGWOCS algorithm: 1) Generate the initial population of gray wolf X (Let population = n). 2) Set vector A, C, and a. 3) Compute the fitness of every gray wolf (search agent) and find the three best gray wolves and assign them name α, β, δ. After that find the magnitude of Xα , Xβ , Xδ 4. While t<iterationmax ( For each gray wolf. 1) Initialize random value of r1 , r2 . 2) Find the value of distance Gα , Gβ , Gδ using the equation (51)
Figure 7. Flow chart of the MGWOCS method
5. �⃗ �⃗ ��⃗ D = 1/γ |G𝑝(i) + G(i)|
(54)
3) Find the positionG1 , G2 , G3 using the equation (52) �G(i ⃗ + 1) = 1/γ {�G𝑝(i) ⃗ + ��⃗ A��⃗ D}
(55)
One more term G4 is find out using the equation (56) �⃗4 = G �⃗t + α⊕Lé vy (λ) G (56) Update the position of prey using the equation (55) �G⃗ (t+1) =
�G⃗1 + �G⃗2 + �G⃗3 + �G⃗4 4
(57)
Update vector A, C, and a, then determine the fitness of all gray wolves (search agent) and discover three best gray wolf and assign them name α, β, δ. After that update the magnitude of Xα , Xβ , Xδ . Increase iteration number t=t+1. 5. Repeat step 4 until criteria fulfills or maximum number of iterations achieved.
20
Implementation of the Proposed Techniques
An experimental model of a 2 MW DFIG-operated wind generation system is used in this work. A Simulink model of the DFIG-based wind generation system is developed in MATLAB to test and validate the proposed controller optimization by the MGWOCS method. A simulation model of DFIG [25, 26, 30] is given in Figure 8. The data for the test DFIG system are given in Table 1. The optimum tuning of the PI controller parameter is obtained by running the MGWOCS algorithm in Matlab. These tuned values of PI controller parameters are fed into the Simulink model of DFIG. In this model, a sudden change in load for t = 4 seconds is created to simulate the fault conditions. And the output waveform is plotted for true power, reactive power, emf, and current, as shown in the result section. The Jaya algorithm and the Whale optimization algorithm are also implemented on the same experimental model Simulink. Results obtained by MGWOCS and are compared with the Jaya algorithm and the Whale optimization algorithm.
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 10. Reactive power variation with respect to time for DFIG with different optimization method Figure 8. Simulink model of DFIG Table 1. Datasheet of the 2 MW DFIG systems Specifications Rotor Equivalent Resistance Rr Stator Equivalent resistance Rs Inductance of Rotor Lr Stator load Ls Mutual-Inductance M Real power proportional gain KP Reactive power proportional gain KQ Rated power P Frequency f Generator (EM) Torque Tem
Values 0.00289Ω 0.00259Ω 0.002586 H 0.002587 H 0.0025 H 2.42 2.34 2MW 50 Hz. 13000 N-M
power of DFIG optimized by MGWO-CS and Jaya optimization is much better than the output with the Whale optimization controller. Figure 10 shows the change in reactive power when the controller is implemented over time. The diagram in Figure 10 illustrates that the output power of DFIG optimized by MGWO-CS and Jaya optimization is much better than the output with the Whale optimization controller. When the machine is loaded, load torque is increased, which results in an increase in reactive power demand, which disturbs the voltage profile of the dc link as well. Figure 11 shows that MGWO-CS [30] gives better results and contains fewer transients than Whale optimization and the Jaya algorithm. It is observed here that there is a large overshoot in the bus voltage. This large overshoot is due to an increase in the sensitivity of the system. It is well known that feedback control systems increase the sensitivity of the system, which increases the overshoot [30–32]. This may be considered the drawback of using the feedback control system with the PI controller, but in the DFIG system, stator winding is linked directly to the grid; hence, our main objective is to control the stator output, which is fulfilled in this work. The DC link is not directly linked to the load or grid, so it is not directly affecting the output power quality.
Figure 9. Variation of the active power
6. Results and Analysis In this section, the result waveforms acquired from the experimental model are plotted and compared for different methods. Values obtained through the Jaya optimization method [27, 28], the MGWOCS method, and Whale optimization [29] are fed into the model of DFIG, and the results are mapped. Figure 9 shows the active power variation when the PI controller is tuned by the MGWO-CS method (blue color), the Jaya optimization method (yellow color), and the Whale optimization controller (green color), respectively. Figure 9 depicts that the output
This transient is also not harmful for the grid or load because this DC link is not directly fed to the grid or load. To control this voltage, there is a capacitor or battery in between two converters that acts as an eliminator for these spikes, and this transient in voltage can also be controlled through a grid-side converter. The comparative results of the proposed methods with other techniques are given in Table 2. Responses obtained by different optimization techniques for the proposed practical wind generation model with the DFIG system are shown in Table 2. Table 2 depicts that WOA, MGWO-CS, and JOA are all three methods improving the response and among the three best responses obtained from PI optimized by the MGWO-CS method. 21
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 2. Comparative results of different methods Parameters Rise Time Settling Time Overshoot Undershoot Peak
Normal PI 0.0885 0.1363 1.9099 0.85 1.6171
PI optimized by WOA [28] 0.0314 0.0535 1.12 0.809 1.415
Figure 11. DC link voltage variation with different optimization method
PI optimized by MGWO-CS 0.0263 0.0386 1.08 0.913 1.378
PI optimized by Jaya Optimization 0.0287 0.0412 1.10 0.956 1.504
Figure 13. THD measurement in DFIG optimized by MGWO-CS Optimization
Figure 12. Effect of γ on convergence plot Figure 14. Graph of Convergence for GWO method The best value of a new factor γ, which is used in the suggested MGWOCS to modify the encircling and position update calculations, is also tested by comparing the convergence speed. It is observed that convergence speed varies with γ. Convergence is compared for different values of γ. The values of γ are taken as 0.93, 1, 1.2, and 1.4. From Figure 12, it can be concluded that the function converges fast when the value of γ is 1.4. Figure 13 shows the THD measurement of output current from DFIG optimized by MGWO-CS Optimization. Figure 13 shows the FFT analysis input load is tested at 5 sec, where the fundamental frequency is 50 Hz, and the achieved THD (total harmonic distortion) is 4.10%, which is less than 5% at the acceptable limit [32, 33]. 22
6.1.
Comparative Analysis in Terms of Convergence and THD Values
The convergence characteristics of GWO, MGWOCS, and Jaya optimization methods are shown in Figures 14, 15, and 16, respectively. THD values of stator current obtained through different controllers are also compared. Table 3 shows that THD values and convergence rate are found to be better for the MGWOCS method than the GWO and Jaya optimization methods.
7.
Conclusion
DFIG-WECS is mathematically modeled, and its rotor-side PI controller is successfully designed. The Novel Modified Hybrid Gray Wolf Cuckoo Search
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
methods. The convergence rate is also compared for different values of γ. It is found that the convergence rate depends on γ, and the convergence rate is highest when the value of γ is 1.4. The results of MGWOCS and GWO are illustrated. THD and convergence speed are compared, and it is concluded that convergence speed is increased. Also, THD values are found to be lowest when using a controller optimized through MGWOCS. The work’s future objective is to analyze the efficacy of the MGWOCS method on other DFIG test systems. Also, this method can be compared with more new Metaheuristic optimization methods.
Figure 15. Graph of Convergence for MGWO-CS method
AUTHORS Ashutosh Kashiv – Department of Electrical Engineering, Prestige Institute of Engineering Management and Research, Indore, Madhya Pradesh 452003, India, e-mail: ashutosh.kashiv@gmail.com. H. K. Verma – Department of Electrical Engineering, Shri Govindram Seksaria Institute of Technology and Science, Indore, Madhya Pradesh 452003, India, e-mail: vermaharishgs@gmail.com. Nagendra Singh∗ – Department of Electrical Engineering, Trinity College of Engineering and Technology, Karimnagar, Telangana 505001, India, e-mail: nsingh007@rediffmail.com. ∗
Corresponding author
References
Figure 16. Graph of Convergence for Jaya Optimization Method Table 3. Comparison of obtained THD values and convergence speed of different techniques Technique used for optimization
THD Values
Number of iterations required to converge
1
Jaya Optimization Method
4.49
71
2
GWO Method
5.52
52
3
MGWO-CS Method
4.10
12
S. No.
Algorithm is applied in this work to obtain the optimum value of the parameters of the controller. The value of gains obtained from the optimization method is fed to the 2 MWDFIG Simulink model. Simulation results are compared with those obtained using the Jaya algorithm and the Whale Optimization algorithm. A sudden transient is created at t = 4 seconds, and the waveform of voltage, current, and power is analyzed. From analyzing the waveforms, it is concluded that MGWOCS gives satisfactory results. Performance is also evaluated in terms of rise time, settling time, overshoot time, and undershoot time, and it was found that MGWOCS is giving much better results than other
[1] R. Tiwari and N.R. Babu, “Recent Developments of Control Strategies for Wind Energy Conversion Systems,” Renewable and Sustainable Energy Reviews, vol.66, 2016, pp. 268–285. [2] M.A. Hannan et al., “Power Electronics Contribution to Renewable Energy Conversion Addressing Emission Reduction: Applications, Issues, and Recommendations,” Applied Energy, vol.251, 2019, pp.113–140. [3] N.L. Panwar, S.C. Kaushik, and S. Kothari, “Role of Renewable Energy Sources in Environmental Protection: A Review, “Renewable and Sustainable Energy Reviews, vol.15, no. 3, 2011, pp. 1513–1524. [4] IEA, “World Energy Outlook,” https://www.iea.org/reports/world-energyoutlook-2022 [5] S. Boubzizi, H. Abid, and A. El hajjaji, “Comparative Study of Three Types of Controllers for DFIG in Wind Energy Conversion System,” Prot Control Mod Power Syst, vol.3, 2018, pp.21–33. [6] M.M. Rezaei, “A Nonlinear Maximum Power Point Tracking Technique for DFIG-based Wind Energy Conversion Systems.” Engineering Science and Technology, an International Journal, vol. 21, no. 5, 2018, pp.901–908. 23
Journal of Automation, Mobile Robotics and Intelligent Systems
[7] S. Ram, O.P. Rahi, and V. Sharma, “A Comprehensive Literature Review on Slip Power Recovery Drives,” Renewable and Sustainable Energy Reviews, vol.73, 2017, pp.922–934. [8] C.Z. El Archi et al., “Real Power Control: MPPT and Pitch Control in a DFIG Based Wind Turbine.” 3rd International Conference on Advanced Communication Technologies and Networking (CommNet), 2020, pp.1–6, [9] H. Chojaa et al., “Nonlinear Control Strategies for Enhancing the Performance of DFIG-based WECS under a Real Wind Profile,” Energies, vol.15, no. 18, 2022, pp.145–155. [10] O.P. Bharti et al., “Controller Design for DFIGbased WT using Gravitational Search Algorithm for Wind Power Generation,” IET Renewable Power Generation, vol.15, no. 9, 2021, pp.1956–67. [11] B. Desalegn, D. Gebeyehu, and B. Tamrat, “Evaluating the Performances of PI Controller (2DOF) under Linear and Nonlinear Operations of DFIGbased WECS: A Simulation Study,” Heliyon, vol.8, no. 12, 2022, pp.247-259. [12] A.G. Abo-Khalil et al., “Current Controller Design for DFIG-based Wind turbines using State Feedback Control,” IET Renewable Power Generation, 2020, vol. 13, pp. 1938–1948. [13] X. Yuan et al., “DFIG control Design based on Internal Model Controller,” CICED 2010 Proceedings, Nanjing, China, 2010, pp.1–6. [14] O. P. Bharti, R.K. Saket, and S.K. Nagar, “Controller Design for Doubly Fed Induction Generator using Particle Swarm Optimization Technique,” Renewable Energy, vol.114, 2017, pp.1394–1406. [15] L. Wu et al., “Identification of Control Parameters for Converters of Doubly Fed Wind Turbines based on Hybrid Genetic Algorithm,” Processes, vol.10, 2022, pp.567–578. [16] S. Soued, et al., “Experimental Behaviour Analysis for Optimally Controlled Standalone DFIG System. IET Electric Power Applications,” vol. 13, 2019, pp.1462-1473. [17] S.K. Almas and S. Nagendra, “Review on Power Quality Improvement for Grid Connected Wind Energy Generation System,” International Journal of Grid and Distributed Computing, vol.13, no. 1, 2020, pp.106-112. [18] J.M. De Queiroz, L.S. Barros, and D. Barbosa, “Analysis of DFIG Differential Protection based on Park Transform,” IEEE PES Innovative Smart Grid Technologies Conference - Latin America (ISGT Latin America), 2021, pp. 1-4. [19] X.-S. Yang and S.Deb,.“Cuckoo Search via Lé vy flights,” World Congress on Nature & Biologically Inspired Computing (NaBIC), 2009, pp. 210-214. 24
VOLUME 20,
N∘ 3
2026
[20] S. Chaine, M. Tripathy, and D. Jain, “Non dominated Cuckoo Search Algorithm Optimized Controllers to Improve the Frequency Regulation Characteristics of Wind Thermal Power System,” Engineering Science and Technology, an International Journal, vol. 20, 2017, pp.1092-1105. [21] S. Mirjalili, and A. Lewis, “Grey Wolf Optimizer. Advances in Engineering Software,” vol.69, 2014, pp.46-61. [22] R.A. Khanum et al., “Two New Improved Variants of Grey Wolf Optimizer for Unconstrained Optimization,” IEEE Access, vol.8, 2019, pp.3080530825. [23] N. Singh et al., “Artificial Intelligence Techniques in Power Systems Operations and Analysis,” Auerbach Publications, 2023, pp.1-234. [24] N. Singh et al.,“A Review. Recent Advances in Power Systems. Lecture Notes in Electrical Engineering,”Springer, vol. 699, 2021; doi: 10.1007/9 78-981-15-7994-3_18 [25] H. Xu, X. Liuand J. Su, “An Improved Grey Wolf Optimizer Algorithm Integrated with Cuckoo Search,” 9th IEEE International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications (IDAACS), vol. l, no. 1, 2017, pp.490-493. [26] C. Lotfi et al., “Optimization of a Speed Controller of a WECS with Metaheuristic Algorithms,” Engineering Proceedings, vol. 29, no. 1, 2023, pp.7–15. [27] E. Ayenew et al., “Modelling and Control of Wind Wnergy Conversion System: Performance Enhancement,” “International Journal of Dynamics and Control, 2023, pp.1-24. [28] R.R. Jaya, “A Simple and New Optimization Algorithm for Solving Constrained and Unconstrained Optimization Problems,” International Journal of Industrial Engineering Computations, 2016, vol. 7, no. 1, pp. 19-34. [29] R.R. Jaya, “An Advanced Optimization Algorithm and its Engineering Applications,” Springer International Publishing, 2019, pp.327-338. [30] A. Kashiv and H. K. Verma, “Performance Analysis of Doubly-fed Induction Generator using PID Controller Optimised by Whale Optimisation,” International Journal of Engineering Systems Modelling and Simulation, 2022; doi: 10.15 04/IJESMS.2022.10048733 [31] F.E.V. Taveiros, L.S. Barros, and F.B. Costa, “Backto-Back Converter State-Feedback Control of DFIG (Doubly-Fed Induction Generator)based Wind Turbines,” Energy, vol.89, 2015, pp.896–906.
Journal of Automation, Mobile Robotics and Intelligent Systems
[32] N. Singh et al., “Analysis of Optimum Cost and Size Of the Hybrid Power Generation System using Optimization Technique,” 2023 IEEE 12th International Conference on Communication Systems and Network Technologies (CSNT), 2023, pp. 284–291; doi: 10.1109/CSNT57126.2023. 10134720
VOLUME 20,
N∘ 3
2026
[33] K. Chatterjee and A. Kumar,“Contribution to Power Quality Improving the DFIG Control Using a Reduced Components of Multi-level Inverter,” 10th IEEE International Conference on Communication Systems and Network Technologies (CSNT), 2021, pp.291-294.
25
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
OPTIMAL PID LEVEL CONTROL OF A NONLINEAR SPHERICAL TANK USING A HYBRID WOA–GA OPTIMIZATION APPROACH Submitted: 4th November 2025; accepted: 2nd February 2026
Abhay Pakhare, Sharad Jadhav, Mukesh Patil DOI: 10.14313/jamris-2026-035 Abstract: Level control of spherical tank systems is challenging due to strong process nonlinearity and varying dynamics across different operating regions. Traditional PID tuning methods and standalone optimization techniques often fail to ensure consistent performance under such conditions. In this paper, a novel hybrid Whale Optimization Algorithm–Genetic Algorithm (WOA–GA)–based approach is proposed for PID controller tuning and implemented on a nonlinear spherical tank model. The proposed hybrid optimization strategy combines the rapid exploitation capability of WOA with the global search ability of GA, enabling effective optimization of PID parameters for nonlinear system dynamics. A nonlinear mathematical model of the spherical tank is developed and utilized for both simulation studies and real-time experimental implementation. The performance of the proposed WOAGA–PID controller is evaluated and systematically compared with WOA–PID and GA–PID controllers using standard time-domain performance indices, including ITAE, rise time, settling time, peak time, and maximum overshoot. Both simulation and experimental results demonstrate enhanced transient response, robustness, and reliability, thereby validating the suitability of the proposed method for nonlinear process control applications. Keywords: spherical tank, Modeling, performance analysis, WOA-GA
1. Introduction In process industries, where maintaining exact liquid levels in storage vessels is essential for product quality, safety, and effective operation, liquid level management is a basic [1–3]. Due to their structural strength and capacity to tolerate high internal pressures, spherical tanks are one of the most popular tank geometries [4, 5]. However, due to the nonlinear relationship between liquid volume and height, where the cross-sectional area varies with the square of the liquid height, spherical tank systems show notable nonlinear dynamics [6]. This significant nonlinearity poses a challenge to traditional control techniques, especially fixed-gain proportional-integral-derivative (PID) controllers, which frequently fail to provide adequate performance throughout the system’s whole 26
operational range because of these fluctuating dynamics [7–9]. Because of its simplicity, ease of implementation, and shown performance in a variety of engineering applications, the PID controller continues to be the foundation of industriallevel control [10–12]. However, adjusting PID parameters for nonlinear systems like spherical tanks is difficult since traditional tuning techniques like Ziegler-Nichols or Cohen-Coon frequently result in less-than-ideal performance, particularly when there are considerable nonlinear behavior and outside disturbances [13, 14]. Furthermore, the performance of fixed-parameter PID controllers may degrade in terms of overshoot, settling time, and robustness against process uncertainty. Metaheuristic optimization techniques have drawn a lot of interest in order to automatically tune PID gains by minimizing specified performance parameters (such as ITAE, ISE, IAE, overshoot, and energy consumption) in order to meet these tuning difficulties [15–17]. Among these, two widely researched nature inspired methods are Whale Optimization Algorithm (WOA) and Genetic Algorithms (GA). While WOA mimics the bubble-net hunting method of humpback whales, exhibiting significant global search capabilities, GA mimics evolutionary selection and genetic operations to effectively traverse search spaces [18]. By combining both strategies, PID parameters in challenging nonlinear control systems can be optimized more successfully by utilizing the resilience of GA and the balance between exploration and exploitation of WOA. The efficiency of hybrid and enhanced metaheuristic algorithms in fine-tuning PID and sophisticated controllers for nonlinear systems has been demonstrated in recent studies. In comparison to traditional tuning techniques, hybrid metaheuristic algorithms have been used for optimal PID tuning in a variety of domains, resulting in better set-point tracking and disturbance rejection. Furthermore, research on spherical tank control indicates that closed-loop performance may be greatly enhanced over traditional controllers by model-based and optimization-driven tuning. Despite these developments, there remains a dearth of research on hybrid WOA-GA optimization
Open Access. © 2026 Abhay Pakhare et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
Journal of Automation, Mobile Robotics and Intelligent Systems
for spherical-tank PID level control. This serves as the driving force behind the current study, which develops a hybrid WOA-GA method to adjust PID settings for a nonlinear spherical tank system. By reducing performance metrics like integrated time-absolute error (ITAE), overshoot, and settling time, the suggested approach seeks to provide reliable and ideal control across the operating range [19]. The hybrid WOA-GA tuned PID controller outperforms single optimization algorithm techniques and traditional tuning methods, according to simulation and real-time experimentation findings. The primary contributions include: - Obtaining the mathematical model of a nonlinear spherical tank system. The complete operating region of the tank is taken into account while modeling. - The obtained model is a second-order type, capturing the inherent nonlinear dynamics of the plant and its impact on level control. - Design and deployment of control strategies like GA-PID, WOA-PID, and proposed control strategies, hybrid WOAGA-PID for controlling the level in a spherical tank. - The performance of the control strategies are evaluated through simulations and real-time deployments, such as variations in reference values, disturbances and robustness. - The performance of control strategies are evaluated based on their transient responses and performance metrics such as ISE, IAE and ITAE, and control effoers providing a comprehensive analysis of their effectiveness. The rest of the paper is organized as follows: Section 2 describes the mathematical modeling of the process. Section 3 discusses the design of various controllers used in the system as well as controllers used for comparison. Section 4 describes the simulation and realtime results discussion. Finally, the paper concludes in Section 5.
N∘ 3
2026
Figure 1. Schematic of a spherical tank height ℎ is expressed as, 𝑉(ℎ) = 𝜋ℎ2 �𝑅 −
ℎ � 3
(2)
The cross-sectional area of the tank at height ℎ is obtained by differentiating the volume with respect to ℎ: 𝑑𝑉 = 𝜋 �2𝑅ℎ − ℎ2 � 𝑑ℎ
(3)
𝑑ℎ 𝑑𝑉 = 𝐴(ℎ) 𝑑𝑡 𝑑𝑡
(4)
𝐴(ℎ) = Using the chain rule,
Substituting this into the mass balance equation gives 𝐴(ℎ)
𝑑ℎ = 𝑓𝑖𝑛 (𝑡) − 𝑓𝑜𝑢𝑡 (𝑡) 𝑑𝑡
(5)
Hence, the nonlinear dynamic model of the spherical tank is 𝑑ℎ 𝑓𝑖𝑛 (𝑡) − 𝑓𝑜𝑢𝑡 (𝑡) = (6) 𝑑𝑡 𝜋 (2𝑅ℎ − ℎ2 ) The outlet flow is governed by gravity through a valve and it is given by,
2. Modeling of System System model identification of a nonlinear spherical tank involves deriving a mathematical representation that captures the tank’s dynamic behavior. Due to the construction geometry, the relationship between the inlet flow rate and liquid level is not linear. Accurate modeling typically requires experimental data and techniques like system identification. A spherical tank of radius 𝑅 containing an incompressible liquid. Let the liquid level at time 𝑡 be ℎ(𝑡). The inlet flow rate is denoted by 𝑓𝑖𝑛 (𝑡) and the outlet flow rate by 𝑓𝑜𝑢𝑡 (𝑡). Figure 1 shows the schematic of spherical tank system. The fundamental mass balance for the tank is given by, 𝑑𝑉 = 𝑓𝑖𝑛 (𝑡) − 𝑓𝑜𝑢𝑡 (𝑡) 𝑑𝑡
VOLUME 20,
𝑓𝑜𝑢𝑡 (𝑡) = 𝐶𝑑 �2𝑔ℎ
(7)
where 𝐶𝑑 is the discharge coefficient, and 𝑔 is the acceleration due to gravity. Substituting equation (7) into the dynamic equation (6) yields, 𝑑ℎ 𝑓𝑖𝑛 (𝑡) − 𝐶𝑑 �2𝑔ℎ = 𝑑𝑡 𝜋 (2𝑅ℎ − ℎ2 )
(8)
At steady state, the liquid level remains constant, and therefore 𝑓𝑖𝑛 = 𝑓𝑜𝑢𝑡 (9) which leads to
(1)
where 𝑉 is the volume of liquid in the tank. For a spherical tank of radius 𝑅, the volume of liquid up to
𝑓𝑖𝑛 = 𝐶𝑑 �2𝑔ℎ𝑠
(10)
where ℎ𝑠 denotes the steady-state liquid level. The inlet flow is manipulated through a pump whose 27
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
dynamics can be approximated by a first-order system: 𝑄𝑖𝑛 (𝑠) 𝐾𝑎 𝐺𝑎 (𝑠) = (11) = 𝑈(𝑠) 𝜏𝑎 𝑠 + 1 where 𝐾𝑎 is the actuator gain and 𝜏𝑎 is the actuator time constant. The overall transfer function from the control input 𝑈(𝑠) to the liquid level 𝐻(𝑠) is obtained by cascading the actuator and the process models: 𝐺(𝑠) =
𝐻(𝑠) = 𝐺𝑝 (𝑠)𝐺𝑎 (𝑠) 𝑈(𝑠)
(12)
Thus, the second-order transfer function of the spherical tank level system is 𝐺(𝑠) =
𝐾𝑎 𝐾𝑝 (𝜏𝑎 𝑠 + 1)(𝜏𝑝 𝑠 + 1)
(13)
Expanding the denominator, the transfer function can be written in standard second-order form as, 𝐺(𝑠) =
𝐾 𝜏𝑎 𝜏𝑝 𝑠 2 + (𝜏𝑎 + 𝜏𝑝 )𝑠 + 1
Figure 2. Block diagram of level setup
(14)
where 𝐾 = 𝐾𝑎 𝐾𝑝
(15)
The spherical tank alone exhibits first-order dynamics. The inclusion of actuator dynamics results in a second-order transfer function, which more accurately represents the practical system behavior and is commonly used for controller design and analysis. 2.1.
Laboratory Setup of Spherical Tank System
The laboratory setup of the system consists of a spherical tank, pump, variable frequency drive (VFD), hydrostatic pressure transmitter, reservoir tank, manual inlet/outlet control valve, interfacing module, and personal computer (PC) as shown in Figures 2 and 3. The hydrostatic pressure transmitter acts as a feedback measurement device to measure the liquid level in the tank, and it converts the level into a 4-20 mA signal. The output is interfaced through the NI DQA 6001 card to the computer using a USB, where it converts it into a voltage in the range of 1-6 volts. The control program was written in Simulink using MATLAB 2013b software. Based on the controller’s gain and the VFD’s frequency, the error signal proceeds to initiate instantaneous control action that creates the necessary water flow in a tank. Figure 2 depicts the spherical tank’s real-time experimental setup with the regulated level. The specifications for the real-time process requirements are listed in Table 1 and are shown below: The System Identification Toolbox in MATLAB provides a comprehensive set of tools for developing dynamic system models using measured input–output data. The procedure usually starts with collecting appropriate input and output signals that capture the system’s behavior under different operating conditions. After data collection, the next step involves preprocessing to eliminate noise, outliers, and unwanted trends that could interfere with accurate system identification. Once the data are refined, an appropriate model structure such as a secondorder transfer function is chosen based on existing system knowledge or exploratory data analysis. 28
Figure 3. Schematic of interfacing of NI DQA USB 6001 module with spherical tank system The MATLAB System Identification Toolbox then applies estimation algorithms, such as least-squares and prediction error minimization methods, to determine the optimal parameters for the selected model. The model’s accuracy is assessed through validation techniques, typically by comparing its predicted output with a separate set of measured data. If the results are unsatisfactory, the process can be repeated by modifying the model structure, estimation method, or preprocessing approach. The final identified transfer function can then be utilized for system analysis, controller design, and simulation tasks. Identifying accurate transfer functions is essential in many engineering fields, as it provides a mathematical model that captures the dynamic characteristics of a system. Such models are vital for designing robust control strategies, forecasting system behavior, and enhancing overall system performance. Figure 4 illustrates the model identification and validation results. The y-axis represents the liquid level in terms of a voltage signal, ranging from 1 V to 6 V, corresponding to a tank liquid level of 0 cm to 22 cm.
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 1. Technical specification of experimental setup Sr.No 1 2 3 4 5 6 7 8
Components Spherical Tank DAQ Device Level Transmitter Dosing Pump VFD Manual outlet valve Input flow rate Output flow rate
Specification SS, H-25 cm, D-25cm, R-12.5 cm NI USB6001, 14 bit, AI-4, AO-2 Two wires, Range 0-250mm, O/P=4 to 20 mA PDDP with adjustable stroke 1-ph 200 VAC, 1.1 A, O/P AC 3-ph, 0-230v 𝐶𝑑 = 19.69 𝑐𝑚2 /𝑠 𝑓𝑖𝑛 = 0 (Min) and 420 (Max) 𝐿𝑃𝐻 𝑓𝑜𝑢𝑡 = It is kept 150 𝐿𝑃𝐻 for study
The system identification process produced a second-order transfer function model that achieved a 94.81% fit with the estimation data. The resulting transfer function is expressed as follows:
controller’s capabilities are constrained. Because of its low cost, it is utilized in industrial processes most frequently. The general control structure of the PID controller is given in Figure 5. Many nonlinear processes are controlled with an industrially proven PID controller. A PID controller may manage several intricate and nonlinear processes and offers physical tweaking. Equation 17 represents the commonly used structure of PID controller. 𝑢(𝑡) = 𝑘𝑐 �𝑒(𝑡) +
1 𝑑𝑒(𝑡) � 𝑒(𝑡) 𝑑𝑡 + 𝑇𝑑 � 𝑇𝑖 𝑑𝑡
(17)
In this representation, the proportional gain is 𝑘𝑝 = 𝑘𝑐 , the integral time is 𝑇𝑖 and 𝑘𝑖 = 𝑘𝑐 /𝑇𝑖 and e(t) is the error value. 3.2.
Figure 4. Open loop response of spherical tank
𝐺(𝑠) =
0.02627𝑠 + 0.0001742 𝑠 2 + 0.064𝑠 + 0.0001549
(16)
This mathematical model (16) is used to design the PID controller.
3. Controller Design This section covers the design and tuning of various controllers, including GA-PID, WOA-PID, and WOAGA-PID controllers. The tuning techniques aim to enhance system stability and performance. A comparison of these methods is provided to evaluate their effectiveness. The analysis includes key performance metrics such as settling time, overshoot, and steadystate error. 3.1.
General PID Controller Structure
PID controllers should be configured to a level where there are no faults between process variables and set points to ensure rapid response. When the system’s parameters change or are unknown, the PID
Whale Optimization Algorithm
The Whale Optimization Algorithm is a metaheuristic optimization algorithm inspired by humpback whales’ distinctive hunting tactic, bubble-net feeding [20, 21]. This sophisticated foraging behavior involves whales creating distinctive bubble formations to encircle and disorient their prey, primarily krill and small fish, before moving towards the concentrated food source [22, 23]. The algorithm emulates this hunting technique through a series of mathematical models that represent the key phases of the bubble-net feeding process, offering a novel approach to solving complex optimization problems across various domains. The Whale Optimization Algorithm’s mathematical formulation is based on the modeling of three primary behaviors: seeking for prey, bubble-net attacking, and encircling prey [24]. As shown in equation 18, an ITAE is employed as the fitness function, 𝑇
ITAE = � 𝑡 |𝑒(𝑡)| 𝑑𝑡
(18)
0
Enclosing Prey: The WOA implies that the target prey is the current best-to-date solution since the prey in the search space is unknown. The other whale will then attempt to reposition themselves in relation to the intended victim. When ∣ 𝑋 ∣> 1, whales search globally instead of attacking. Distance calculation Figure 5. Structure of PID controller
𝑍 = | 𝑌 ⋅ 𝑊rand (𝑡) − 𝑊(𝑡) |
(19)
29
Journal of Automation, Mobile Robotics and Intelligent Systems
Position update 𝑊(𝑡 + 1) = 𝑊rand (𝑡) − 𝑋 ⋅ 𝑍
(20)
where, 𝑋, 𝑌=Vectors which controls exploration and exploitation. 𝑊(𝑡)=Current position of whale, 𝑊rand (𝑡) Best solution (Position of Prey). and 𝑡 is current iteration 𝑋 = 2 ⋅ 𝛼 ⋅ (rand) − 𝛼 𝑌 = 2 ⋅ (rand) where, 𝑋 and 𝑌 are coefficient vectors, 𝛼 decreases linearly from 2 to 0 and (rand) is random vector in [0,1]. The selection of this operator depends on the coefficient vector 𝑋 and a random number 𝑟𝑎𝑛𝑑, where 𝑟𝑎𝑛𝑑 is a random number in [0, 1]. For each individual, if 𝑟𝑎𝑛𝑑 < 0.5 and |𝑋| < 1, the position is updated by encircling prey. Attacking Phase: If 𝑟𝑎𝑛𝑑 > 0.5, the bubble-net predation stage begins. The mathematical model of the search is as follows. 𝑊(𝑡 + 1) = 𝐿′ ⋅ 𝑒 𝑐𝑙 ⋅ cos(2𝜋𝑙) + 𝑊rand (𝑡)
(21)
Where, 𝐿′ =|𝑊rand (𝑡) − 𝑊(𝑡)|, ‘𝑙’ is a random value between -1 and 1, ‘𝑐’ is the constant which denotes the shape of the logarithmic spiral.
VOLUME 20,
N∘ 3
Step 1: Establish the starting settings, including Popsize, MaxGen, lower bound, upper bound, and number of variables. Next, a starting population is chosen at random. Step 2: Calculate the fitness value of each individual using Equation (18) and find best solution 𝐾𝑝 , 𝐾𝑖 , 𝐾𝑑 . Step 3: Update the populations by WOA. If 𝑟𝑎𝑛𝑑 < 0.5 and |𝑋| < 1, individuals are updated by encircling prey using Equations (19)–(20). If 𝑟𝑎𝑛𝑑 < 0.5 and |𝑋| ≥ 1, individuals are updated by searching for prey using Equation (22). If 𝑟𝑎𝑛𝑑 > 0.5, individuals are updated by bubble-net attacking using Equation (21). Step 4: Use GA to update the population the related operators make use of mutation, crossover, and the roulette wheel selection approach. Step 5: Calculate the fitness value of each individual and update best solution 𝐾𝑝 , 𝐾𝑖 , 𝐾𝑑 . Step 6: If t < MaxGen, then t = t + 1, update a, 𝑋, 𝑌, 𝑐, 𝑙 and 𝑟𝑎𝑛𝑑, and return to step 3. If t = MaxGen, then the best fitness and global-best 𝐾𝑝 , 𝐾𝑖 , 𝐾𝑑 are obtained. The main optimization parameters for GA, WOA, and the hybrid WOA–GA, such as limits, population size, and number of generations, are compiled in the Table 2. Crossover and mutation for GA and the control parameter 𝛼 dropping from 2 to 0 for WOA and WOA–GA are highlighted. The Table 3 displays differences in 𝐾𝑝 , 𝐾𝑖 , and 𝐾𝑑 resulting from various search strategies and compares PID controller gains acquired using GA, WOA, and hybrid WOA–GA optimization.
Searching for Prey: If 𝑟𝑎𝑛𝑑 < 0.5 and |𝑋| ≥ 1, the individuals need to expand their exploration areas. This phase uses a randomly generated set of solutions 𝑊rand (𝑡) to update the individuals’ locations, which allows the WOA algorithm to perform a global search. The rules of the update solution are as follows. 𝑊(𝑡 + 1) = 𝑊rand (𝑡) − 𝑋|𝑌 ⋅ 𝑊rand (𝑡) − 𝑊(𝑡)| (22) 3.3.
Genetic Algorithm
The GA is a random global search and optimization technique that was created by mimicking the natural biological evolution process. During the search process, it may automatically gather information about the search space, and adaptively manage the search procedure to find the optimal answer. Every member of the population is a potential solution to the optimization issue. The fitness function evaluates an individual’s capacity for adaptability. Those with high levels of fitness are able to procreate while those with low levels of fitness are removed. To create a new population during reproduction, selection, crossover, and variation are required. Finally, a better answer will be found [25]. 3.4.
Hybrid WOA-GA Algorithm
Figure 6 depicts the flowchart of the proposed WOA-GA method, and Figure 7 shows the block diagram of WOAGA-PID controller. Figure 6. WOA-GA algorithm process flowchart
30
2026
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 2. Optimization Parameters Parameter Lower bound [𝐾𝑝 , 𝐾𝑖 , 𝐾𝑑 ] Upper bound [𝐾𝑝 , 𝐾𝑖 , 𝐾𝑑 ] No. of Generations No. of Population Crossover Fraction Mutation Fraction 𝛼 decrease
GA [1, 0.01, -20] [5, 3, 5] 20.0 20.0 0.8 0.2 –
WOA [1, 0.01, -20] [5, 3, 5] 20.0 20.0 – – 2 to 0
WOA-GA [1, 0.01, -20] [5, 3, 5] 20.0 20.0 0.8 0.2 2 to 0
The optimizer discovered a derivative action that mitigates to counteract excessive damping introduced by the plant as indicated by the negative 𝐾𝑑 in GA-PID. Figure 8 compares the convergence characteristic curve of level control of spherical tank obtained by GA, WOA and WOAGA for a population size of 20. The convergence curves confirm that the control method presented based on GA, WOA and WOAGA reaches a sensible degree of performance in controlling level. It is clearly noticeable that WOAGA converges to a small error, compared to GA and WOA.
4. Results and Discussion This section presents the simulation and realtime results of liquid level control in a spherical tank considering its nonlinear mathematical model. The performance of each control strategy was verified through simulation and real-time for different scenarios where the control variable was affected. For a quantitative assessment of the results of each controller in terms of response quality and error value, which allows for comparing them with one another, several Performance Indices were calculated. A MATLAB R2023b/SIMULINK software is used in simulation and real-time analysis to evaluate the output performance of the liquid level of a spherical tank under various controller configurations. The tank’s liquid level range (0–22 cm) is used to control each region’s level.
Figure 7. Block diagram of proposed WOA-GA based PID controller Table 3. Controller tuning values Sr.No. 1 2 3
Controller Type GA-PID WOA-PID WOA-GA-PID
Kp 1.15 2.56 1.92
Ki 0.03 0.17 0.2
Kd -34.34 1.31 -7.27
Figure 8. Convergence curves of GA, WOA and WOAGA with population size 20
4.1.
Simulation Results
The performance of the designed controllers are tested to regulate the liquid level in a spherical tank. All controllers are designed using the obtained mathematical model of the spherical tank. The desired set values of the level is 19.30 cm, and the simulation is run for 1000 seconds. The closed-loop responses for the designed controllers are shown in Figure 9 (a) and (b) depict the simulation output responses and control efforts taken by all designed controllers, respectively. The response of the controllers are recorded in terms of rise time, peak time, settling time, and overshoot percentage, also the performance indices ISE, IAE, and ITAE are captured, which are shown in Tables 4 and 5, respectively. Table 4 allows for a more precise analysis of the behavior of the control strategies by verifying their transient response. The simulation results demonstrate that the WOAGA optimized PID controller outperforms other tuning methods, achieving the shortest rise time of 193 seconds, the fastest peak time 302 seconds, minimal overshoot of 1.25%, and fastest settling time of 251 seconds. While GA-PID and WOA-PID exhibit comparable rise times of 254 and 237 seconds, respectively, WOAGA-PID’s superior dynamic response highlights its effectiveness in balancing speed and stability. Notably, GA-PID suffers from excessive overshoot of 3.21%, emphasizing challenges in conventional tuning for complex systems. WOAGA-based methods consistently reduce overshoot compared to GA and slitely more than WOA approaches, validating their robustness in optimizing 31
Journal of Automation, Mobile Robotics and Intelligent Systems
(a) Simulation responses
VOLUME 20,
N∘ 3
2026
(b) Simulation control efforts
(c) Real-time responses
(d) Real-time control efforts
Figure 9. Simulation and real-time control efforts and responses
Figure 10. Setpoint tracking Table 4. Time domain comparison between simulation and real-time implementation Set Point
Specification
19.30 cm Level
Rise Time Peak Time Settling Time Overshoot %
GA-PID 254 504 670 3.21
Simulation WOA-PID WOAGA-PID 237 193 374 302 307 251 1.04 1.25
control performance across multiple time-domain criteria. Table 5 reports the performance indices calculated for this scenario in a time of 1000 seconds. The novel WOAGA-optimized PID controller achieves the lowest ISE (518), IAE (300), and ITAE (2604), confirming its superior error minimization and transient response compared to GA-PID and WOA-PID methods.
32
GA-PID 305 512 680 2.78
Real-Time WOA-PID WOAGA-PID 275 210 398 347 325 274 1.56 1.67
It is also observed WOAGA-PID also outperforms than GA-PID and WOA-PID controllers in IAE and ITAE, emphasizing WOAGA desiged PID efficacy in reducing cumulative errors. GA-PID exhibits the poorest performance ITAE (7004), likely due to oscillatory behavior, while WOA-PID shows balanced but suboptimal metrics.
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 5. Comparison of performance parameters in simulation and real-time implementation Set Point
Controller
19.30 cm Level
GA-PID WOA-PID WOAGA-PID
Simulation ISE IAE ITAE 624 462 7004 614 361 3732 518 300 2604
ISE 648 657 638
Real-Time IAE ITAE 492 8422 412 4757 398 3228
Table 6. Comparison of control effort in simulation and real-time Norm norm 2 norm ∞
GA-PID 143.79 6
Simulation WOA-PID WOAGA-PID 141.34 139.51 5.6 5.6
GA-PID 145.18 5.7
Real-Time WOA-PID WOAGA-PID 143.71 142 5.7 5.2
Figure 11. Disturbnce rejection These results reinforce WOAGA-optimized controllers as robust solutions for precision and stability in level control systems. The Figure 10 depicts the setpoint tracking performance of all designed controllers and observed that WOAGA-PID outperforming than GA-PID and WOA-PID. GA-PID response is sluggish to achieve the changed setpoint. This comparison shows that the WOAGA-PID controllers superior robustness. 4.2.
Real-time implementation results
To prove the efficacy of designed controllers, the simulation results are validated through real-time implementation on a laboratory set-up of a nonlinear spherical tank system. The closed loop performance of the designed GA-PID, WOA-PID, and hybrid WOAGA optimized-PID controllers are presented to support the claim. Figure 9 (c) and (d) depicts the real-time responses and their control efforts for reference tracking performance for the setpoint of 19.30 cm level, respectively. From this it’s observed that the WOAGAPID controller is performing best among the rest of the controllers. The hybrid WOAGA-PID controller shows better performance in a real-time environment, achieving the fastest rise time of 210 seconds, the lowest settling time of 274 seconds, and minimal overshoot of 1.67% as given in Table 4, while giving optimal error metrics (ISE: 638, IAE: 398, ITAE: 3228),
outperforming all counterparts. WOA-PID shows moderate overshoot 1.56% and reduced ITAE (4757) compared to GA-PID (ITAE: 8422) reported in Table 5. Notably, WOAGA-PID’s balanced performance combining rapid response, stability, and precision validates WOAGA adaptability to real world disturbances. While WOAGA and WOA methods exhibit trade-offs GA-PID’s high sluggishness, WOAGA optimized control consistently bridges speed and accuracy, proving its robustness in practical implementations. from analysis it is observed that the values recoded in real time are slightly more than simulation, because of inherent hardware limitations and dynamics. Another hardware limitation is that it does not work for control signals more than six volts. Its working range is one to six volts, which converts it into 0 to 50 Hz to control the speed of the pump. Based on the data in Table 6, the hybrid WOAGA-PID controller consistently exhibits the lowest control effort in both simulation and real-time implementation, indicating improved energy efficiency compared to GA-PID and WOA-PID controllers. Moreover, the reduced norms in realtime operation demonstrate the robustness and practical effectiveness of the proposed hybrid approach under experimental conditions. The WOAGA-PID controller demonstrates superior disturbance rejection by exhibiting the smallest deviation from the setpoint 33
Journal of Automation, Mobile Robotics and Intelligent Systems
and the fastest recovery following the disturbance. In contrast, GA-PID and WOA-PID show larger transient dips and slower settling, indicating reduced robustness under disturbance conditions as depicted in Figure 11.
5. Conclusion This paper presents the successful design and implementation of an optimal PID level control strategy for a nonlinear spherical tank system using a hybrid WOA–GA. The proposed hybrid approach effectively combines the fast exploitation capability of WOA with the global search strength of GA to achieve reliable PID tuning for nonlinear process dynamics. Comparative simulation studies demonstrate that the WOAGA–PID controller outperforms WOA–PID and GA–PID controllers in terms of transient response, robustness, and standard performance indices. The effectiveness of the proposed method is further confirmed through real-time experimental validation on a laboratory-scale spherical tank, showing excellent set-point tracking and disturbance rejection under varying operating conditions. Overall, the results establish the proposed hybrid WOA–GA–based PID controller as an effective and practical solution for nonlinear level control applications, with strong potential for extension to other complex industrial processes. Future research will extend this approach to adaptive or predictive control frameworks with comprehensive uncertainty, noise, and large-scale experimental validation.
AUTHORS Abhay Pakhare∗ – Department of Instrumentation Engineering, Ramrao Adik Institute of Technology, D. Y. Patil Deemed to be University, Navi Mumbai, India, e-mail: abhay.pakhare@rait.ac.in. Sharad Jadhav – Department of Instrumentation Engineering, Ramrao Adik Institute of Technology, D. Y. Patil Deemed to be University, Navi Mumbai, India, e-mail: sharad.jadhav@rait.ac.in. Mukesh Patil – Department of Instrumentation Engineering, Ramrao Adik Institute of Technology, D. Y. Patil Deemed to be University, Navi Mumbai, India, e-mail: mukesh.patil@rait.ac.in. ∗
Corresponding author
References [1] C. Urrea et al., “Design and Performance Analysis of Level Control Strategies in a Nonlinear Spherical Tank”, Processes, vol. 11, no. 3, 2023, PP. 720. [2] C. Urrea et al., “Design and Comparison of Strategies for Level Control in a Nonlinear Tank”, Processes, vol. 9, no. 5, 2021, PP. 735. [3] S. Yu et al., “Liquid Level Tracking Control of Three-Tank Systems”, International Journal of Control, Automation and Systems, vol. 18, no. 10, 2020, PP. 2630–2640. 34
VOLUME 20,
N∘ 3
2026
[4] C. Sreepradha et al., “Synthesis of Fuzzy Sliding Mode Controller for Liquid Level Control in Spherical Tank”, Cogent Engineering, vol. 3, no. 1, 2016, PP. 1222042. [5] C. Priya et al., “Particle Swarm Optimisation Applied to Real Time Control of Spherical Tank System”, International Journal of Bio-Inspired Computation, vol. 4, no. 4, 2012, PP. 206–216. [6] K. Sundaravadivu et al., “Design of Fractional Order PID Controller for Liquid Level Control of Spherical Tank”. In: 2011 IEEE International Conference on Control System, Computing and Engineering, 2011, PP. 291–295. [7] R. Arivalahan et al., “Liquid Level Control in Two Tanks Spherical Interacting System with Fractional Order Proportional Integral Derivative Controller using Hybrid Technique: A Hybrid Technique”, Advances in Engineering Software, vol. 175, 2023, PP. 103316. [8] A. Ashwini et al., “Quadruple Spherical Tank Systems with Automatic Level Control Applications using Fuzzy Deep Neural Sliding Mode FOPID Controller”, Journal of Engineering Research, vol. 13, no. 1, 2025, PP. 68–83. [9] P. Kamalakkannan et al., “Optimising Liquid Level in a Spherical Two Tank Interacting System with Fractional Order PID Control via Hybrid POA-RERNN Approach”, International Journal of Heavy Vehicle Systems, vol. 32, no. 2, 2025, PP. 163–188. [10] R. Anandanatarajan et al., “Limitations of a PI Controller for a First-Order Nonlinear Process with Dead Time”, ISA Transactions, vol. 45, no. 2, 2006, PP. 185–199. [11] N. Ajlouni et al., “Enhancing PID Control Robustness in CSTRs: A Hybrid Approach to Tuning under External Disturbances with GA, PSO, and Machine Learning”, Neural Computing and Applications, vol. 37, no. 18, 2025, PP. 12153–12177. [12] R. P. Borase et al., “A Review of PID Control, Tuning Methods and Applications”, International Journal of Dynamics and Control, vol. 9, no. 2, 2021, PP. 818–827. [13] P. Mohindru et al., “Review on PID, Fuzzy and Hybrid Fuzzy PID Controllers for Controlling Non-Linear Dynamic Behaviour of Chemical Plants”, Artificial Intelligence Review, vol. 57, no. 4, 2024, PP. 97. [14] M. Kowsalya and K. Sundaravadivu, “Design and Implementation of Non-Linear System Using Gain Scheduled pi Controller”. In: Procedia Engineering, vol. 38, 2012, PP.2568–2575. [15] J. Garicano-Mena et al., “Nature-inspired Metaheuristic Optimization for Control Tuning of
Journal of Automation, Mobile Robotics and Intelligent Systems
Complex Systems”, Biomimetics, vol. 10, no. 1, 2024, PP. 13. [16] S. W. Kareem et al., “Metaheuristic Algorithms in Optimization and Its Application: A Review”, JAREE (Journal on Advanced Research in Electrical Engineering), vol. 6, no. 1, 2022. [17] O. Kozlov et al., “Swarm Optimization of the Drone’s Intelligent Control System: Comparative Analysis of Hybrid Techniques”. In: Proceedings of the 12th International Conference Information Control Systems & Technologies (ICST 2024). Available at: https://ceur-ws. org, vol. 3790, 2024. [18] W.-Y. Wang et al., “An Online GA-Based OutputFeedback Direct Adaptive Fuzzy-Neural Controller for Uncertain Nonlinear Systems”, IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), vol. 34, no. 1, 2004, PP. 334–345. [19] C. Kadu et al., “ITAE Based Robust Controller for Spherical Tank Level Control System”, Engineering Research Express, vol. 7, no. 3, 2025, PP. 035408.
VOLUME 20,
N∘ 3
2026
[21] M. A. Thanoon et al., “Boost Converter Control using Proportional-Integral-Derivative Controller Optimized by Whale Optimization Algorithm”, International Journal of Robotics & Control Systems, vol. 5, no. 3, 2025. [22] A. Hassan et al., “Tuning PID Controller for Inverted Pendulum using Whale Optimization Algorithm”, Nile Journal of Engineering and Applied Science, vol. 2, no. 2, 2025, PP. 0–0. [23] H. Zhang et al., “Fractional-order PID Control of Two-Wheeled Self-Balancing Robots via MultiStrategy Beluga Whale Optimization”, Fractal and Fractional, vol. 9, no. 10, 2025, PP. 619. [24] S. Mirjalili et al., “The Whale Optimization Algorithm”, Advances in Engineering Software, vol. 95, 2016, PP. 51–67. [25] D. Pradeepkannan et al., “Control of a Nonlinear Spherical Tank Process using GA Tuned PID Controller”, International Journal of Innovative Research in Science, Engineering and Technology, vol. 3, no. 3, 2014, PP. 580–586.
[20] R. K. Mahto et al., “Controller Design using PSO and WOA Algorithm for Enhanced Performance of Vector-Controlled PMSM Drive”, International Journal of Modelling and Simulation, vol. 45, no. 5, 2025, PP. 1841–1851.
35
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
COMPARATIVE STUDY OF LQR AND NEURAL NETWORK CONTROLLERS FOR QUADCOPTER ROLL ANGLE STABILIZATION Submitted: 4th September 2025; accepted: 2nd February 2026
Hiba Arif, Marouane Kadi, Aymane Bidah, Omar Zakary DOI: 10.14313/jamris-2026-036 Abstract: This paper presents a comparative analysis between a conventional Linear Quadratic Regulator (LQR) and a neural network-based approach for quadcopter roll angle stabilization. While the LQR controller delivers optimal performance in nominal conditions, its accuracy deteriorates when subjected to model uncertainties and persistent external disturbances, causing steady-state deviations. To mitigate these issues, we develop a neural network controller trained via imitation learning. The proposed architecture inherently learns an integral action, enabling perfect rejection of constant disturbances. Extensive simulations show that our neural controller achieves superior performance in realistic operating conditions while maintaining similar stability margins to LQR in ideal cases. This study demonstrates how machine learning techniques can enhance classical control paradigms for aerial robotics. The results suggest that neural network controllers offer a viable path toward more adaptive and robust autonomous drone systems in challenging environments. Keywords: Quadrotor control, Linear Quadratic Regulation, Deep learning, UAV stabilization, Robust control, Behavioral cloning
1. Introduction Modern unmanned aerial systems, particularly quadrotor platforms, have seen exponential growth in applications across surveillance, precision agriculture, and disaster management [5]. These versatile systems nevertheless present significant control challenges due to their inherent nonlinear dynamics, coupled states, and open-loop instability [7]. Effective disturbance rejection and precise attitude maintenance under environmental uncertainties remain critical requirements for their operational reliability. The quest for optimal control strategies has led to extensive research into various methodologies. While conventional PID controllers remain popular for their simplicity, modern optimal control approaches like the Linear Quadratic Regulator (LQR) have gained prominence for multivariable systems [3]. LQR’s theoretical foundation in state-space optimization makes it particularly suitable for aerial vehicle control [1]. 36
By minimizing a quadratic cost function that balances state regulation against control effort, LQR generates optimal feedback gains that ensure system stability while considering multiple state variables simultaneously [2]. Despite its mathematical elegance, LQR exhibits several practical limitations [4]. Its performance assumptions - perfect system modeling, linear dynamics, and Gaussian noise - rarely hold in real-world flight conditions. These constraints manifest particularly in persistent steady-state errors and inadequate disturbance rejection, compromising the quadcopter’s positioning accuracy [16]. The controller’s static gain structure further limits its adaptability to varying operational conditions. Recent advances in machine learning have introduced new paradigms for control system design. Neural networks, with their universal approximation capabilities, offer compelling alternatives by learning control policies directly from operational data [8]. This data-driven approach circumvents the need for precise system modeling while maintaining robustness against uncertainties. Hybrid architectures combining LQR with neural components (LQR-RNN) have demonstrated particular promise, merging classical control theory with adaptive learning [9]. Such systems leverage LQR’s optimization framework while employing recurrent networks (RNNs) to dynamically adjust control parameters based on real-time performance [11]. This work presents a comparative evaluation of classical LQR and neural network-based controllers for quadrotor roll stabilization. Focusing on the fundamental roll dynamics described by states 𝑥1 = 𝜙(𝑡) ̇ and 𝑥2 = 𝜙(𝑡) [6], we examine each controller’s performance under modeled disturbances and parametric uncertainties. Our simulation-based analysis highlights the relative strengths of each approach, with particular emphasis on the neural controller’s potential advantages in realistic operational scenarios.
2.
Theoretical Background
Modern technological progress in quadrotor UAV systems has transformed numerous industrial and commercial domains, including cinematography, delivery services, and precision farming [5]. At the core of these applications lies the critical need for reliable flight control architectures that ensure
Open Access. © 2026 Hiba Arif et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
Journal of Automation, Mobile Robotics and Intelligent Systems
operational safety and performance consistency [6]. Among various control challenges, maintaining precise attitude regulation under environmental disturbances remains a fundamental research problem in aerial robotics [7]. Roll angle dynamics (𝜙), representing the rotational motion about the longitudinal axis, constitute a fundamental component of quadrotor stabilization [16]. Conventional control methodologies, particularly linear approaches, remain popular due to their well-established theoretical framework. The Linear Quadratic Regulator stands out in this category, offering an optimal control solution through quadratic optimization of state deviations and control inputs [1]. However, this optimality is strictly contingent upon accurate system modeling and the absence of unaccounted disturbances - conditions seldom met in practical implementations [4]. Recent developments in artificial intelligence have introduced novel control paradigms through machine learning techniques. Neural networkbased controllers present distinct advantages by learning control policies directly from operational data rather than relying on explicit system models [9]. This characteristic enables them to capture complex nonlinear relationships and adapt to various uncertainties, including parameter variations (e.g., changes in moment of inertia) and external perturbations (such as wind gusts) [17]. Reinforcement learning frameworks further enhance this adaptability through continuous performance optimization [10].
VOLUME 20,
4.
N∘ 3
2026
System Dynamics Formulation
In developing our control framework, we adopt a simplified dynamic model that captures the essential roll dynamics (𝜙) of the quadrotor system. This singleaxis approximation, frequently employed in control literature [5], provides several advantages: - Isolation of a fundamental stabilization challenge - Reduction of system complexity while preserving key dynamic characteristics - Enables clearer comparative analysis of control approaches 4.1.
Roll Dynamics Formulation
The rotational motion of a quadrotor around its longitudinal axis follows the fundamental principles of rigid body dynamics, specifically Euler’s rotation equation derived from Newton’s second law for rotational systems [16]. When considering only the x-axis rotation, we obtain the governing differential equation: ̈ 𝐼𝑥𝑥 𝜙(𝑡) = 𝜏𝑥 (𝑡) + 𝛿(𝑡)
(1)
where the physical quantities are defined as: - 𝜙(𝑡): Roll angle (rad) - the primary controlled output variable ̈ - 𝜙(𝑡): Angular acceleration (rad/s2 ) - second time derivative of roll angle - 𝐼𝑥𝑥 : Moment of inertia about x-axis (kg⋅m2 ) - characterizes the vehicle’s rotational inertia [7]
Our research conducts a systematic performance comparison between conventional LQR and neural network-based controllers for roll angle stabilization. The study demonstrates that while LQR achieves theoretical optimality under ideal conditions, its performance deteriorates significantly when facing persistent disturbances and modeling inaccuracies. In contrast, properly trained neural controllers exhibit superior robustness and disturbance rejection capabilities, as evidenced through extensive simulation scenarios [12].
- 𝜏𝑥 (𝑡) (or 𝑢(𝑡)): Control torque input (N⋅m) - generated by differential thrust from rotors - 𝛿(𝑡): Disturbance torque (N⋅m) - aggregates various unmodeled effects including: - Aerodynamic disturbances (wind gusts, ground effects)
The paper is organized as follows: Section 2 develops the system’s mathematical model. Section 3 elaborates on both control architectures (LQR and neural network). Section 4 presents comparative simulation results and analysis, followed by concluding remarks on the findings’ significance.
4.2.
3. System Modeling For the purpose of control system design and analysis, we consider a simplified representation of the quadrotor dynamics by focusing exclusively on its rotational behavior about the longitudinal axis (roll angle 𝜙). This common simplification in aerial vehicle control studies enables focused examination of a crucial degree of freedom while maintaining analytical tractability [6].
- System imperfections (motor imbalances, structural asymmetries) [17] - Measurement noise and modeling inaccuracies State-Space Formulation
Modern control theory extensively employs statespace representations as they provide a comprehensive framework for analyzing and designing control systems using linear algebra techniques [3]. For our second-order roll dynamics system, we define the state vector x(𝑡) to completely characterize the system’s dynamic behavior: 𝜙(𝑡) 𝑥 (𝑡) x(𝑡) = � 1 � = � ̇ � 𝑥2 (𝑡) 𝜙(𝑡)
(2)
where 𝑥1 (𝑡) represents the roll angle and 𝑥2 (𝑡) corresponds to the angular velocity. Through differentiation and substitution of the governing dynamics (1), we derive the following coupled first-order differential equations: 37
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
- Diagonal elements determine relative importance of different states 𝑥̇ 1 (𝑡) = 𝑥2 (𝑡) 1 1 𝑢(𝑡) + 𝛿(𝑡) 𝑥̇ 2 (𝑡) = 𝐼𝑥𝑥 𝐼𝑥𝑥
(3) (4)
These equations can be compactly expressed in matrix form as: x(𝑡) ̇ = Ax(𝑡) + B𝑢(𝑡) + E𝛿(𝑡)
(5)
with the system matrices defined as:
A=�
0 0
1 �, 0
0 B = � 1 �, 𝐼𝑥𝑥
- Prevents actuator saturation in practical implementations For our roll stabilization problem, we employ diagonal weighting matrices: 𝑞 Q = � 11 0
𝐼𝑥𝑥
5. Control System Architectures Building upon the established dynamic model, this section develops two distinct control methodologies for quadrotor roll stabilization. The first represents a well-established model-based optimal control technique (LQR), while the second employs a contemporary data-driven paradigm using artificial neural networks. - Linear Quadratic Regulator (LQR): A classical optimal control approach that requires precise system modeling - Neural Network Controller: A modern machine learning technique that learns control policies from operational data This comparative study examines both approaches’ theoretical foundations and implementation strategies, highlighting their respective advantages and limitations for aerial vehicle stabilization.
(6)
0
The weighting matrices in the cost function serve distinct purposes: - Q (positive semi-definite): Governs state regulation performance
R = [𝑟]
(7)
- Performance degrades with nonlinear effects (e.g., aerodynamic drag, motor saturation) [16] - Local linearization may not capture full operational envelope behavior 2) Model Dependency: - Optimality requires exact knowledge of system parameters (𝐼𝑥𝑥 in our case) - Inaccurate inertia estimates lead to suboptimal gain selection [3] - Sensitivity to unmodeled dynamics and parameter variations 3) Disturbance Rejection Deficiency: - Pure state feedback lacks inherent integral action - Persistent disturbances cause steady-state offsets [1] - Limited robustness to constant torque disturbances These structural limitations, particularly the inability to achieve perfect disturbance rejection, necessitate the development of more robust control paradigms that can maintain performance under realistic operating conditions. 5.2.
∞
𝐽 = � �x𝑇 (𝑡)Qx(𝑡) + 𝑢𝑇 (𝑡)R𝑢(𝑡)� 𝑑𝑡
0 �, 𝑞22
where 𝑞11 and 𝑞22 respectively weight roll angle ̇ errors. The optimal gain (𝜙) and angular velocity (𝜙) matrix K is derived through solution of the Algebraic Riccati Equation (ARE), a fundamental result in optimal control theory [3]. Inherent Limitations of LQR Control: While the LQR controller offers theoretical optimality for linear systems, its practical implementation faces several fundamental constraints [4]: 1) Linearity Assumption: - Designed exclusively for linear system dynamics
LQR Control Methodology
Theoretical Foundation: The Linear Quadratic Regulator represents a cornerstone of optimal control theory for linear time-invariant systems [2]. This approach computes an optimal state feedback control law 𝑢(𝑡) = −Kx(𝑡) by minimizing a quadratic performance index:
38
- R (positive definite): Controls input energy expenditure - Higher values prioritize control efficiency over response speed [4]
0 E=� 1 �
The state matrix A captures the inherent system dynamics, while the input matrix B describes how the control torque 𝑢(𝑡) affects the state evolution. The disturbance matrix E quantifies the impact of external perturbations on the system. This formulation provides the mathematical foundation for both conventional and neural network-based controller designs [1].
5.1.
- Larger values enforce faster convergence to equilibrium [1]
Neural Network-Based Control Architecture
Conceptual Framework: Motivated by the identified limitations of model-based LQR control, we develop an alternative approach leveraging artificial neural networks (ANNs). This data-driven methodology aims to: - Learn optimal control policies directly from operational experience rather than analytical models
Journal of Automation, Mobile Robotics and Intelligent Systems
- Capture complex nonlinear relationships inherent in real-world quadrotor dynamics - Automatically adapt to system uncertainties and disturbances The proposed architecture employs supervised learning to approximate an enhanced control law that combines the strengths of conventional state feedback with learned disturbance rejection capabilities. Unlike the LQR’s fixed-gain structure, the neural controller develops adaptable mappings between system states and control actions through exposure to diverse operational scenarios. Expert Controller Design and Behavioral Cloning: Training Strategy: Our approach employs imitation learning, requiring an expert controller capable of optimal performance across diverse operational scenarios. This expert must overcome the LQR’s key limitation by achieving perfect disturbance rejection through integral action. LQI Controller Architecture: We implement a Linear Quadratic Integral (LQI) controller as our expert system. The state vector is augmented with an integral term: 𝜙(𝑡) x(𝑡) ̇ � xaug (𝑡) = � � = � 𝜙(𝑡) 𝜉(𝑡) 𝑡 ∫0 𝜙(𝜏)𝑑𝜏
(8)
The extended system dynamics become: ẋaug (𝑡) = Aaug xaug (𝑡) + Baug 𝑢(𝑡)
(9)
where: 0 Aaug = �0 1
1 0 0
0 0� , 0
0 Baug = �1/𝐼𝑥𝑥 � 0
(10)
∞
𝐽aug = � �x𝑇aug (𝑡)Qaug xaug (𝑡) + 𝑢𝑇 (𝑡)Raug 𝑢(𝑡)� 𝑑𝑡 (11)
The solution yields an optimal gain matrix: Kaug = �K𝑝
𝐾𝑖 � = �𝑘1
𝑘2
𝑘3 �
(12)
(13)
Data Generation Process: The training dataset is generated through extensive simulations that systematically vary: - Moment of inertia (𝐼𝑥𝑥 ) - Initial conditions - Disturbance profiles
2026
For each scenario, the expert controller computes optimal control actions, creating state-action pairs (xaug , 𝑢expert ). The neural network learns to approximate the mapping 𝑓 ∶ xaug → 𝑢expert across the entire operational envelope. Network Architecture and Training Methodology: The proposed controller employs a multilayer perceptron (MLP) structure, designed according to established deep learning principles [8]: - Input Layer: - Receives the 3-dimensional augmented state vector (𝜙, 𝜙,̇ ∫ 𝜙𝑑𝑡) - Normalizes inputs to ensure stable training dynamics - Hidden Layers: - Two fully-connected layers with 64 units each - ReLU activation functions providing necessary nonlinear transformations [10] - Batch normalization between improved training stability
layers
for
- Output Layer: - Single linear neuron generating torque commands 𝑢(𝑡) - No activation function to allow unbounded control outputs The training process minimizes the mean squared error (MSE) between predicted and expert control actions using the Adam optimization algorithm [9]. Key training characteristics include: - Large-scale exposure to diverse operational scenarios - Iterative weight updates through backpropagation
- Early stopping based on validation performance Through this data-driven approach, the network develops an adaptive control policy that generalizes across various operating conditions and disturbance profiles, demonstrating superior robustness compared to conventional LQR methods [12].
6.
implementing the control law: 𝑢expert (𝑡) = −K𝑝 x(𝑡) − 𝐾𝑖 𝜉(𝑡)
N∘ 3
- Learning rate scheduling for convergence refinement
Optimal Control Formulation: The LQI controller minimizes an augmented cost function:
0
VOLUME 20,
Performance Evaluation and Comparative Analysis
To rigorously assess the effectiveness of both control strategies, we implemented a comprehensive simulation study under demanding operational conditions. This test scenario was specifically designed to: - Expose the inherent vulnerabilities of conventional LQR control - Validate the enhanced robustness of our neural network-based approach - Provide quantitative comparison metrics under identical test conditions
39
Journal of Automation, Mobile Robotics and Intelligent Systems
The experimental framework incorporates multiple challenging aspects including parameter uncertainties, persistent disturbances, and nonlinear effects, thereby creating a realistic evaluation environment for both control paradigms. 6.1.
Experimental Configuration
The comparative evaluation employs the following test conditions to rigorously assess controller performance: - Parameter Uncertainty: - LQR design based on nominal inertia 𝐼𝑥,nominal = 0.1 kg⋅m2 - Actual system inertia 50% higher (𝐼𝑥,actual = 0.15 kg⋅m2 )
- The control input converges to −0.5 N⋅m, exactly counterbalancing the disturbance but requiring a non-zero angular deviation [1] - This behavior confirms the theoretical inability of pure state feedback to completely reject constant disturbances Neural Network Controller Performance The ANNbased controller (green solid lines) demonstrates notable advantages: - Transient Phase: - More aggressive initial response with maximum deviation limited to 0.044 rad - Higher control authority utilization (peaking at −0.65 N⋅m)
- External Disturbances: - Initial equilibrium condition (𝜙 = 0, 𝜙̇ = 0)
- Faster error correction dynamics [10] - Steady-State Behavior: - Complete disturbance by 𝑡 = 4s
- Models sustained aerodynamic forces (e.g., crosswinds) [17] The evaluation compares two distinct control architectures: - Conventional LQR (2-state feedback) - ANN-based controller (3-state feedback with integral action) [6] Performance Evaluation
The comparative analysis of both control strategies under the specified test conditions yields the following observations. Disturbance Rejection Capabilities Figure 1 demonstrates the controllers’ behavior when simultaneously subjected to parameter uncertainty and external disturbances. Figure 1: Perf)rma(ce C)mparis)(: LQR vs. Augmen.e ANN A(gle Response Un er Constant Disturbance
Roll Angle ϕ (rad)
0.04
rejection
achieved
- Final control input matches the disturbance magnitude (−0.5 N⋅m) - Perfect setpoint recovery demonstrating learned integral action [12] An initial transient artifact visible before 𝑡 = 1s warrants further investigation in subsequent analysis (see Figure 2) [8]. Regulation Performance in Nominal Conditions Figure 2 examines the controllers’ stabilization capabilities from a non-zero initial condition (𝜙(0) = 0.5 rad) without external disturbances [6]. Figu)e 2: Cont)olle) Pe) o)mance in Ideal Conditions (Stabili−ation )om Initial Angle) Angle Response
0.5
Roll Angle ϕ (rad)
- Step disturbance 𝑑(𝑡) = 0.5 N⋅m applied at 𝑡 = 1s
0.05
2026
- Upon disturbance application at 𝑡 = 1s, the system develops a persistent steady-state error of 0.048 rad
- Represents realistic payload variations (e.g., equipment changes) [15]
6.2.
N∘ 3
VOLUME 20,
LQR Cont)olle) Augmented ANN Cont)olle)
0.4 0.3 0.2 0.1 0.0
−0.1
Cont)ol E o)t
0.03 0
0.01
−1
Cont)ol To)(ue u(t)
0.02
−2
0.00
LQR w/ Mismatch & Dist/rbance A/gmente ANN Contro&&er
−0.01
−3 −4 −5
C)(tr)l Eff)rt LQR C)(tr)l Eff)rt Augme(te ANN C)(tr)l Eff)rt
0.0
Con.rol Torque u(t)
−0.1
−6
LQR Cont)ol E o)t Augmented ANN Cont)ol E o)t
−7 0
−0.2 −0.3
2
4
Time (s)
6
8
10
Figure 2. Comparative regulation performance in disturbance-free conditions
−0.4 −0.5 −0.6 0
2
4
Time (s)
6
8
10
Figure 1. Comparative response of LQR and neural network controllers under model mismatch and constant disturbance conditions LQR Controller Performance The conventional LQR approach (shown in red dashed lines) exhibits characteristic limitations predicted by control theory [3]: 40
Reference LQR Performance The LQR controller (red dashed line) demonstrates its theoretical optimality for linear systems [2]: - Critically damped response with no overshoot - Minimum-time convergence given the system constraints - Smooth control input trajectory
Journal of Automation, Mobile Robotics and Intelligent Systems
Neural Network Controller Behavior The ANN controller (green solid line) exhibits remarkable characteristics: - Near-perfect emulation of optimal LQR response [9] - Identical convergence rate and stability properties - Superimposed state and control trajectories [10] Key observations: - The transient artifact present in Figure 1 disappears completely - Confirms the network’s ability to switch between: - Disturbance rejection mode (Figure 1) - Optimal regulation mode (Figure 2) [8] - Demonstrates contextual adaptation based on system conditions [1] Robustness Evaluation Across Parameter Variations To systematically assess controller performance under diverse operating conditions, we performed a parametric study examining: - Inertia variations: 𝐼𝑥 ∈ [0.10, 0.25] kg ⋅ m2 (representing ±50% nominal value)
N∘ 3
VOLUME 20,
2026
- Nominal Operation (𝑑 = 0 N⋅m): - Comparable performance for both controllers at 𝐼𝑥 = 0.15 kg⋅m2 [6] - Identical settling time of 2.0 seconds - Moderate Disturbance (𝑑 = 0.5 N⋅m): - ANN maintains robust stability across full inertia range [12] - LQR exhibits: - Significant 0.06 rad overshoot at nominal inertia [1] - Persistent 0.06 rad steady-state error for increased inertia [4] Parameter Sensitivity Analysis - Neural Network Controller: - Minimal 4% variation in settling time across tested inertia range [8] - Consistent disturbance elimination within 5 second window [10] - LQR Controller Limitations: - Severe 210% settling time degradation at maximum inertia [3] - Emerging instability (oscillations) for: - Disturbances ≥ 0.8 N⋅m [16]
- Disturbance range: 𝑑 ∈ [0.0, 0.8] N ⋅ m (covering mild to severe perturbations) Figure 3 presents the comprehensive evaluation matrix, with: - Rows corresponding to different inertia values
- Inertia values ≥ 0.20 kg⋅m2 [2] Comparative Performance Assessment
- Columns representing increasing disturbance magnitudes
Table 1. Statistical comparison of control performance across experimental trials Evaluation Criterion Settling Time (2% threshold) Peak Overshoot Magnitude Steady-State Precision Control Energy Expenditure
- ANN responses shown in solid green trajectories - LQR baselines displayed as red dashed curves Performance of ANN Contro er (s LQR o(er Different Inertia and Disturbance Va ues d = 0.0
0.3
d = 0.2
d = 0.5
d = 0.8
I_x = 0.10 Ang e (rad)
LQR ANN
0.2 0.1
Neural Network 3.2 ± 0.4 s
Conventional LQR 5.8 ± 2.1 s
12 ± 3 %
34 ± 18 %
0.00 ± 0.00 rad
0.07 ± 0.05 rad
0.45 ± 0.08 Nm
0.52 ± 0.12 Nm
0.0
Neural Network Architecture Benefits The enhanced control performance of the artificial neural network originates from three fundamental design characteristics: Extended State Formulation The controller architecture incorporates an integrated error state, providing the network with complete state information [9]:
I_) = 0.12 Ang e (rad)
0.3 0.2 0.1 0.0
I_) = 0.15 Ang e (rad)
0.3 0.2 0.1 0.0
I_) = 0.18 Ang e (rad)
0.3 0.2 0.1 0.0
𝑡
̇ 𝑢𝑁𝑁 (𝑡) = 𝒩 �𝜙(𝑡), 𝜙(𝑡),
I_) = 0.20 Ang e (rad)
0.3 0.2 0.1
� 𝜙(𝜏)𝑑𝜏 � 0 �������
(14)
error accumulation
0.0 0
2
4
6
8
10
0
2
4
6
8
10
0
2
4
6
8
10
0
2
4
6
8
10
Figure 3. Parametric performance comparison showing ANN (green) vs LQR (red) responses across inertia-disturbance combinations. The grid organization enables direct comparison of robustness to parameter variations and disturbance rejection capabilities Performance Insights Disturbance Rejection Characteristics
where 𝒩(⋅) represents the neural network mapping function. Adaptive Nonlinear Mapping The multilayer perceptron structure automatically learns to compensate for system nonlinearities through [11]: 𝑢adapt (𝐼𝑥 , 𝑑) = Γ (𝐼𝑥 , 𝑑(𝑡)) 𝜕Γ satisfying < 0, 𝜕𝐼𝑥
𝜕Γ >0 𝜕𝑑
(15)
41
Journal of Automation, Mobile Robotics and Intelligent Systems
demonstrating proper sensitivity to parameter variations and disturbances. Broad Operational Generalization The training methodology employs comprehensive parameter variations [17]: - Moment of inertia range: 0.08 to 0.25 kg⋅m2 - Disturbance magnitude: -1.5 to +1.5 N⋅m ensuring reliable performance across unencountered operational scenarios [15]. Disturbance Rejection Performance Evaluation To quantitatively assess the neural controller’s disturbance rejection capabilities, we conducted systematic tests across multiple disturbance levels while maintaining constant inertia (𝐼𝑥 = 0.15 kg⋅m2 ). The evaluation considered four distinct disturbance magnitudes: 𝑑 ∈ {0.0, 0.2, 0.5, 0.8} N⋅m. Figure 4 presents the superimposed temporal responses [6]. Neural Network Co troller Respo se u der Var(i g Disturba ces 0.30
Disturba ce = 0.0 Disturba ce = 0.2 Disturba ce = 0.5 Disturba ce = 0.8 Refere ce
Roll A gle (rad)
0.25 0.20
VOLUME 20,
N∘ 3
2026
Table 2. Neural controller performance across disturbance levels Performance Metric Settling Time (s) Max Deviation (rad) Steady-State (rad) Control Effort (Nm)
0.0 Nm 2.25 0.00 0.000 0.12
0.2 Nm 2.85 0.015 0.000 0.28
0.5 Nm 3.01 0.025 0.000 0.53
0.8 Nm 3.40 0.040 0.000 0.81
- Disturbance Rejection Phase (1-3s): Active detection and compensation of external perturbations [8] - Steady-State Regulation Phase (>3s): Precise maintenance of desired setpoint The controller dynamically modulates its effective gain parameters [10]: K𝑒𝑓𝑓 = �1.2 ± 0.3
0.4 ± 0.1
0.8 ± 0.2�
(17)
with observed integral gain enhancement proportional to disturbance magnitude [4].
0.15 0.10 0.05 0.00
−0.05 0
2
4
Time (s)
6
8
10
Figure 4. Neural network controller performance under increasing disturbance levels. The superimposed responses demonstrate consistent stability and tracking precision across all tested conditions Performance Evaluation Regulation Characteristics - Undisturbed Operation (blue curve): - Optimal transient response with 2.25 second settling time [1]
Figure 5. ANN controller stabilization duration versus disturbance intensity for fixed inertia (𝐼𝑥 = 0.15 kg⋅m2 ). The results demonstrate consistent performance degradation with increasing disturbance while maintaining guaranteed stability
- No observable overshoot - Minimal RMS tracking error (0.002 rad) - Moderate Disturbance (red curve, 𝑑 = 0.5 N⋅m): - Peak deviation of 0.04 rad occurring at 1.5 seconds [17] - Complete disturbance rejection within 3.01 seconds [12] - Zero steady-state error Adaptive Control Behavior The neural controller exhibits intelligent compensation through [9]: ̇ 𝐾𝑝 𝜙(𝑡) + 𝐾𝑑 𝜙(𝑡) 𝑢(𝑡) = � ̇ 𝐾𝑝 𝜙(𝑡) + 𝐾𝑑 𝜙(𝑡) + 𝐾𝑖 ∫ 𝜙𝑑𝑡 + 𝑢𝑐𝑜𝑚𝑝 (𝑑)
pre-disturbance post-disturbance (16)
where 𝑢𝑐𝑜𝑚𝑝 (𝑑) represents the automatically learned disturbance compensation component. Quantitative Performance Metrics Neural Controller Operational Phases Analysis of the neural network’s control strategy reveals three distinct operational regimes [11]: - Initial Stabilization Phase (0-1s): Pure state regulation from non-zero initial conditions 42
Figure 5 quantifies the relationship between disturbance magnitude and stabilization duration. The neural controller exhibits a predictable monotonic increase in response time with higher disturbance amplitudes while preserving complete stability. Specific stabilization durations measure 2.25 seconds for the undisturbed case (𝑑 = 0.0), progressing to 2.85 seconds (𝑑 = 0.2), 3.01 seconds (𝑑 = 0.5), and 3.40 seconds (𝑑 = 0.8), verifying the controller’s adaptive capability under varying operational conditions. The detailed comparison of simulation outcomes enables us to derive significant insights regarding the fundamental characteristics and performance capabilities of both control methodologies [16]. The Limitations of a Rigid Model. The experimental results obtained with the LQR controller, while theoretically sound, clearly reveal the constraints inherent to model-dependent control strategies [1]. The controller’s optimal performance proves highly sensitive to precise system modeling, as its effectiveness diminishes when faced with deviations from its design assumptions [4]. The observed performance
Journal of Automation, Mobile Robotics and Intelligent Systems
reduction in the presence of unmodeled disturbances represents an intrinsic property of its mathematical formulation rather than an implementation deficiency [2]. Consequently, the LQR approach remains primarily applicable to well-characterized systems operating in controlled environments [3]. The Learned Intelligence of the ANN Controller. The superior performance of the ANN-based controller stems not simply from its universal approximation capacity [8], but rather from its demonstrated ability to develop an adaptive and context-sensitive control policy, enabled by two critical architectural features [10]: 1) Learning integral action: The incorporation of the integral state ∫ 𝜙𝑑𝑡 proved instrumental [9]. This augmentation provided the network with temporal information about error accumulation. Through training, the ANN autonomously discovered the fundamental relationship between growing integral values and persistent disturbances, recognizing that even small but sustained angle deviations require compensatory control actions [12]. Remarkably, the network developed integral control functionality without explicit programming of this control law [11], as clearly evidenced in Figure 1. 2) Behavior versatility: The contrasting behaviors observed between Figures 1 and 2 provide compelling evidence of the controller’s sophisticated learning capability [15]. The network has effectively mastered two distinct operational modes [6]: - Under disturbance-free conditions (Figure 2), with minimal integral term activity, it functions comparably to an optimal PD controller, matching LQR performance [1]. - When disturbances are present (Figure 1), triggering integral term growth, it automatically engages its learned integral compensation mechanism to eliminate steady-state error [17]. This context-dependent behavior, acquired purely through data-driven learning, represents a fundamental advancement over conventional LQR and demonstrates superior suitability for real-world applications [8]. In conclusion, this research illustrates how moving beyond simple “black box” approaches through thoughtful architectural design and comprehensive training can yield significant benefits [9]. By combining established control theory principles (state augmentation) with modern machine learning techniques, we have developed a controller that exhibits not only robustness but also behavioral versatility and operational transparency [12].
7. Conclusion This research has systematically evaluated and compared the effectiveness of conventional LQR control with a neural network-based approach for quadrotor roll stabilization under challenging operational conditions [10]. Through extensive simulation
VOLUME 20,
N∘ 3
2026
studies, we have obtained definitive and insightful answers to our original research questions [1]. Our results conclusively demonstrate that while the LQR controller achieves theoretical optimality in ideal scenarios, it exhibits fundamental limitations when confronted with real-world operational challenges [4]. The controller’s structural inability to completely reject persistent disturbances, manifesting as steady-state errors, combined with its sensitivity to parameter inaccuracies, significantly restricts its practical utility in applications demanding robust performance [6]. In contrast, the carefully designed neural network controller has demonstrated exceptional performance characteristics [8]. By incorporating an augmented state representation that includes error integration and through rigorous model selection from multiple training iterations, the ANN controller has achieved dual capabilities: matching LQR’s efficiency in nominal conditions while developing superior disturbance rejection strategies [12]. The controller’s ability to autonomously recognize persistent disturbance patterns and adaptively modulate its control effort to achieve perfect error cancellation represents an emergent integral-like behavior that was learned purely from data [9]. This adaptive capability, emerging naturally from the learning process, strongly validates the potential of such approaches for advanced control applications [11]. This study naturally leads to several promising directions for future investigation [15]: - Physical implementation and validation: Transitioning from simulation to real-world deployment by testing the ANN controller on actual quadrotor hardware would provide crucial validation of its practical effectiveness - Full attitude control expansion: Generalizing the current single-axis approach to comprehensive three-dimensional attitude control would represent a significant step toward fully learning-based flight control systems - Advanced disturbance modeling: Investigating the controller’s performance against more complex, time-varying disturbances such as turbulent wind fields would further explore the boundaries of its robustness [17] - Alternative learning paradigms: Exploring reinforcement learning methodologies could potentially eliminate the need for expert demonstrations while maintaining or enhancing controller performance In summary, this work provides compelling evidence for the transformative potential of neural networks in modern control system design [9]. The successful integration of classical control theory principles with powerful deep learning techniques has yielded an intelligent, adaptive control solution capable of meeting the stringent demands of contemporary robotic applications [16].
43
Journal of Automation, Mobile Robotics and Intelligent Systems
AUTHORS Hiba Arif∗ – Faculty of Sciences Ben M’Sick, Laboratory of Analysis, Modeling and Simulation, Hassan II University of Casablanca, Morocco, e-mail: arifhiba1@gmail.com. Marouane Kadi – Faculty of Sciences Ben M’Sick, Laboratory of Analysis, Modeling and Simulation, Hassan II University of Casablanca, Morocco, e-mail: Kadi.marouan@gmail.com. Aymane Bidah – Faculty of Sciences Ben M’Sick, Laboratory of Analysis, Modeling and Simulation, Hassan II University of Casablanca, Morocco, e-mail: Aymanebidah2@gmail.com. Omar Zakary – Faculty of Sciences Ben M’Sick, Laboratory of Analysis, Modeling and Simulation, Hassan II University of Casablanca, Morocco, e-mail: zakaryma@gmail.com. ∗
Corresponding author
References [1] B. D. O. Anderson and J. B. Moore, Optimal Control: Linear Quadratic Methods. Courier, 2007. [2] A. E. Bryson and Y. C. Ho, Applied Optimal Control. Taylor & Francis, 1975. [3] G. F. Franklin, J. D. Powell, and A. Emami-Naeini, Feedback Control of Dynamic Systems, 7th ed. Pearson, 2015. [4] K. Zhou, J. C. Doyle, and K. Glover, Robust and Optimal Control. Prentice Hall, 1996. [5] G. M. Hoffmann, H. Huang, S. L. Waslander, and C. J. Tomlin, “Quadrotor helicopter flight dynamics,” in Proc. AIAA Guidance, Navigation and Control Conference, 2007, pp. 1–20. [6] D. Mellinger and V. Kumar, “Minimum snap trajectory generation and control for quadrotors,” in Proc. IEEE Int. Conf. Robot. Autom., 2011, pp. 2520–2525. [7] R. Mahony, V. Kumar, and P. Corke, “Multirotor aerial vehicles,” IEEE Robot. Autom. Mag., vol. 19, no. 3, pp. 20–32, 2012. [8] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436–444, 2015. [9] A. Y. Ng and S. J. Russell, “Algorithms for inverse reinforcement learning,” in Proc. Int. Conf. Machine Learning, 2000, pp. 663–670. [10] J. Schulman, S. Levine, P. Moritz, and M. I. Jordan, “Trust region policy optimization,” in Proc. Int. Conf. Machine Learning, 2015, pp. 1889–1897. [11] H. Schaub, J. L. Junkins, and R. D. Robinett, “RNN attitude control,” J. Guid. Control Dyn., vol. 21, no. 3, pp. 464–468, 1998. [12] P. Abbeel, A. Coates, M. Quigley, and A. Y. Ng, “An application of reinforcement learning to aerobatic helicopter flight,” Int. J. Robot. Res., vol. 29, no. 13, pp. 1608–1639, 2010. 44
VOLUME 20,
N∘ 3
2026
[13] T. Zhang, G. Kahn, S. Levine, and P. Abbeel, “Learning deep control policies for autonomous aerial vehicles with MPC-guided policy search,” in Proc. IEEE Int. Conf. Robot. Autom., 2016, pp. 528–535. [14] S. Levine, C. Finn, T. Darrell, and P. Abbeel, “Endto-end training of deep visuomotor policies,” J. Mach. Learn. Res., vol. 17, no. 1, pp. 1334–1373, 2016. [15] G. Loianno, C. Brunner, G. McGrath, and V. Kumar, “Estimation, control, and planning for aggressive flight with a small quadrotor,” IEEE Robot. Autom. Lett., vol. 1, no. 2, pp. 404–411, 2016. [16] H. K. Khalil, Nonlinear Systems, 3rd ed. Prentice Hall, 2015. [17] S. L. Waslander and C. Wang, “Wind disturbance estimation and rejection for quadrotors,” in Proc. AIAA Infotech@Aerospace Conference, 2009, pp. 1–12. [18] S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proc. Int. Conf. Artificial Intelligence and Statistics, 2011, pp. 627–635.
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
A HYBRID LSTM-DNN MODEL WITH FUZZY INFERENCE FOR ADAPTIVE DISPATCHER CONTROL OF INDUSTRIAL INFORMATION-CONTROL SYSTEMS Submitted: 20th November 2025; accepted: 2nd February 2026
Barnokhon Temerbekova, Gulnora Bekimbetova, Ulugbek Mamanazarov, Bakhodir Bekimbetov DOI: 10.14313/jamris-2026-037 Abstract: Erroneous and noisy measurements significantly distort the assessment of the stability of technological lines in industrial information and control systems. This paper proposes an algorithm for operational dispatch control of technological complexes based on the assessment of noise immunity using a hybrid LSTM-DNN model and a fuzzy set apparatus. The hybrid model predicts time dependencies and failure probabilities, while the fuzzy inference system combines objective sensor data and subjective expert assessments into a single reliability indicator. The center of gravity method’s use of trapezoidal membership functions and defuzzification ensures a smooth mapping between linguistic and numerical variables. The study then carries out an analysis of the conditional monotonic behavior of the interference immunity coefficient with respect to system performance under the adopted modeling assumptions. Modelling has shown that, compared to probabilistic methods, the proposed neuro-fuzzy approach can improve the adaptability and stability of a system under non-stationary disturbances. The model developed here can potentially be integrated into industrial dispatch control systems, subject to further validation. Keywords: operational dispatch control, information and control systems, interference immunity assessment, fuzzy logic/fuzzy set apparatus, hybrid LSTM-DNN neural network model, adaptive control, intelligent control systems, stability of technological processes
1. Introduction Industrial information and control systems (ICS) operate under conditions of constantly changing modes and a high degree of uncertainty in measurement information. The resulting errors and noise in the signals significantly distort the calculated stability indicators of technological lines and lead to inadequate control decisions [1]. The greatest risks arise in the control center, where not only is an accurate assessment of the current state required, but also a timely system response to external and internal disturbances. The methods of interference immunity analysis used in practice are usually based on probabilistic models and equipment failure statistics. However, such approaches are only effective with complete and reliable information, which is rarely achievable in real
production conditions. In addition, they do not take into account subjective factors — in particular, expert observations and intuitive assessments by operators based on practical experience. This limits the system’s ability to adapt to non-stationary situations and reduces the reliability of control. Modern trends in the development of intelligent control systems involve the use of machine learning and fuzzy logic methods to combine objective and subjective data in a single analytical circuit [2]. Hybrid neural network architectures, in particular LSTMDNN models, allow for identifying temporal dependencies between technological process parameters and predicting the probability of deviations [3]. The mathematical apparatus of fuzzy set theory provides processing of incomplete and contradictory information, including the expert judgements and heuristic knowledge of specialists. This paper presents an algorithm for operational and dispatch control of a technological complex based on the assessment of the noise immunity of an information and control system using a hybrid LSTM-DNN neural network and elements of fuzzy logic. The aim of this study is to form an integrated neuro-fuzzy model that ensures the stability and adaptation of the system when exposed to disturbances and noisy data. It provides a formal analysis of the dependence of the interference immunity coefficient on production characteristics; describes the structure of the rule base; and demonstrates the possibility of practical integration of the developed algorithm into real-time industrial systems. The present study is focused on the development and formal analysis of a model- and simulation-based algorithm for assessing the interference immunity of industrial information and control systems. The proposed approach is intended to provide a rigorously defined indicator and an integration framework combining neural network forecasting with fuzzy inference, rather than a fully validated industrial deployment. All results reported in this paper are obtained under explicitly-stated modeling assumptions and simulation scenarios, which define the scope and limitations of applicability of the proposed method. The main contribution of this work lies in the integration pattern, which combines a hybrid LSTM-DNN risk prediction module with a Mamdani-type fuzzy inference system to form an interference immunity
Open Access. © 2026 Barnokhon Temerbekova et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
45
Journal of Automation, Mobile Robotics and Intelligent Systems
indicator suitable for dispatch-level decision support. The contribution is methodological and model-based, focusing on structured integration and interpretability rather than on deployment-specific optimization. Unlike classical ANFIS-based or tightly coupled neuro-fuzzy systems, the proposed architecture follows a modular integration paradigm. The LSTMDNN block is used exclusively for temporal risk prediction, while the fuzzy inference system operates as an independent decision-level fusion layer. This separation preserves interpretability, enables flexible rule base design, and allows explicit incorporation of expert knowledge, distinguishing the proposed approach from end-to-end neuro-fuzzy learning schemes commonly reported in the literature.
2. Theoretical foundations Assessing the noise immunity of information and control systems is one of the central tasks in the design and operation of complex technological complexes [4]. Most existing approaches are based on probabilistic analysis methods that use statistical data from equipment failures and external disturbances [5]. Such methods yield integral reliability indicators, but require a large number of reliable observations, which limits their applicability in conditions of incomplete or noisy measurements. In recent years, research related to the use of artificial intelligence and adaptive algorithms has been developing. Machine learning methods, in particular recurrent and deep neural networks, demonstrate high efficiency in forecasting non-stationary time series. Architectures such as Long Short-Term Memory (LSTM) and Deep Neural Network (DNN) make it possible to take into account long-term dependencies and nonlinear relationships between the parameters of technological processes [6]. However, neural network models are “black box” in nature, and do not provide sufficient interpretability when making management decisions. Another class of approaches is based on the application of fuzzy set theory and fuzzy logic. Unlike probabilistic models, these allow for the formalization of subjective expert assessments and the use of nonnumerical information. Fuzzy inference systems effectively describe the uncertainty and partial reliability of data, which is especially important when analyzing measurements that are subject to noise and hardware failures. However, when used independently, fuzzy logic does not take into account dynamic dependencies and is unable to predict the behavior of the system over time. In this regard, hybrid neuro-fuzzy models, which combine the learning and prediction capabilities of neural networks with the interpretability of fuzzy systems, are of particular interest. Such a synthesis provides the ability to adapt to changes in technological parameters and to explain control decisions. The use of hybrid architectures combining LSTM-DNN and Fuzzy Inference System (FIS) blocks allows for the formation of stability criteria that take into account 46
VOLUME 20,
N∘ 3
2026
both objective measurements and subjective expert knowledge [7]. The generalized architecture of a hybrid neurofuzzy information and control system is shown in Fig. 1. This article uses this approach as the basis for developing an algorithm for assessing interference immunity and adaptive operational control of technological complexes [8].
3.
Methodology
3.1.
Definitions and notation
In this study, all variables and operators are defined explicitly to ensure formal consistency and reproducibility. Let x∈[0, 100] denote the normalized load level of a technological line, expressed on a unified percentage scale; let R∈[0, 100] represent the predicted risk value obtained from the LSTMDNN neural network model; and let 𝑅𝐸𝑋𝑃 ∈ [0, 100] denote the expert risk assessment provided by the dispatcher. The interference immunity indicator 𝑍𝑎 ∈ [0, 100] is defined as the defuzzified output of a Mamdani-type fuzzy inference system (FIS), constructed on the basis of the input variables x, 𝑟,̂ and 𝑅𝐸𝑋𝑃 . The fuzzy inference process includes fuzzification using trapezoidal membership functions; rule evaluation using min-max aggregation operators; and defuzzification performed by the centroid (center-of-gravity) method. All variables are mapped to a common numerical range [0,100] through predefined normalization procedures introduced at the model design level. This range represents a unified scale for fuzzy inference compatibility and ensures commensurability of heterogeneous inputs without implying physical equivalence of their original units. 3.2.
Model Concept
The system under development belongs to the class of intelligent ICS and is designed for adaptive control of a technological complex in the presence of noise, incomplete data and uncertainty [9]. The algorithm combines two complementary approaches: - Neural network forecasting of temporal dependencies of parameters based on a hybrid LSTMDNN architecture. - Fuzzy-logical assessment of noise immunity, allowing for expert judgements and subjective characteristics of the system state to be taken into account. At the upper level, the LSTM subsystem analyzes historical measurements of technological parameters, forming a forecast of the probability of failures and the risk coefficient 𝑅𝐿𝑆𝑇𝑀 . At the lower level, the Fuzzy Inference System (FIS) combines the results of neural network forecasting, current sensor data, and expert assessments by the dispatcher 𝑅𝐸𝑋𝑃 to form an integral interference immunity indicator 𝑍𝑎 . The resulting value is used to correct operational tasks and select control actions in real time [10].
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Supervisory Level SCADA / DCS / MES
confirmation / approval of actions
recommendations, explanations
Fuzzy Logic Layer • Rule base (Gaussian MFs) • Inference + defuzzification • Explanations / recommendations
modes / manual override, alarms / telemetry
control actions u(t) PLC / Control Loop SIMATIC / etc. • Controllers • Interlocks
e(t)=y-ŷ
Sensor Layer Process measurement signals y(t)
Neural Network Layer LSTM-DNN • Prediction ŷ(t+Δ) • Confidence ρ
y(t) (data flow)
predictions / confidence metrics raw/cleaned time series
Historian / Data Lake Archiving, ETL, data preparation
model weight updates (CI/CD) ŷ(t+Δ), ρ
datasets, features, metrics
adaptation based on Z_a
rule / membership boundary adjustment based on Z_a
physical feedback → new y(t)
setpoints, e(t), events
update of rules / MF parameters
Actuation Level Valves, pumps, drives
measurements y(t)
u(t)
ŷ(t+Δ), ρ
decisions / explanations
Training / MLOps Offline/online retraining Validation, model/rules deployment
y(t), σ² of channels
update of Z_a thresholds Adaptation loop: • When Z_a falls below thresholds: - Sensor re-weighting - Fine-tuning / incremental retraining of LSTM - Adjustment of fuzzy membership functions and rules
Quality & Reliability Assessment Z_a = f(ρ, e, σ², …)
Solid arrows - operational data/control flow Dotted arrows - training / adaptation (offline & online)
Figure 1. Architecture of a hybrid neuro-fuzzy IUS with a feedback adaptation loop Table 1. LSTM-DNN architecture and training parameters Component LSTM layers Hidden units Activation Optimizer Loss function Training epochs Validation
over activated rules, is the aggregation operation; 𝑋𝑎 (𝑄) is a set of parameters determining the current state of the lines; and 𝜇𝑒 (𝑥) is a membership function characterizing the degree of stability according to parameter x. In this work, 𝑍𝑎 is interpreted as the defuzzified output of a Mamdani-type fuzzy inference system constructed for the inputs x, R, and 𝑅𝐸𝑋𝑃 .
Description 1-2 stacked layers e.g., 32–64 tanh (LSTM), ReLU (DNN) Adam MSE e.g., 100 sliding window
𝑛
To ensure reproducibility of the simulation-based experiments, the main architectural and training parameters of the LSTM-DNN model used in this study are summarized in Table 1. 3.3.
Formal Model for Assessing Interference Immunity
𝜇𝑒 (𝑥) =
⎧ � 𝜇 (𝑥 ), 𝑖 𝑖 ⎪ ⎪ 𝑖=1 for the probabilistic interpretation, ⎨ max 𝜇 (𝑥 ), 𝑖 𝑖 ⎪ 𝑖 for the logical-linguistic ⎪ interpretation. ⎩ (2)
The interference immunity indicator Z𝑎 is computed within a Mamdani-type fuzzy inference framework. The monotonic behavior under discussion is assumed to hold only within specific operating regimes characterized by the absence of downstream bottlenecks and within the parameter ranges defined by the adopted fuzzy rule base. Outside these regimes, no monotonic relationship is claimed. Aggregation of fuzzy rules is performed using the standard max-min scheme, and no alternative aggregation operators are considered in this study: 𝑥∈𝑋𝑎 (𝑄)
Based on expressions (1) and (2), an analytical dependence of the interference immunity index on the outputs of lines 𝑄𝑖 is established. Under the assumed structure of the fuzzy rule base and the adopted ordering of membership functions, the interference immunity indicator 𝑍𝑎 exhibits a non-decreasing behavior with respect to increasing line performance in the absence of limiting interconnections. This property is conditional, and holds within the specified operating regime and fuzzy inference configuration considered in this study:
where ⊗ denotes that the Mamdani rule aggregation operator, implemented as the maximum (max)
𝑍𝑎 (𝑄𝑘+1 ) ≥ 𝑍𝑎 (𝑄𝑘 ),
𝑍𝑎 =
⊗ 𝜇𝑒 (𝑥),
(1)
(3) 47
Journal of Automation, Mobile Robotics and Intelligent Systems
Table 2. Comparative analysis of the membership functions for input and output variables Term Parameters trapmf Interpretation Low [0, 0, 0, 40] low load / low risk Medium [30, 45, 55, 70] medium range High [60, 100, 100, 100] high load/risk Output variable: 𝑍𝑎 – degree of interference immunity (0-100). Term trapmf parameters Interpretation Low [0, 0, 0, 50] low system stability medium level Med [25, 40, 60, 75] of protection High [50, 100, 100, 100] high stability
This statement is conditional and is reported as a property of the chosen rule configuration and operating regime, rather than as a general analytical theorem. This reflects the conditional monotonic behavior of the indicator under the adopted fuzzy rule configuration. In particular cases, this dependence can be expressed differentially: 𝑑𝑍𝑎 ≥ 0, 𝑑𝑄𝑖
(4)
which confirms the correctness of the model’s behavior during transient processes and ensures the stability of control decisions when loads change. It has been shown that when the line capacity, which is not limited by other nodes, increases, the noise immunity coefficient increases, and when there are limiting connections between lines, it remains the same or grows more slowly. This property ensures the correct behavior of the model during transient processes and serves as the basis for adaptive control. 3.4.
Fuzzy Evaluation Subsystem
To describe the input and output variables in the fuzzy inference system, trapezoidal membership functions (trapmf) are used, which are defined by four parameters [a, b, c, d], where [b, c] determines the “plateau” of certain membership, and [a, d] defines a smooth transition between states [11, 12]. This choice ensures stability and continuity during fuzzification. The parameters and interpretations of the membership functions used for the input and output variables are summarized in Table 2. Input variables (range 0-100): - 𝑄𝑖 - line load - 𝑅𝐿𝑆𝑇𝑀 - predicted risk based on neural network data - 𝑅𝐸𝑋𝑃 - expert risk assessment according to the dispatcher’s opinion Expert risk scores are treated as calibrated ordinal-to-interval assessments provided within a predefined linguistic and numerical framework. The model assumes a bounded level of consistency for expert judgments within the considered operating period, without claiming invariance across different experts or time horizons. 48
VOLUME 20,
N∘ 3
2026
The fuzzy rule base forms the dependency 𝑍𝑎 = 𝑓(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ) in the form of a set of IF-THEN production expressions [13, 14]. An example of a rule base fragment is shown below:
The fuzzy rules presented here illustrate the general structure of the rule base. In practical applications, the rule base is designed to be extensible, and can be adapted to specific technological contexts, expert knowledge, and operational constraints. The rules are aggregated using the maximum operation, after which defuzzification is performed using the centroid method. The result of 𝑍𝑎 is used in the adaptation block to adjust dispatch decisions. This approach allows statistical forecasts and expert assessments to be combined into a single cognitive-intellectual model of production process stability [15–17]. 3.5.
Experimental Implementation
The experimental study is simulation-based and was implemented in MATLAB R2024a using the Fuzzy Logic Toolbox and Deep Learning Toolbox (Fig. 2). Input signals were generated synthetically to represent typical operating regimes of technological lines, and all input variables were normalized to the range [0,100], as described in Section 3. Measurement noise was introduced as additive stochastic disturbances, with noise levels varying from 0% to 15% of the signal amplitude. For each noise level, the simulation was repeated across multiple runs and the reported results were averaged. The evaluation metrics included the relative estimation error of the interference immunity indicator and the adaptation time under changes in input conditions [18].
4.
Results and Discussion
4.1.
Analytical Dependence and Interpretation of the Model
During numerical modelling, the dependence of the interference immunity index 𝑍𝑎 on the input variables [19] was investigated: line load 𝑄𝑖 , neural network model forecast 𝑅𝐿𝑆𝑇𝑀 , and expert assessment 𝑅𝐸𝑋𝑃 . Based on the fuzzy inference system described in Section 3, an approximation of the response surface 𝑍𝑎 = 𝑓(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ) was constructed. All observations presented in this section are obtained under the simulation assumptions defined in Section 3 and should be interpreted within this modeling framework. The general functional form is represented as: 𝑍𝑎 = 𝐹(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ),
(5)
Where 𝐹(⋅) describes the composition of fuzzification, application of the rule base, and defuzzification by the center-of-gravity method.
Journal of Automation, Mobile Robotics and Intelligent Systems
N∘ 3
VOLUME 20,
2026
MATLAB Platform LSTM-DNN (Deep Learning Toolbox)
Q_i, R_LSTM
FIS module (Fuzzy Logic Toolbox)
Z_a
• compute Z_a • compare/visualize (e.g., curve of Z_a vs. time)
Z_a = f(Q_i, R_LSTM, R_EXP) (visualization / logging)
R_EXP (experimental/ground truth)
Figure 2. Diagram of the model implementation in the MATLAB R2024a environment using the Fuzzy Logic Toolbox and Deep Learning Toolbox packages For individual components, the system allows for partial derivatives [20] that determine the sensitivity of noise immunity to each of the parameters: 𝜕𝑍𝑎 > 0, 𝜕𝑄𝑖
𝜕𝑍𝑎 < 0, 𝜕𝑅𝐿𝑆𝑇𝑀
𝜕𝑍𝑎 < 0, 𝜕𝑅𝐸𝑋𝑃
(6)
which means that an increase in load leads to an increase in stability only up to a certain threshold, after which the increase in risk according to the model or expert assessment reduces the final indicator 𝑍𝑎 . This interpretation is model-dependent and does not represent a general analytical sensitivity result independent of the fuzzy inference structure. Further analysis showed that the combined influence of objective (neural network) and subjective (expert) factors can be approximated by a linear composition with weighting coefficients 𝛼 and 𝛽: 𝑍𝑎 = 𝛼 ⋅ 𝑍𝐿𝑆𝑇𝑀 + 𝛽 ⋅ 𝑍𝐸𝑋𝑃 ,
𝛼 + 𝛽 = 1,
(7)
where 𝑍𝐿𝑆𝑇𝑀 and 𝑍𝐸𝑋𝑃 are the partial outputs of the two branches of the fuzzy system, corresponding to the machine forecast and expert opinion, respectively [20]. The parameter α is adjusted adaptively based on explicit uncertainty estimators for the neural and expert branches: 𝛼=
𝜎𝐸𝑋𝑃 , 𝜎𝐿𝑆𝑇𝑀 + 𝜎𝐸𝑋𝑃
center-of-gravity method [21]. The graph (Figure 3) shows the areas of stability of the technological complex in the input parameter space. The zones of maximum values of 𝑍𝑎 correspond to a combination of moderate line load and low risks according to the neural network and expert data. The minima of the surface occur with a simultaneous increase in 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 and exceeding the threshold load [22]. The surface 𝑍𝑎 = 𝑓(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ) at 𝑅𝐸𝑋𝑃 = 0.5. All input variables are specified in the range [0,100], which corresponds to the percentage scales of load and risks: 𝑞 = 𝑄𝑖 ,
𝑟𝑙 = 𝑅𝐿𝑆𝑇𝑀 ,
where 𝜎𝐿𝑆𝑇𝑀 is computed as the root-meansquare prediction residual of the LSTM-DNN model on a sliding validation window, and 𝜎𝐸𝑋𝑃 is computed as the normalized dispersion (variance) of expert risk scores obtained from repeated or multiple expert assessments. Both estimators are normalized to the same scale prior to their use in (8) to allow commensurability within the adopted modeling framework. The proposed weighting mechanism represents an algorithmic heuristic embedded in the model design, and is not presented as an optimal or theoretically guaranteed adaptation rule. As the uncertainty of the neural network forecast increases, this heuristic results in an empirically observed adjustment of the relative weight of the expert component, facilitating a balance between statistical and cognitive sources of information. The response surface 𝑍𝑎 = 𝑓(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ) is obtained using the Mamdani model, with trapezoidal membership functions and defuzzification using the
𝑞, 𝑟𝑙 ,
𝑟𝑒 ∈ [0, 100].
(9)
Each of them is described by a trapezoidal membership function: 𝜇𝑡𝑟𝑎𝑝 (𝑥; 𝑎, 𝑏, 𝑐, 𝑑) = max �min �
𝑥−𝑎 𝑑−𝑥 , 1, � , 0� 𝑏−𝑎 𝑑−𝑐
𝑥 ∈ [0, 100]. (10)
The activation of each rule j is calculated as the minimum of the membership degrees of all three premises [23]: (𝑡 )
(8)
𝑟𝑒 = 𝑅𝐸𝑋𝑃 ,
(𝑡 )
(𝑡 )
𝐿 𝐸 (𝑟𝑙 ), 𝜇𝐸𝑋𝑃 (𝑟𝑒 ), 𝑎𝑗 = min(𝜇𝑄 𝑄 (𝑞), 𝜇𝐿𝑆𝑇𝑀
(11)
where 𝑡𝑄 , 𝑡𝐿 , 𝑡𝐸 ∈ {Low, Medium, High} . For each output term Low, Medium, High, the maximum of all activated rules is calculated, after which defuzzification is performed using the center of gravity method: 𝑍𝑎 ≈
∑𝑧 𝑧𝜇𝑍 (𝑧) , ∑𝑧 𝑧𝜇𝑍 (𝑧)
𝑧 ∈ [0, 1].
(12)
In a discrete implementation, the calculation is performed on a uniform grid, and if necessary, a zeroorder Sugeno surrogate is applied: 𝑍𝑎 ≈
∑𝑗 𝑎𝑗 𝑐𝑗 ∑𝑗 𝑎𝑗
,
(13)
where 𝑐𝑗 are the centers of the output terms (e.g., 𝑐𝐿 = 0.2, 𝑐𝑀 = 0.5, 𝑐𝐻 = 0.8). The obtained response surface provides smooth transitions between interference immunity levels, and is therefore suitable for qualitative sensitivity analysis and operational interpretation within the simulation framework. In this study, no surrogate regression model was used to support the main claims. 49
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 3. Three-dimensional response surface of the dependence of the degree of interference immunity 𝑍𝑎 on the parameters 𝑄𝑖 (line load) and 𝑅𝐿𝑆𝑇𝑀 (predicted risk) at a fixed expert risk level 𝑅𝐸𝑋𝑃 = 50 (normalized range of input parameters [0,100]) 4.2.
Interpretation of the Surface Shape
When analyzing the shape of the surface 𝑍𝑎 = 𝑓(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ), pronounced non-linearity and asymmetry are observed. In the low value range of 𝑅𝐿𝑆𝑇𝑀 and 𝑅𝐸𝑋𝑃 , the function 𝑍𝑎 increases almost linearly with increasing load 𝑄𝑖 , reflecting a growth in system efficiency with moderate resource use. After reaching a critical level 𝑄𝑖 ≈ 0.6, a stability plateau appears, where a further increase in load does not lead to an increase in noise immunity [23]. When 𝑅𝐿𝑆𝑇𝑀 or 𝑅𝐸𝑋𝑃 increases, the surface slopes, indicating an asymmetric influence from expert and model risk; an increase in expert risk assessment reduces 𝑍𝑎 faster than a similar change in the neural network forecast. This is explained by the greater weight of the subjective component when data uncertainty is high [24, 25]. Thus, the response surface clearly demonstrates the adaptive nature of the model with the transition from the area of stability growth to the area of saturation, as well as further decline when the threshold levels of risk and load are exceeded. 4.3.
Visualization of the Response Surface
The three-dimensional surface 𝑍𝑎 = 𝑓(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ) has a concave shape with local minima at highrisk values and technological line overloads. 50
In the 𝑄𝑖 ∈ [40, 60], 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ∈ [30, 50] zone, there is a stable area where the value of 𝑍𝑎 > 70 corresponds to the optimal operating mode of the complex. As 𝑅𝐿𝑆𝑇𝑀 increases (increased probability of failures predicted by the neural network), the surface 𝑍𝑎 descends along the axis 𝑅𝐿𝑆𝑇𝑀 , maintaining symmetry along the axis 𝑅𝐸𝑋𝑃 . This confirms the consistency of the model — the influence of the two risk factors is equivalent within the acceptable range. Fig. 3 shows the three-dimensional surface of the function 𝑍𝑎 = 𝑓(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ) constructed for the range of input parameters 𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ∈ 100 at a fixed value of 𝑅𝐸𝑋𝑃 = 50. The surface reflects the dependence of the degree of interference immunity on the current load of the production line and the predicted risk determined by the LSTM-DNN neural network model [26]. The surface has a concave shape with local minima at high-risk values and technological line overloads. In the 𝑄𝑖 ∈ [40, 60], 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ∈ [30, 50] region, there is a stable zone where the values of 𝑍𝑎 > 70, which corresponds to the optimal operating mode of the complex. As 𝑅𝐿𝑆𝑇𝑀 (probability of failures predicted by the neural network) increases, the surface descends along the 𝑍𝑎 axis, maintaining symmetry
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 4. Comparison of the average noise immunity estimation error for probabilistic and neuro-fuzzy models at different levels of input data noise along the 𝑅𝐸𝑋𝑃 axis. This confirms the consistency of the model: the influence of neural network and expert risk factors remains equivalent within the acceptable range [27]. 4.4.
Comparison with the Probabilistic Method
To evaluate the effectiveness of the proposed approach, a comparison was made with the classical probabilistic stability model based on failure statistics [28]. In both cases, identical input data were used, but the proposed neuro-fuzzy model showed higher noise immunity; the average estimation error at 15% noise decreased from 12.4% to 4.8%. As can be seen in Fig. 4, with an increase in the noise level, the error of the probabilistic model increases almost linearly, reaching ≈ 12.4% at 15% noise, while the neuro-fuzzy model demonstrates stability, with error not exceeding 4.8%. In addition to the comparison with a classical probabilistic model, it should also be noted that a standalone LSTM-DNN architecture was evaluated during preliminary experiments. While the LSTMDNN model demonstrated improved performance compared to probabilistic approaches, its estimates exhibited higher sensitivity to noise and reduced interpretability in decision-making scenarios. The incorporation of the fuzzy inference layer enables structured integration of expert knowledge and provides additional robustness and transparency, and this cannot be achieved by the neural network alone. This confirms the effectiveness of integrating LSTM-DNN and FIS [29]: joint processing of temporal and expert features provides adaptive noise filtering and accurate estimation 𝑍𝑎 . In addition, the adaptation time to new data decreased by 18% due to the use
of the recurrent component of LSTM [30], which provides a long-term dependence between parameters. Table 2 shows the results of a comparative analysis of two approaches — the probabilistic model and the proposed neuro-fuzzy model — under identical modelling conditions. The indicators include the average error in noise immunity estimation at a 15% noise level; the time required for the system to adapt to new data; resistance to measurement gaps; and the presence of a mechanism for processing expert factors. The data presented confirm that the use of a hybrid LSTM-DNN architecture with a fuzzy logic block provides a comprehensive improvement in key characteristics, including a 61% reduction in error, an 18% reduction in adaptation time, a 20% increase in assessment reliability, and the ability to take subjective expert assessments into account when forming the final stability criteria. Thus, the model provides: - Increased reliability of stability assessment with incomplete data; - Adaptation to changes in the statistical characteristics of signals; and - Integration of subjective factors into a single stability criterion.
5.
Discussion
The results confirm that the proposed approach combines the advantages of neural network and fuzzy logic methods. The use of the LSTM-DNN component enables the prediction of dynamic parameter deviations, while the fuzzy set apparatus provides interpretability and facilitates the incorporation of expert knowledge. 51
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 3. Comparative analysis of the characteristics of probabilistic and neuro-fuzzy models Indicator Average error at 15% noise, % Adaptation time, s Resistance to data omission, % of reliable estimates Processing of expert factors
The present study is based on simulation experiments conducted in a controlled modeling environment. One limitation of the proposed approach, however, is that no validation using real industrial process data is included in the current work. Unlike probabilistic methods, which require complete statistical information, the presented algorithm maintains stability and reliability even with high data uncertainty. This makes it applicable to industrial automated operational dispatch control systems (IAODCs), where both response speed and cognitive interpretation of system states are important. 5.1.
Limitations and Directions for Further Research
The experimental part of the study was modelbased and was performed in MATLAB R2024a using the Fuzzy Logic Toolbox and Deep Learning Toolbox packages. The simulation covered the approximation of the response surface 𝑍𝑎 = 𝑓(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ) and a comparative analysis of error at different levels of data noise [31]. The present study is based on simulation experiments conducted in a controlled modeling environment. No validation using real industrial process data is included in the current work. This represents a limitation of the present study. Factors such as sensor delays, correlated noise between measurement channels, actuator constraints, and operational disturbances inherent to real production environments are not explicitly modeled, and may affect the behavior of the proposed approach in practice. The proposed model relies on a set of design-level assumptions regarding input normalization, expert score calibration, and operating regime characterization. These assumptions are explicitly introduced to enable tractable fuzzy inference, and do not imply universal validity outside the defined modeling framework [32]. Further research will focus on validating the proposed approach using real technological systems; implementing a prototype in a real-time environment; and introducing online training procedures for the parameters of the LSTM-NN neural network block [33, 34]. Future work will focus on validating the proposed approach using real industrial process data; extending the experimental setup to include time delays and correlated noise; and implementing a prototype within a real-time control environment.
6. Conclusion This paper presents a model-based and simulation-validated approach for assessing the interference immunity of industrial information and 52
Probabilistic model 12.4 1.00 T 71 None
Neuro-fuzzy model 4.8 0.82 T 91 Implemented
Improvement ↓ 61 ↓ 18 + 20 -
control systems using a hybrid LSTM-DNN and fuzzy inference framework. The results demonstrate the feasibility of integrating temporal prediction and expert knowledge within a formally defined indicator under conditions of noisy and incomplete data. The use of the recurrent LSTM component made it possible to generate forecasts of technological line parameters and compensate for the effect of accumulated noise, while the use of the DNN block ensured the extraction of relevant features from noisy signals. The built-in fuzzy logic module (FIS) implements adaptive assessment of noise immunity 𝑍𝑎 = 𝑓(𝑄𝑖 , 𝑅𝐿𝑆𝑇𝑀 , 𝑅𝐸𝑋𝑃 ), combining objective and subjective factors. The simulation confirmed a 61% increase in the accuracy of stability assessment compared to the probabilistic method, with data noise of up to 15%. The adaptation time to new conditions was reduced by 18%, which made it possible to apply the algorithm in real time as part of the IAODC. The practical significance of this work lies in the creation of an intelligent monitoring tool capable of self-adaptation when technological conditions change and data is partially lost. Further research is planned to expand the functionality of the system by introducing neuro-fuzzy controllers into the control loop and integrating with cloud computing platforms for multizone optimization of production processes.
AUTHORS Barnokhon Temerbekova∗ – Department of Information Technologies and Automation of Technological Processes and Production, Almalyk Branch of the National University of Science and Technology MISIS, Almalyk, 110100, Uzbekistan, e-mail: misis_temerbekova@mail.ru. Gulnora Bekimbetova – Department of Economics and Management, Tashkent State University of Economics, Tashkent 100007, Uzbekistan, e-mail: gmkbbd@gmail.com. Ulugbek Mamanazarov – Department of Information Technologies and Automation of Technological Processes and Production, Almalyk Branch of the National University of Science and Technology MISIS, Almalyk, 110100, Uzbekistan, e-mail: m67811@mail.ru. Bakhodir Bekimbetov – Department of General Professional and Economic Sciences of Almalyk State Technical Institute, Almalyk, 110100, Uzbekistan, e-mail: bakhodir.bekimbetov@inbox.ru (B.B). ∗
Corresponding author
Journal of Automation, Mobile Robotics and Intelligent Systems
References [1] Arufe, L.; Rasconi, R.; Oddi, A.; Varela, R.; Gonzá lez, M.A. New coding scheme to compile circuits for Quantum Approximate Optimization Algorithm by genetic evolution. Applied Soft Computing 2023, 144, 110456. https://doi.org/10.1016/j.asoc.2023.110456 [2] Jang, J.-S.R. ANFIS: Adaptive-Network-Based Fuzzy Inference System. IEEE Transactions on Systems, Man, and Cybernetics 1993, 23(3), 665–685. https://doi.org/10.1109/21.256541 [3] Hochreiter, S.; Schmidhuber, J. Long ShortTerm Memory. Neural Computation 1997, 9(8), 1735–1780. https://doi.org/10.1162/neco.199 7.9.8.1735 [4] Wang, W.; Shao, J.; Jumahong, H. Fuzzy inferencebased LSTM for long-term time series prediction. Scientific Reports 2023, 13, 20359. https://doi. org/10.1038/s41598-023-47812-3 [5] Gulyamov, S.M.; Temerbekova, B.M.; Mamanazarov, U.B. Noise immunity criterion for the development of a complex automated technological process. E3S Web of Conferences 2023, 452, 03014. ht t ps : //doi.org/10.1051/e3sconf/202345203014 [6] Sevinov, J.; Temerbekova, B.; Bekimbetova, G.; Mamanazarov, U.; Bekimbetov, B. Hybrid LSTM– DNN Architecture with Low-Discrepancy Hypercube Sampling for Adaptive Forecasting and Data Reliability Control in Metallurgical InformationControl Systems. Processes 2026, 14(1), 147. htt ps://doi.org/10.3390/pr14010147 [7] Sarkar, S.; Pramanik, A.; Maiti, J. An integrated approach using rough set theory, ANFIS, and Z-number in occupational risk prediction. Engineering Applications of Artificial Intelligence 2023, 117, Part A, 105515. https: //doi.org/10.1016/j.engappai.2022.105515 [8] Temerbekova, B.M.; Mamanazarov, U.B.; Bekimbetov, B.M.; Ibragimov, Z.M. Development of integrated digital twins of control systems for ensuring the reliability of information and measurement signals based on cloud technologies and artificial intelligence. Chernye Metally 2023, No. 4, 39–46. https://doi.org/10.17580/chm.2 023.04.07 [9] El-Nagar, A.M.; Zaki, A.M.; Soliman, F.A.S.; ElBardini, M. Hybrid deep learning controller for nonlinear systems based on adaptive learning rates. International Journal of Control 2023, 96(7), 1710–1723. https://doi.org/10.1080/ 00207179.2022.2067080 [10] Guo, D.; Zhang, C.; Cang, N.; Zhang, X.; Xiao, L.; Sun, Z. New fuzzy zeroing neural network with noise suppression capability for time-varying
VOLUME 20,
N∘ 3
2026
linear equation solving. Artificial Intelligence Review 2025, 58, 126. https://doi.org/10.100 7/s10462-024-11026-4 [11] Sun, Z.; Zhou, Y.; Tang, S.; Luo, J.; Zhao, B. Noise suppression zeroing neural network for online solving the time-varying inverse kinematics problem of four-wheel mobile manipulators with external disturbances. Artificial Intelligence Review 2024, 57, 211. https://doi.org/10.1007/ s10462-024-10804-4 [12] Jin, J.; Lei, X.; Chen, C.; Lu, M.; Wu, L.; Li, Z. A fuzzy zeroing neural network and its application on dynamic Hill cipher. Neural Computing and Applications 2025, 37, 10605–10619. https://do i.org/10.1007/s00521-024-10599-z [13] Chen, M.-S. Neuro-fuzzy approach for online message scheduling in real-time systems. Engineering Applications of Artificial Intelligence 2015, 38, 59–69. https://doi.org/10.1016/j.en gappai.2014.10.002 [14] Avazov, K.; Sevinov, J.; Temerbekova, B.; Bekimbetova, G.; Mamanazarov, U.; Abdusalomov, A.; Cho, Y.I. Hybrid Cloud-Based Information and Control System Using LSTM-DNN Neural Networks for Optimization of Metallurgical Production. Processes 2025, 13(7), 2237. https://doi.or g/10.3390/pr13072237 [15] Moreira, H.A.M.; Gomes, H.P.; Villanueva, J.M.M.; Bezerra, S.T.M. Real-time neuro-fuzzy controller for pressure adjustment in water distribution systems. Water Supply 2021, 21(3), 1177–1187. https://doi.org/10.2166/ws.2020.379 [16] He, B.; Zhu, G.; Han, L.; Zhang, D. Adaptive-NeuroFuzzy-Based Information Fusion for the Attitude Prediction of TBMs. Sensors 2021, 21(1), 61. htt ps://doi.org/10.3390/s21010061 [17] Huang, H.; Arogbonlo, A.; Yu, S.S.; Kwek, L.C.; Lim, C.P. Adaptive neuro-fuzzy inference systembased active force control with iterative learning for trajectory tracking of a biped robot. International Journal of Systems Science 2025, 56(6), 1171–1188. https://doi.org/10.1080/002077 21.2024.2420069 [18] Espitia, H.; Machó n, I.; Ló pez, H. Proposal of a Compact Neuro-Fuzzy Adaptive Controller for Filling Regulation of Two Coupled Spherical Tanks. International Journal of Fuzzy Systems 2025, 27, 391–409. https://doi.org/10.1007/ s40815-024-01782-4 [19] Carvalho, F.C.; de Oliveira, M.V.F.; Lara-Molina, F.A.; Cavalini, A.A., Jr.; Steffen, V., Jr. Fuzzy robust control applied to rotor supported by active magnetic bearings. Journal of Vibration and Control 2021, 27(7–8), 912–923. https://doi.org/10.1 177/1077546320933734 53
Journal of Automation, Mobile Robotics and Intelligent Systems
[20] Turgunbaev, A.Y.; Temerbekova, B.M.; Usmanova, Kh.A.; Mamanazarov, U.B. Application of the microwave method for measuring the moisture content of bulk materials in complex metallurgical processes. Chernye Metally 2023, No. 4, 23–28. https://doi.org/10.17580/chm.2023. 04.04 [21] Casari, M.; Kowalski, P.A.; Po, L. Optimisation of the adaptive neuro-fuzzy inference system for adjusting low-cost sensors PM concentrations. Ecological Informatics 2024, 83, 102781. https: //doi.org/10.1016/j.ecoinf.2024.102781 [22] Wang, Z.-J. Eigenvector-driven interval priority derivation and acceptability checking for interval multiplicative pairwise comparison matrices. Computers & Industrial Engineering 2021, 156, 107215. https://doi.org/10.1016/j.cie.2021 .107215 [23] Temerbekova, B.M. Application of systematic error detection method to integral parameter measurements in sophisticated production processes and operations. Tsvetnye Metally 2022, No. 5, 79–86. https://doi.org/10.17580/tsm .2022.05.11 [24] Kiani Mavi, N.; Brown, K.A.; Fulford, R.; Goh, M. Forecasting project success in the construction industry using adaptive neuro-fuzzy inference system. International Journal of Construction Management 2024, 24(14), 1550–1568. htt ps://doi.org/10.1080/15623599.2023.226667 6 [25] Chicaiza Salazar, W.; Ortiz Machado, D.; Gallego Len, A.J.; Escañ o Gonzalez, J.M.; Bordons Alba, C.; de Andrade, G.A.; Normey-Rico, J.E. Neuro-Fuzzy Digital Twin of a High-Temperature Generator. IFAC-PapersOnLine 2022, 55(9), 466–471. https: //doi.org/10.1016/j.ifacol.2022.07.081 [26] Nguyen, T.T.T.; Na, J.; Nguyen, L.T.; Wang, X. Neuro-Fuzzy Network-Based Nonlinear Hybrid Active Noise Control Systems. Entropy 2025, 27(2), 138. https://doi.org/10.3390/e27020 138 [27] Kü the, S.; Alonso Oñ a, I.; Glaser, B. Adaptive Neuro-Fuzzy Inference System–Long Short-Term Memory Hybrid Model to Forecast Castability of Al-Killed Steel Prior to Continuous Casting. Steel Research International 2025, 96(8), 2400220. https://doi.org/10.1002/srin.202400220
54
VOLUME 20,
N∘ 3
2026
[28] Kumar, M.; Mondal, S. Advancements and prospects of fuzzy-based adaptive unscented Kalman filters for nonlinear systems: A review. Applied Soft Computing 2025, 177, 113297. https://doi.org/10.1016/j.asoc.2025.113297 [29] Nik-Khorasani, A.; Mehrizi, A.; Sadoghi-Yazdi, H. Robust hybrid learning approach for adaptive neuro-fuzzy inference systems. Fuzzy Sets and Systems 2024, 481, 108890. https://doi.org/10 .1016/j.fss.2024.108890 [30] Safari, A.; Ghaemi, S. NeuroFuzzyMan: A hybrid neuro-fuzzy BiLSTM stacked ensemble model for financial forecasting and analysis: Dataset case studies on JPMorgan, AMZN and TSLA. Expert Systems with Applications 2025, 266, 126037. htt ps://doi.org/10.1016/j.eswa.2024.126037 [31] Kandasamy, J.; Ramachandran, R.; Guizani, S.; Hamam, H. Adaptive fuzzy-recurrent neural network tuned fractional-order distributed control for robust frequency regulation in multimicrogrid systems. Scientific Reports 2025, 15, 33042. https://doi.org/10.1038/s41598-02518242-0 [32] Inga Espinoza, C.H.; Palma, M.T. A Coordinated Neuro-Fuzzy Control System for Hybrid Energy Storage Integration: Virtual Inertia and Frequency Support in Low-Inertia Power Systems. Energies 2025, 18(17), 4728. https://doi.org/ 10.3390/en18174728 [33] Gu, W.; Lan, J.; Mason, B. Online neuro-fuzzy model learning of dynamic systems with measurement noise. Nonlinear Dynamics 2024, 112, 5525–5540. https://doi.org/10.1007/s11071024-09360-x [34] Jiang, Z.; Xue, H.; Yue, H.; Bao, X.; Zhu, J.; Wang, X.; Zhang, L. A Review of Artificial Intelligence-Driven Active Vibration and Noise Control. Machines 2025, 13(10), 946. https://doi.org/10.3390/machines13100946
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
MICRO-ROS SYSTEM INTEGRATION ONTO THE TURTLEBOT3 MOBILE ROBOT PLATFORM Submitted: 8th September 2025; accepted: 20th December 2025
Bartłomiej Stadnik, Artur Wymysłowski DOI: 10.14313/jamris-2026-038 Abstract: Micro-ROS extends Robot Operating System (ROS) capabilities to resource-constrained environments, enabling seamless integration into robotic systems. This paper focuses on the integration of the micro-ROS system into the TurtleBot3 platform, porting the existing functionalities of the mobile robot into the new architectural layout, reducing the complexity of the system, and obtaining benefits by leveraging advantages linked to the updated software architecture. The performed experiments demonstrate that the new architecture maintains the performance expected from a ROS 2-based platform while significantly reducing power consumption, simplifying the overall design, and lowering jitter. Keywords: micro-ROS, ROS, OpenCR, real-time systems
1. Introduction The Robot Operating System (ROS) has served as the backbone for numerous robotics applications over the past decades. Although ROS 2 has introduced significant enhancements [1], such as distributed communication via Data Distribution Service (DDS) and improved real-time capabilities, its deployment on resource-constrained microcontrollers remains challenging. Prior work [2] demonstrated that the micro-ROS system can extend ROS 2 capabilities to devices with limited computational resources (e.g., ESP32). In this paper, we present the results of the research on the integration approach for mobile robots by porting micro-ROS to the TurtleBot3 platform on the OpenCR1.0 board [3] with ARM Cortex M7, including a floating-point unit (FPU). Traditionally, TurtleBot3 architecture incorporates a Raspberry Pi to run ROS 2 nodes and an OpenCR1.0 board for low-level control of Dynamixel servos, and communication between these components is achieved using rosserial over a USB connection. However, this dual-controller scheme may increase architectural complexity and power consumption. Our objective is to consolidate control and communication functions on a single hardware platform, improving security, reducing power consumption, and simplifying the overall design while using the new software. Here we propose an update to the general framework of the robot by the introduction of the microROS system, installing it directly onto the
OpenCR1.0 board. The updated software stack introduces several key capabilities, including support for the DDS, a node-based modular architecture, and standardized communication interfaces. In our streamlined approach, the Raspberry Pi is removed entirely, with the ROS 2 node running directly on the microcontroller while simultaneously managing low-level control. Furthermore, to maintain the complete functionality of the initial system, wireless communication was implemented in two variants: first including Grove Bluetooth module and UART communication; second including Grove WiFi-UART module and UDP protocol.
2.
TurtleBot3 Overview
The TurtleBot3 platform, developed by Robotis in collaboration with Open Robotics, represents a widely adopted standard for research and education in mobile robotics. Designed as a low-cost modular system compatible with ROS and ROS 2, it balances affordability with robust functionality, enabling applications ranging from autonomous navigation to human–robot interaction studies. Among its two primary variants: the Waffle, optimized for manipulation tasks with its broader chassis, and the Burger, a compact model prioritizing agility, this research focuses on the Burger variant (Figure 1). Its minimalist design reduces mechanical complexity while retaining core capabilities, making it an ideal candidate to evaluate embedded ROS 2 solutions. At its core, the TurtleBot3 Burger relies on a bifurcated architecture that segregates high-level computation from low-level control. The higher layer traditionally employs a single board computer (SBC), a Raspberry Pi (e.g., Pi 4), which hosts ROS 2 nodes for tasks such as simultaneous localization and mapping (SLAM), path planning, and sensor fusion. This subsystem runs on a Linux distribution and communicates with the actuators of the robot through a secondary microcontroller board, the OpenCR1.0. Beyond actuation, TurtleBot 3 is equipped with the platform sensing suite, including a LiDAR (LDS-01) for environmental mapping and a 9-axis inertial measurement unit (IMU) for orientation estimation. Communication between the SBC and the microcontroller occurs via serial communication using USB connectivity.
Open Access. © 2026 Bartłomiej Stadnik and Artur Wymysłowski, published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP. 4.0 License
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives
55
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 2. OpenCR1.0 layout [3] Figure 1. TurtleBot3 general layout [3] More resource-intensive tasks, such as RViz visualization, can instead be executed on a separate networked computer that also supports ROS communication.
3. OpenCR1.0 Microcontroller Board The OpenCR1.0 (Figure 2), developed by Robotis, serves as a bridge between software abstraction and physical hardware. Built around an STM32F746NGH6U microcontroller (216 MHz ARM Cortex-M7 core, 1 MB Flash, 320 KB SRAM), it executes real-time control loops for the Dynamixel AX-12A smart servos that drive the robot wheels. These actuators integrate positional feedback via built-in encoders, enabling closed-loop control. Communication between OpenCR1.0 and Dynamixel servos occurs over a half-duplex TTL bus. Consequently, OpenCR1.0 serves as the embedded control center within the TurtleBot3 project. The microcontroller utilizes an Arduino-based framework through the Arduino IDE. This software provides dedicated libraries and allows for a simplified development process and rapid prototyping with easy firmware updates via a Device Firmware Upgrade (DFU). At its core, OpenCR1.0 interfaces with Dynamixel AX-12A servos through the dedicated Dynamixel2Arduino library of Robotis, which abstracts the complexities of the Dynamixel TTL communication protocol (Protocol 2.0). This library provides methods for configuring servo parameters (e.g. operating modes, velocity limits), reading positional feedback from built-in encoders, and executing synchronized motion commands. For instance, wheel velocity control is achieved by converting ROS geometry_msgs/Twist messages into Dynamixel-specific angular velocity targets, leveraging the function of the provided library. The OpenCR1.0 firmware parses incoming Twist commands from the Raspberry Pi into Dynamixel actuator instructions while simultaneously aggregating sensor data (e.g., IMU readings, wheel odometry) into ROS sensor_msgs and nav_msgs formats for upstream processing.
56
However, this architecture imposes inherent constraints: Rosserial lacks native support for ROS 2 Quality of Service (QoS) policies, and its serialization/de-serialization routines introduce latency during high-frequency data exchanges.
4.
RPi Based Architecture
In the default configuration, a Raspberry Pi communicates with an OpenCR 1.0 board using Rosserial [4], a legacy protocol inherited from ROS 1. Operating over a USB–UART connection, Rosserial converts ROS messages into a serialized byte stream with a custom protocol, establishing a client–server communication model. Messages from the standard ROS network are serialized and transmitted via a serial link to a Rosserial client using the defined packet format; the reverse process occurs when data from the microcontroller are sent to the Rosserial server. The protocol implements custom methods for data negotiation to determine publish–subscribe pairs by querying embedded devices for necessary information such as topic names, message types, and other metadata. In the Rosserial frame structure, a typical message begins with a synchronization flag, followed by protocol version information, a payload length indicator, and then additional fields such as CRC length, topic ID, serialized payload, and finally CRC message. Furthermore, time synchronization between devices is achieved by transmitting std_msgs::Time messages and applying the necessary clock offset. Although functional, the protocol has several drawbacks that could limit its usage. For example, it lacks native compatibility with ROS 2, which introduces significant challenges when attempting to integrate legacy systems with the modern ROS 2 ecosystem. The architectural changes in ROS 2, such as the introduction of DDS-based communication and enhanced security features, render Rosserial less adaptable. For instance, the adopted standard for microcontroller interfacing with ROS 2 is the eProsima Micro XRCE-DDS allowing for communication eXtremely Resource Constrained Environments (XRCEs) with the DDS framework. A detailed comparison with Micro XRCE-DDS reveals some of the inherent shortcomings of Rosserial.
Journal of Automation, Mobile Robotics and Intelligent Systems
Although rosserial lacks standardized framing techniques, Micro XRCE-DDS over serial transport employs HDLC framing, providing a more defined and reliable structure. In Rosserial, the CRC used for message verification is based on a simplistic modulo 256 calculation, which is less rigorous compared to the CRC-16-CCITT standard adopted by Micro XRCE-DDS. These differences reflect a larger issue: Rosserial error-checking and data integrity measures are not as robust, leaving room for potential reliability issues in more demanding applications. Memory usage is another area where Rosserial shows limitations. The serialization process in Rosserial requires the hardware buffer to be as large as the serialization buffer, leading to potentially higher memory overhead. In contrast, Micro XRCE-DDS makes use of a more efficient buffer management strategy, reducing memory consumption by decoupling the size of the hardware buffer from that of the serialization buffer. This not only optimizes resource usage, but also allows for larger message sizes dictated by the uxr_stream rather than the physical limitations of the hardware buffer [5].
5. Researched Micro-ROS Based Architecture The micro-ROS system extends the core concepts of ROS 2 to resource-constrained devices by introducing its lightweight but fully compatible implementation. It utilizes a layered architecture that reuses many of the foundational components of ROS 2, ensuring consistency in node design, communication protocols, lifecycle management, and system interoperability. The following points summarize the key aspects of micro-ROS: - Resource-Constrained Adaptation: Unlike traditional ROS 2 implementations that require significant computational resources, micro-ROS is specifically tailored for microcontrollers and embedded systems. It achieves this by employing a design that minimizes dynamic memory allocation after the initialization phase, thereby ensuring efficient and deterministic operation in real-time environments. - Distributed and Modular Design: micro-ROS embraces a decentralized architecture of the traditional ROS 2 system. Using the discovery system, the nodes broadcast their presence via the DDS API, allowing a modular fault-tolerant system where each component can operate independently. - Middleware Integration: At its core, micro-ROS uses Micro XRCE-DDS middleware to interface with DDS communication [6]. This middleware layer enables micro-ROS to adopt the DDS communication model, including support for publish– subscribe and client–service interactions. However, discovery processes (node broadcasting its presence via the DDS API) and Quality of Service (QoS) mechanisms are offloaded by the ROS Agent, as they would be too computationally expansive for an embedded device. As such, the ROS Agent is expected to run on a host capable of running standard ROS 2 application.
VOLUME 20,
N∘ 3
2026
Figure 3. Architecture of the micro-ROS stack [7]
- Lightweight Abstraction Layers: To bridge the gap between resource-limited devices and the fullfledged ROS 2 ecosystem, micro-ROS introduces the RCLC library, a compact C API that mirrors the functionality of the more comprehensive RCLCPP. This enables developers to implement executors, manage node lifecycles, and schedule callbacks with minimal overhead. The architectural design of micro-ROS not only preserves the familiar concepts of ROS 2 but also introduces innovations that make it suitable for embedded systems, as seen in the architecture layout in the Figure 3. This balance of compatibility and optimization supports diverse robotic uses, ranging from small-scale sensors to mobile platforms, effectively bridging the gap between high-level robotics frameworks and low-level hardware constraints. However, when using micro-ROS, the microcontroller is always the client within the network, while the host role is taken by the ROS Agent. It imposes some limitations on system usage; it is particularly impossible for the microcontroller to use ROS 2 software to communicate directly with another MCU through peer-to-peer architecture.
6.
Proposed Architecture
The proposed architecture (Figure 4) eliminates the necessity of the Raspberry Pi SBC by embedding complete ROS specific components directly on the OpenCR1.0 board. For that, the micro-ROS software was installed on the existing OpenCR1.0 board and combined with the already present software for lowlevel control, together fulfilling application needs as a sole hardware component. The OpenCR1.0 board falls within the scope of micro-ROS supported devices [7], with the restriction of official support limited to Arduino framework with serial communication protocol for interfacing with the ROS Agent. Since the board lacks the ability for wireless communication, it was complemented with the necessary modules in two variants: One with the Grove series Bluetooth v3; Second with Grove UART Wifi V2.
57
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Listing 1. Binding connection to virtual serial port rfcomm connect /dev/rfcomm0 <MAC > # Bluetooth 2 socat -d -d PTY ,raw ,echo=0,link =/dev/ ttyWiFi UDP6 -LISTEN :8888 , reuseaddr WiFi 1
Figure 4. Original architecture compared to proposed one
Each module enables the transmission of serial data by connecting it to the RX/TX pins of the microcontroller (Serial 1), passing the data wirelessly via Bluetooth or WiFi, thus enabling message exchange with the ROS 2 application running on a host system. For the Bluetooth variant, the host system was paired using the rfcomm [8] utility. This involved binding the module MAC address to a virtual serial port (e.g., /dev/rfcomm0, listing 1), enabling the ROS Agent to interact with the device as a standard serial interface. On the microcontroller side, the Bluetooth module was pre-configured via AT commands to operate at a baud rate of 115200 bps, with persistent settings for device visibility and pairing mode. For ROS integration, the micro-ROS agent was initialized with the designated virtual port, treating the Bluetooth link as a conventional serial channel for publishing sensor data and subscribing to control topics. Following the Bluetooth setup, interchangeable steps can be prepared for the WiFi variant. The Grove UART WiFi V2 module was configured through AT commands executed during microcontroller initialization. The configuration sequence set the module to station mode, established persistent WiFi credentials, verified network parameters, and initiated UDP connectivity to a designated host IP and port. Crucially, transparent transmission mode was enabled to allow continuous data streaming. On the host system, the socat utility bridged the UDP packets to a virtual serial interface (e.g., /dev/ttyWiFi, listing 1), which served as the communication endpoint for the ROS system. The ROS Agent was subsequently launched against this virtual port at 115200 bps, identical to the Bluetooth configuration, ensuring consistent ROS API exposure regardless of the underlying wireless technology.
58
#
Although the micro-ROS Agent natively supports UDP transport, the official OpenCR1.0 micro-ROS client distribution is provided and tested for serial transport only. To use an IP-based link (WiFi/UDP) from the OpenCR board, we implemented a custom transport that carries the Micro XRCE-DDS wire protocol over an ESP-AT based WiFi module. The micro-ROS middleware allows for interfacing with the lowest-level transport layer at runtime, essentially making the Micro XRCE-DDS wire protocol transmittable over any communication mechanism. As such, a custom transportation can be tailored specifically for our use case. Following the official guide [9], it is required to implement external transport callbacks as in the listing 2 in the form of functions for opening, closing, reading from, and writing to the custom link. Integration was made by reusing the upstream official micro-ROS Arduino codebase [10], where WiFi specific code was provided. The repository has compatibility with ESP-AT modules upon defining the BOARD_WITH_ESP_AT macro. Our Grove UART WiFi V2 module falls into this category, as it is configured via AT commands over a serial interface. This code utilizes an underlying library to manage the WiFi module via AT commands, establishing a UDP client that connects to the IP address and port of the host running the micro-ROS agent. Consequently, the agent can be launched in UDP mode, allowing the OpenCR1.0 board to communicate directly over WiFi without an intermediate serial-to-UART bridge. Listing 2. Custom transport definition rmw_uros_set_custom_transport ( false , // Framing disabled here 3 (void *) &locator , 4 arduino_wifi_transport_open , 5 arduino_wifi_transport_close , 6 arduino_wifi_transport_write , 7 arduino_wifi_transport_read 8 ); 1 2
7.
Environment Setup
With Arduino as the main framework of the board, the entirety of the application was hosted and built using the PlatformIO IDE solution for VS Code, enabling easier dependency management with support for micro-ROS libraries. As OpenCR 1.0 is not officially supported by PlatformIO, we applied a set of community-contributed patches following the approach described in the Dragonbot blog [11]. This required manually creating package metadata files in PlatformIO directories to register the modifications. After applying these changes, PlatformIO recognizes OpenCR as a board, meaning projects targeting OpenCR could be initialized and built using platformio init –b opencr commands.
Journal of Automation, Mobile Robotics and Intelligent Systems
The listing 3 directs PlatformIO to use the patched STM32 platform with OpenCR board metadata, including the Arduino core. The lib_deps entries explicitly reference both the micro-ROS PlatformIO integration and the Dynamixel2Arduino library, ensuring hardwarespecific servo control capabilities. Enabling deep+ library resolution ensures that nested includes from micro-ROS are discovered correctly. Specifying the micro-ROS distribution as Humble aligns the firmware expectations with the host-side packages and tools; in practice, this setting propagates into build scripts so the correct client libraries and settings are selected for Humble, which is a recommended ROS 2 release. By default, the transport of micro-ROS is set to serial and can be switched to a custom one that enables a UDP connection via the WiFi module. Furthermore, with the board framework using its own boot loader called opencr_ld, the upload command needs to be customized by directly invoking the OpenCR tool rather than a generic STM32 programmer. Listing 3. Example configuration in platformio.ini file [env:opencr] platform = ststm32 3 board = opencr 4 framework = arduino 5 lib_deps = 6 https:// github.com/micro -ROS/ micro_ros_platformio 7 robotis -git/Dynamixel2Arduino @ ~0.7.0 8 jandrassy/WiFiEspAT@ ^2.0.0 9 lib_ldf_mode = deep+ 10 board_microros_distro = humble 11 board_microros_transport = custom 1
VOLUME 20,
8. Application Development Application development starts with mirroring the code of the TurtleBot3 repository to have the original codebase as a baseline. The first task is to identify and remove all code, libraries, and configuration tied to rosserial-based communication while preserving the core control loops and the low-level hardware abstraction (motor drivers, sensor reads, IMU handling, and encoder processing). Once the Rosserial references are stripped out, the remaining code is organized into clear modules: - Hardware Abstraction Module: contains low-level drivers for Dynamixel servos, IMU, encoders, and any serial peripherals (e.g., Bluetooth/WiFi module). These routines expose simple C/C++ functions for reading sensors and commanding motors.
2026
- micro-ROS Client Module: encapsulates initialization of the micro-ROS node, RCLC support setup, allocator configuration, publishers/subscribers, timers, and executors. This module handles all interactions with the ROS Agent via XRCE-DDS over the chosen serial transport. - Control Logic Module: implements control loops (e.g., odometry calculations, goal-velocity updates, safety checks), which invoke hardware abstraction functions and publish sensor data or act on incoming commands. - Connectivity and Diagnostics Module: manages transport setup, monitors ROS Agent connectivity, handles reconnection logic, and manages LED/error indicators for status feedback. After organizing into the four modules, we proceed with incremental integration. First, hardware abstraction is verified standalone. Then, we embed the microROS setup in a compact initialization function. Listing 4 shows a minimal publisher that replaces Rosserial communication: Listing 4. Example of micro-ROS publisher definition
2
On the host side (a personal computer, e.g., RPi SBC), the application setup for communication with the micro-ROS is resolved with Docker images [12]. Firstly, for interfacing with the microcontroller, a ROS Agent needs to be present; as such, running the command with the specified serial port (or UDP mode with the port number) is enough to have a functional application. Secondly, to run ROS 2 applications on the TurtleBot3, a container with installed ROS 2 dependencies was prepared, along with packages providing functionalities for common messages, visualization, and a teleoperation interface.
N∘ 3
1 2 3 4 5 6
RCCHECK( rclc_publisher_init_best_effort ( &publisher_sensor_state , &node , ROSIDL_GET_MSG_TYPE_SUPPORT ( turtlebot3_msgs , msg , SensorState), "sensor_state "));
In order to address bottlenecks, the communication type was configured to a predefined setting of best_effort. This guaranties the delivery of the most upto-date data and avoids the retransmission mechanics that could impose a significant strain on the network, particularly under unstable conditions. The control loops were resolved using RCLC timers paired with dedicated publishers, enabling scheduled publications on different types of topics (listing 5). Timers need to be assigned to the executor; then, an executor needs to have a specified number of handlers for proper initialization: Each timer or subscription added will increment this number. In our case, we grouped publishing timers into one executor and created another executor dedicated to subscription handling; Thus, we minimize callback interference with more consistent loop timing and guaranty easier scalability. Listing 5. Example timer definition with executor initialization 1 2 3 4 5
RCCHECK( rclc_timer_init_default ( &timer_drive_info , &support , RCL_MS_TO_NS (1000 / PUBLISH_FREQUENCY ), publishDriveInformationCallback ));
6 7
8
RCCHECK(rclc_executor_init (& executor , & support.context , 3, &allocator)); RCCHECK( rclc_executor_add_timer (& executor , &timer_drive_info));
To maintain a common interface of messages native to the TurtleBot3, a package containing their definitions is required to be present at build time. 59
Journal of Automation, Mobile Robotics and Intelligent Systems
Because micro-ROS rebuilds message interfaces through its own framework, the TurtleBot3 message package had to be included in the workspace before building. To automate this, custom or vendor-specific packages can be placed inside extra_packages directory at the repository root, either by adding them directly or by creating an extra_packages.repos file with entries pointing to their GitHub repository URLs. During the build, the micro-ROS setup fetched and compiled these definitions so that the common TurtleBot3 interfaces were maintained seamlessly. As of today, microROS does not fully support the standard transformation package, which is called the TF2 – its usage is limited to the underlying tf2_msgs and elementary math utilities, assuming their placement in extra_packages directory. To work around this, one can implement transform broadcasting manually: instantiate a TransformStamped message, populate its header (stamp, frame IDs) and transform fields (translation and quaternion), then publish it on the /tf topic at the desired rate. In practice, this means creating a dedicated micro ROS publisher and timer callback that periodically fills out and emits the TransformStamped message, replicating the behavior of TransformBroadcaster without requiring the full TF2 library. The connection to the ROS Agent was confirmed via the built-in rmw_uros_ping_agent command, which allowed the application setup to block until the agent responded. If the ping fails, the system automatically retries and waits, resuming initialization as soon as the agent is reachable. Finally, the connection with the ROS Agent was complemented with calls to the rmw_uros_sync_session function, enabling time synchronization between devices using built-in tooling.
9. Experimental Results The application assumes an active connection to the TurtleBot 3 ecosystem. In addition to running the ROS Agent and the robot itself, we leveraged key components of the TurtleBot 3 software stack, such as RViz visualization, the robot description package, teleoperation modules, and other related packages for seamless compatibility. Using the ros2 launch turtlebot3_bringup rviz2.launch.py we start all core TurtleBot 3 drivers, TF broadcasters, and the RViz visualization. Next, in another terminal, we run ros2 run turtlesim turtle_teleop_key with the proper topic remap to send keyboard commands to the /cmd_vel topic. Each key press is picked up by the robot’s micro ROS subscriber, translated by the hardware abstraction layer into Dynamixel servo instructions, and executed by the motors. The robot successfully reads and processes the message and sends back all the required data, including the odometry and sensory information. Figure 5 shows the output of the RViz tool, in which the odometry of the moving robot is displayed, reconstructing the robot trajectory relative to its starting pose.
60
VOLUME 20,
N∘ 3
2026
Figure 5. Visualized odometry of robot moving within 1x1m square
For an experiment, we quantified the end-to-end transport delay using the built-in CLI utility ros2 topic delay. This tool subscribes to a topic and computes the time difference between each message header/stamp and the host reception time. With enough message count, the statistics are provided as follows: minimum, maximum, mean, and standard deviation, presenting a compact view of latency and jitter without custom parsing. Because accurate one-way measurements require synchronized clocks, the micro-ROS client was configured to perform time synchronization with the ROS Agent (via the rmw_uros_sync_session calls), so header/stamp values emitted on the OpenCR1.0 are directly comparable to host timestamps. While the TurtleBot3 runs the full set of topics and TF broadcasters (each with best_effort QoS policy), we deliberately selected a single representative telemetry stream for the transport comparison. The tested topic is /sensor_state (turtlebot3_msgs/SensorState), a custom message that aggregates multiple sensor readings and status fields into a single payload; because it contains many fields, it provides a realistic measure of both serialization overhead and network behavior for typical robot telemetry. For the experiments, this topic publish rate was set at 10 Hz, with the window of 10 000 samples. The test was prepared for every variant of the application: - custom micro-ROS transport with WiFi module and direct UDP protocol, - serial micro-ROS transport with WiFi module and Socat mapping of UDP to virtual serial port, - serial micro-ROS transport with Bluetooth module and rfcomm mapping to virtual serial port. Each configuration was evaluated in two variants: first, via direct topic analysis, and then with a constant 30 Hz background input stream published to the robot /cmd_vel topic to emulate a sustained teleoperation load and observe its effect on latency and jitter. The results (Table 1) show that native UDP transport yields the lowest mean latency and the tightest dispersion in our setup, whereas the Socat based WiFi bridge introduces substantially higher and more variable delays.
Journal of Automation, Mobile Robotics and Intelligent Systems
Table 1. One-way latency statistics for /sensor_state (times in milliseconds) Test (ms) UDP UDP + input UDP (Socat) UDP (Socat) + input Bluetooth Bluetooth + input
Mean 16 28 90 80 78 123
Min 14 14 72 76 32 41
Max 37 114 284 225 169 256
Std 1.72 15.60 32.53 9.50 20.39 35.14
Bluetooth results lie between the two, with moderate baseline latency but greater sensitivity to conditions. Introducing a constant 30 Hz command stream increased mean latency and jitter on multiple transports, indicating that link contention and bridging/buffering behavior materially affect end-to-end timing by prioritizing the input stream. Interestingly, the Socat bridge showed slightly lower latency under additional load, likely due to buffering and flushing effects in the pseudo-terminal mapping rather than a genuine transport improvement.
10.
Comparison with Original TurtleBot3 Layout
To quantify the effect of reintroducing the Raspberry Pi, we set up a hybrid configuration: the OpenCR1.0 board still runs the identical micro-ROS application but now communicates over USB serial to a Raspberry Pi 4, which hosts the ROS Agent. The RPi is connected via WiFi to the host PC running RViz and diagnostic tools. This mimics the original TurtleBot3 split architecture (with micro-ROS instead of rosserial) and allows us to directly compare performance against the single-board design. In this setup, the micro-ROS node on OpenCR1.0 and the ROS Agent on the RPi form a chained data path, so messages from the MCU must traverse the USB link to the RPi, and then the wireless network to the host. Accurate time synchronization across all devices is critical for valid latency measurements. We configured the host PC as an NTP time server (Chrony [13]) and the Raspberry Pi as its client. On the host, we added a network-wide allow rule in /etc/chrony/chrony.conf ; on the Pi, we added server <host_ip> iburst and restarted Chrony on both machines. This ensures the RPi’s clock tracks the host’s clock. Concurrently, the OpenCR1.0 microcontroller uses the built-in time-sync mechanism of micro-ROS (rmw_uros_sync_session) to align its timestamps with the ROS Agent. With these steps, all components share a common time base, and the topic delay utility can yield accurate one-way latency values. We then repeated the latency experiment exactly as before. The OpenCR1.0 micro-ROS node publishes the /sensor_state topic at 10 Hz. On the host PC, we ran the topic delay command to measure one-way message latency. Tests were performed both under idle conditions and with a sustained 30 Hz command stream on /cmd_vel to simulate high-load operation.
VOLUME 20,
N∘ 3
2026
Table 2. One-way latency statistics for RPi intermediary vs. best uC standalone (times in milliseconds) Test (ms) RPi RPi + input uC UDP uC UDP + input
Mean 14 18 16 28
Min 6 13 14 14
Max 415 334 37 114
Std 16.03 3.86 1.72 15.60
All timestamps originate from the MCU micro-ROS publisher, and due to our synchronization setup, the experimental results directly reflect the true transit time. We obtained latency statistics suitable for direct comparison by collecting the same number of samples as in the previous experiment. Analyzing the results (Table 2), the single-board OpenCR1.0 with native UDP transport achieves more consistent low-latency behavior and avoids OS-induced spikes, supporting the claim that removing the Raspberry Pi simplifies the architecture and improves determinism. The RPi intermediary shows competitive or better average latency under load (because it can absorb bursts); however, the worst response time is significantly larger. In addition to communication performance and system complexity, power consumption plays a critical role in mobile robotics. To quantify the impact of the architectural simplification, we conducted direct power measurements of both configurations under typical runtime conditions, as presented in Table 3. Energy measurements made with a USB power tester show that the OpenCR1.0 + UART-WiFi configuration draws, on average, 1.19 W (≈237 mA at 5 V), whereas the OpenCR1.0 with a Raspberry Pi 4 intermediary draws 3.77 W (≈754 mA at 5 V). This corresponds to an ≈68.6% reduction in average power when the Raspberry Pi is removed, saving about 2.58 Wh per hour. Taken together, these results show that the single-board OpenCR 1.0, with a native UDP design driven by the micro-ROS framework, is a viable alternative to the RPi intermediary.
11.
Conclusion
This study demonstrates that micro ROS can be effectively ported onto the TurtleBot3 OpenCR1.0 board, enabling the consolidation of processing and control into a single embedded system. The resulting architecture reduces system complexity, power consumption, and preserves the performance expected from a ROS 2 based platform. Our transport experiments between application variants indicate that native UDP provides the best latency characteristics for time-sensitive telemetry, while serial bridges (Bluetooth or socat-based WiFi) introduce larger and more variable delays that should be considered when designing control loops. Furthermore, Bluetooth typically consumes less power and is easy to set up for short-range, low-bandwidth control links. However, it has a limited range and throughput, and it can perform poorly in noisy RF environments. Compared to Bluetooth, WiFi offers substantially higher throughput, longer range, and stronger standardized security (WPA2/WPA3), making it better for high-rate telemetry and remote operation, at the expense of more complex configuration. 61
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 3. Measured energy and average power/current during representative runs Configuration OpenCR1.0 + UART-WiFi OpenCR1.0 + RPi4
Energy (Wh) 0.291 0.559
Future research will further investigate the scalability of this approach, with particular emphasis on communication robustness and the integration of additional microcontrollers for extended functionality. With reliable low-latency links and additional sensors, the platform can naturally support higher-level functionalities, such as map-based goal sending and local obstacle handling.
Avg. Power (W) 1.19 3.77
Avg. Current (mA) 237 754
[5] “Micro XRCE-DDS compared to Rosserial”. https: //micro.ros.org/docs/concepts/middleware/r osserial/ [6] I. Ben Abdallah, Electronic Research Archive, vol. 31, 2023, pp. 5083–5103. [7] “Official Site”. https://micro.ros.org/ [8] “Linux Man Page”. https://linux.die.net/man/1/ rfcomm
AUTHORS Bartłomiej Stadnik – Wroclaw University of Science and Technology, Faculty of Electronics, Photonics and Microsystems, ul. Janiszewskiego 11/17, 50-372 Wrocław, Poland, e-mail: bartlomiej.stadnik@pwr.edu.pl. Artur Wymysłowski∗ – Wroclaw University of Science and Technology, Faculty of Electronics, Photonics and Microsystems, ul. Janiszewskiego 11/17, 50-372 Wrocław, Poland, e-mail: artur.wymyslowski@pwr.edu.pl. ∗ Corresponding author
References [1] V. DiLuoffo, W. Michalson, and B. Sunar, International Journal of Advanced Robotic Systems, vol. 15, 2018. [2] B. Stadnik and A. Wymysłowski, Przegląd Elektrotechniczny, 2024, doi: 10.15199/48.2024.12.48. [3] “TurtleBot3 e-Manual”. https://emanual.robotis. com/docs/en/platform/turtlebot3/overview/ [4] “Rosserial Package Documentation”. https://wiki .ros.org/rosserial
62
[9] “Creating Custom Micro-ROS Transports”. https:// micro.ros.org/docs/tutorials/advanced/create _custom_transports/ [10] https://github.com/micro-ROS/micro_ros_ardu ino/blob/humble/src/micro_ros_arduino.h [11] Eclipse Foundation. “DragonBotOne Egg Hatching with Zenoh and Zenoh-Pico”. https://micro.ros.or g/, February 2022 [12] micro ROS Dockers. https://hub.docker.com/r/m icroros/micro-ros-agent [13] R. Curnow and M. Lichvar. “Chrony project”. https://chrony-project.org/
VOLUME 20, N° 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
Research on Fuzzy-PID Control System for Mecanum Robot Submitted: 11th July 2024; accepted: 3rd September 2024
Trong-Tai Nguyen, Quang-Tho Le, Quang-Phuoc Pham, Tran-Long Le DOI: 10.14313/jamris-2026-039 Abstract: In this paper, a comprehensive control system for precise motion control and trajectory tracking of Four Mecanum-Wheeled Mobile Robots (FMWMRs) is presented. The proposed approach addresses issues including nonlinearities, uncertainties, and external disturbances by combining a fuzzy logic controller for orientation control and a fuzzy-PID controller for position control. The dynamic model of the FMWMR is summarized. For localization and heading angle estimates, an IMU sensor and a camera are leveraged, and a complementary filter method is utilized for the best possible sensor fusion. A real-world FMWMR prototype and extensive Matlab/ Simulink simulations show how reliable and effective the control system is at providing accurate trajectory tracking with low RMS errors under a range of loading scenarios. With possible applications in material handling, mobile robotics, and industrial automation, the suggested approach provides a dependable solution for trajectory tracking and motion control in FMWMRs. Keywords: Mecanum wheel, Fuzzy logic control, FuzzyPID control, trajectory tracking, sensor fusion
1. Introduction Nowadays, the demand for mobile robots is rising as they have the potential to free humans from repetitive, dangerous, and labor-intensive tasks [1–4]. In industries such as manufacturing, logistics, and healthcare [5, 6], mobile robots are increasingly utilized for tasks like assembly line work, warehouse management, and patient transport. By taking over these duties, mobile robots not only enhance efficiency and productivity but also reduce the risk of injuries and allow human workers to focus on more complex and creative activities. Thanks to their remarkable omnidirectional mobility, four Mecanum-Wheeled Mobile Robots (FMWMRs) have garnered significant interest across various sectors, including military and industrial applications [7–9]. With special design of mecanum wheel, it allows the robots to navigate confined and intricate spaces with ease, making them highly useful in various fields such as industrial automation, warehouse operations, and search and rescue efforts [10]. Through the independent control of speed and direction for each wheel, this configuration enables FMWMRs to maneuver instantly in any direction from any starting orientation, eliminating the requirement
for pre-orientation. As a result, compared to conventional platforms, these vehicles have excellent mobility and are well-suited for operations in confined places or crowded areas [7]. However, achieving accurate trajectory tracking and motion control of such platforms remains a challenge due to parameter uncertainties, non-linearities, external disturbances and the need for robust control strategies. These factors can further degrade the control performance. As a result, there are many problems that need to be solved when implementing a Mecanum robot project, such as control algorithm problems, position and rotation angle determination problems and so on [9, 11–14]. Many research works have been conducted to develop and implement different kinds of control techniques , such as: PID control [15, 16], backstepping control [17], sliding mode control [18], intelligent control based on neural networks [19, 20]. Fuzzy PID controllers are applied in various fields, including robotics, industrial automation, and process control, where system conditions are constantly changing and require adaptive control strategies. A Fuzzy PID controller combines the principles of a conventional PID (Proportional-Integral-Derivative) controller with fuzzy logic to improve performance in systems with non-linearity, uncertainties, or varying dynamics. Compared to conventional PID controllers, the Fuzzy PID controller demonstrates superior characteristics such as better handling of non-linearities, robustness in conditions of system uncertainty, and effective adaptation to dynamic environments without needing a precise mathematical model. Thanks to these advantages, Fuzzy PID controllers have been extensively investigated in numerous research studies [16, 21–26]. In [16, 21–23, 27], the Fuzzy PID controller was employed to control the mobile robot in condition of uncertainty factors such as friction, disturbance, and the robot’s nonlinearity. While in [24–26], the fuzzy PID is applied in some other fields such as a permanent magnet synchronous motor, tranformer coil temperature or active suspension. For effective trajectory tracking and motion control in FMWMRs or mobile robots, precise localization and heading angle estimation need to be determined. Robot positioning and heading angle fusion methods play a crucial role in enhancing the accuracy and reliability of robotic navigation systems. Common positioning approaches include GPS [2], inertial sensors - IMU [28, 29], and vision systems [28, 30], LiDAR [31–33]. Especially, for indoor
Open Access. © 2026 Trong-Tai Nguyen et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP. This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 License.
63
Journal of Automation, Mobile Robotics and Intelligent Systems
robot positioning, various methods have been explored, including IMU (Inertial Measurement Unit) and its sensor fusion techniques, ultrasonic sensors, infrared sensors, vision-based systems, radio frequency technologies, and LiDAR. The technologies, along with the advantages and drawbacks of each method, of these methods are discussed in [31, 34, 35]. Additionally, various fusion algorithms, such as Kalman filters [36], hidden Markov model [37, 38] or Bayesian estimation [2, 28, 39–41], are utilized to merge sensor data and reduce errors. Besides, to measure the robot’s orientation, the IMU, which combines a magnetometer, accelerometer, and gyroscope, is extensively used. Wellknown methods for calculating the heading angle with an IMU include the Complementary filter [29], Extended Kalman filter [28, 41, 42], Madgwick filter [43, 44] and Mahony filter [45]. However, IMUs have intrinsic limitations, including drift overtime due to cumulative sensor errors and vulnerability to external environmental factors like magnetic interference. To address these drawbacks, sensor fusion methods are employed to combine IMU data with other sensor inputs, enhancing overall reliability and performance. In this paper, a comprehensive control system is investigated to achieve precise tracking control of FMWMRs. The controller structure is composed of a Fuzzy logic controller for orientation control and a Fuzzy-PID controller for position control. Through this integration, the trajectory tracking and motion control performance of FMWMRs are significantly improved. Velocities along the x and y axes are computed by the Fuzzy-PID controller based on position error and its derivative, while angular velocity is determined by the Fuzzy logic controller using the error and derivative of the heading angle. These computed velocities are subsequently utilized as inputs to inverse kinematic equations to derive the necessary wheel velocities. System nonlinearities and uncertainties are effectively managed and compensated by the fuzzy rule base through expert tuning. To evaluate the control performance, extensive simulations were conducted in the Matlab/Simulink environment. The PID and Fuzzy controllers
Figure 1. Overall structure of the proposed control system 64
VOLUME 20,
N° 3
2026
were also implemented to assess the superiority of the Fuzzy-PID approach. Additionally, a real-life FMWMR prototype was constructed to verify the effectiveness of the proposed controller under practical conditions. For obtaining the FMWMR’s position and heading angle, a vision technique based on ArUco markers [46, 47] is employed. However, challenges such as sensitivity to lighting conditions and slow response rates are encountered with this approach. To address these limitations, an IMU is integrated with the ArUco marker detection results using a complementary filter [29], thereby improving the response and accuracy of the heading angle estimation. The overall of control approach and sensor fusion methodology can be illustrated in a comprehensive way as shown in Figure 1. The simulation and experimental results demonstrate the superiority of the proposed control system in achieving precise trajectory tracking and motion control for FMWMRs in both simulation and realworld scenarios. The subsequent parts of this paper are organized as follows: Section II summarizes the mathematical model of a FMWMR. Section III details all the steps in Fuzzy and Fuzzy-PID design process. The method of positioning and orientation determination by sensor fusion is then illustrated in Section IV. Simulation, experiment results and discussion are shown in Section V. Finally, Section VI concludes our research and provides future directions for development.
2. Mathematical model
In this research, the mathematical model of FMWMR in [48] is used. In which, the kinematic model and dynamic model of FMWMR can be summarized as: The robot kinematic model:
J w r rw
The robot’s inverse kinematic model:
(1)
r rw J w (2)
Journal of Automation, Mobile Robotics and Intelligent Systems
In world frame – as shown in Figure 2, the robot motion is determined as:
q r (3)
where: φq=[xq yq ϕ]T is a vector of robot motion in inertia frame; φr=[xr yr ϕ]T is a vector robot motion in robot frame; θw=[θw1 θw2 θw3 θw4]T is a vector of rotation angle of robot wheels; ϕ is robot heading angle; rw is radius of mecanum wheel. J is geometry matrix of robot; J+ is inverse matrix of J
1 1 J 1 1
a b 1 1 a b ;J 1 1 a b 1 1 a b a b 1
1 1
1 1
1 a b
1 a b
1 1 1 a b
a and b are haft of width and length of the robot flatform; ℜ(ϕ) is frame convert matrix, which is determined as:
cos sin 0 sin cos 0 (4) 0 0 1
The dynamic model of FMWMR is determined as:
M w D w w F (5)
where: τ =[τ1 τ2 τ3 τ4]T is a vector of applied torque generated by robot wheel motors; F = [F1 F2 F3 F4]; with Fi = fi sgn wi ; i ÷ 1 : 4 represents the static friction on mecanum wheels.
3. Controller Design
For FMWMRs to operate well, trajectory tracking and motion control need a resilient and flexible control system that can deal with uncertainties, nonlinearities, and external disturbances. A hierarchical of control architecture that includes a fuzzy-PID controller for position control with a fuzzy logic controller for orientation control is investigated in order to overcome these issues. While the Fuzzy-PID controller is responsible for regulating the robot’s position, the Fuzzy logic controller is designed to control the robot’s heading angle, guaranteeing precise orientation tracking. The overall block diagram of the control system is illustrated in Figure 3. Each controller’s design process will be detailed in the following subsections, beginning with the fuzzy controller used for orientation control.
3.1. Fuzzy logic controller for orientation control To control the robot’s orientation, a fuzzy logic controller is designed with two inputs and one output, as shown in Figure 4. The inputs consist of the heading angle error between the desired orientation and the robot’s orientation (eϕ = ϕd − ϕ) and its
VOLUME 20,
N° 3
2026
derivative (deϕ). The output is the desired angular velocity of the robot. Since the inputs and output of the fuzzy logic controller operate within different physical ranges, significant time is required to adjust the fuzzy logic for each variable in varying situations. To address this, in this research, all inputs and outputs of the fuzzy controller are normalized to the standard range [-1, 1] using normalization gains. This normalization simplifies the formulation of the fuzzy rule base and provides flexibility in adjusting input and output ranges during performance evaluation and tuning. By adjusting the normalization gains, the expected range can be modified, enabling the designed membership functions to automatically adapt to the new range without requiring a complete redesign.
Fuzzification The design approach for the heading angle fuzzy controller begins with fuzzification. In this step, the crisp input values are converted into fuzzy sets using membership functions. Trapezoidal and triangular membership functions are used in this design due to their simplicity and ability to clearly represent the input and output spaces. Based on experience from simulation and experiment results, the fuzzy inputs are fuzzified using seven membership functions: BN (Big Negative), MN (Medium Negative), SN (Small Negative), ZE (Zero), SP (Small Positive), MP (Medium Positive), and BP (Big Positive). The distribution of the membership functions of each input is illustrated in Figure 5 and Figure 6. Similarly, for the fuzzy output, the angular velocity is fuzzified using seven membership functions: BN (Big Negative), MN (Medium Negative), SN (Small Negative), ZE (Zero), SP (Small Positive), MP (Medium Positive), BP (Big Positive). The distribution of these membership functions is illustrated in Figure 7. The normalization gains for inputs and outputs of orientation controller are selected and presented as Table 1.
Figure 2. Robot in world frame and robot frame
65
Journal of Automation, Mobile Robotics and Intelligent Systems
Figure 3. Block diagram of the control system
Figure 4. Inputs and output of heading angle Fuzzy logic controller
Figure 5. Design of membership functions for heading angle error’s fuzzy input
Figure 6. Design of membership functions for derivative of heading angle error fuzzy input
Figure 7. Design of membership functions for angular velocity fuzzy output 66
VOLUME 20,
N° 3
2026
Journal of Automation, Mobile Robotics and Intelligent Systems
Table 1. Normalization gains of orientation fuzzy controller Gain Ke−yaw Kde−yaw Kω
Description
Value
Orientation error fuzzy input normalization gain Orientation derivative of error fuzzy input normalization gain
Angular speed’s fuzzy output normalization gain
1
1
π
π
4π
10
Fuzzy Rule Base The fuzzy rule base comprises a set of IF-THEN rules that determine the appropriate control action based on the linguistic values of the input variables. The rules consider the combination of error and rate of change of error to determine the appropriate velocities for the robot. Each rule specifies the linguistic values of the inputs and the desired linguistic value of the output. These fuzzy rules were derived from the experience of numerous simulations and experiments, aiming to achieve the most effective control system as shown in Table 2. Fuzzy Inference System Once the fuzzy rule base is established, the fuzzy inference system combines the fuzzy rules with the fuzzified input values to determine the appropriate Table 2. Fuzzy rule base for heading angle control E
ωz BN de
MN SN ZE SP
MP BP
VOLUME 20,
Defuzzification The defuzzification process converts the fuzzy output into a crisp output value of the angular velocity. In this implementation, the Center of Gravity (COG) method is leveraged for defuzzification, which calculates the centroid of the aggregated output fuzzy set.
3.2. Fuzzy-PID controller for position control This section presents the use of fuzzy logic to tune the parameters of PID controllers for trajectory tracking along the x and y axes. Two fuzzy PID controllers are independently designed for each axis. The structure of the fuzzy system, shown in Figure 8, includes two inputs (position error and its derivative) and three outputs (the PID controller gains). Similar to the orientation controller, all the inputs and outputs of fuzzy logic controller in this section are designed in a normalized range of [-1, 1] with inputs and [0, 1] with outputs. The details of the fuzzy tuning process are described as follows:
MN
SN
ZE
SP
MP
BP
BN
BN
MN
MN
SN
ZE
SP
BN
MN MN SN ZE
BN BN
MN SN ZE SP
BN SN SN ZE SP
MP
2026
control action. In this paper, the Mamdani’s Fuzzy Inference method with Max–Min composition is applied. It involves combining the minimum of the degrees of membership for each rule during aggregation and considers the maximum value of these minimum values during defuzzification.
BN BN
N° 3
BN SN ZE SP
MP BP
MN ZE SP SP
MP BP
SN SP
MP BP BP BP
ZE
MP MP BP BP BP
Figure 8. Structure of inputs and outputs Fuzzy logic design
67
Journal of Automation, Mobile Robotics and Intelligent Systems
For the inputs of fuzzy logic, each tracking position error along each axis (ex = xr − x or ey − yr − y) is fuzzified using five membership functions: NB (Negative Big), NS (Negative Small), ZE (Zero), PS (Positive Small), and PB (Positive Big). Similarly, the derivative of the position tracking error (dex or dey) is fuzzified using the same five membership functions: NB (Negative Big), NS (Negative Small), ZE (Zero), PS (Positive Small), and PB (Positive Big). Figure 9 and Figure 10 illustrate the membership functions fuzzy logic inputs. The normalization gains are set to 0.02 for the position error and 0.04 for the error derivative. The fuzzy outputs include three parameters of PID controller Kp, Ki and Kd, which are normalized in range of [0, 1]. These output values are fuzzified using five membership functions: S (Small), MS (Medium Small), M (Medium), MB (Medium Big), B (Big). The membership functions of these outputs are designed similarly as illustrated in Figure 11. Based on the experience through extensive simulations and experiments, for best control performance, PID controller gains are chosen to vary in ranges of: Kp =[9, 11], Ki =[7, 12], Kd =[0.05, 1]. In fuzzy outputs, these parameters are normalized into the range of [0, 1] by equation (6). K x'
K x K x min (6) K x max K x min
VOLUME 20,
2026
where Kx stands for the reality PID gains Kp, Ki or Kd; and stands for the normalized PID gains (Fuzzy outputs) K 'p , Ki' or K d' . From (6), the value of PID parameter gains can be calculated from fuzzy output by equation (7). K x = K 'x (K x max - K xmin ) + K xmin (7)
Thus, as a result, the value of Kp, Ki or Kd can be detailed as: K p = 2 K 'p + 9; K d = 0.05K d' + 0.05; Ki = 5Ki' + 7 . Based on tuning and evaluation of control of performance, the rule base for three outputs of the fuzzy logic is established in a similar manner and is illustrated as in Table 3. To determine the fuzzy outputs, the Mamdani Fuzzy Inference Method with Max–Min composition and COG defuzzification, is applied. This method is identical to the design process used for the fuzzy controller in orientation control mentioned earlier.
4. Experiential Setup and Heading Angle Fusion Technique
4.1. Experimental Setup The structure of the experimental control system is illustrated in Figure 12. In this setup, the visionbased ArUco marker detection method and trajectory tracking Fuzzy PID controllers are executed on a laptop computer. The ATmega 2560 is responsible
Figure 9. Position error’s membership function
Figure 10. Derivative of position error’s membership function
Figure 11. Membership function for fuzzy outputs - the PID normalized gains K 'p , Ki' , K 'd 68
N° 3
Journal of Automation, Mobile Robotics and Intelligent Systems
for calculating the robot’s orientation (Magdwick and Complementary filters), executing the heading angle Fuzzy controller, and managing the wheel speed PID controllers. Communication between the laptop and the ATmega 2560 is facilitated via a Bluetooth module (HC-05). The block diagram showing the functional blocks of the real-world implementation for mecanum robot control is presented in Figure 13.
VOLUME 20,
N° 3
2026
4.2. Heading Fusion Technique Accurate heading angle estimation is crucial for precise trajectory tracking and motion control in FMWMRs. To achieve reliable and robust heading angle estimates, a sensor fusion approach that combines data from a camera-based localization system and an Inertial Measurement Unit (IMU) is investigated.
Table 3. Fuzzy rule base for controlling KP, KI, KD
e
K 'x
de
NB
NS
ZE
PS
PB
S
MS
MB
NB
B
MB
S
ZE
MB
MS
S
NS PS
PB
B
MB M
M
MS S
S S
S
MS M
MB
M
MB B B
Figure 12. Structure of the experimental execution
Figure 13. Overall block diagram of real-life FMWMR control 69
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
2026
An ArUco marker affixed to the robot is detected by the camera-based localization system, employing an overhead camera. Through analysis of the marker’s pose, the system establishes the robot’s position and orientation in the global coordinate system. Despite its utility, ArUco marker detection is susceptible to limitations such as sensitivity to lighting conditions and a slow detection… To mitigate these constraints, IMU data are concurrently utilized to augment the accuracy of heading angle detection. The IMU integrates signals from accelerometers, magnetometers, and gyroscopes, facilitating the Madgwick filter [44] in estimating a drift-free and responsive heading angle. Subsequently, the robot’s heading angle is determined by combining data from both the camera and IMU using the complementary filter [43]. Here, the IMU supplies rapid angle variations via a high-pass filter, complemented by the camera’s provision of slower variations through a low-pass filter. Figure 14 illustrates the complementary filter integration. The mathematical models of the complementary filter using first order filter in both Laplace and Z domains are expressed in Eq. (8) and (9), respectively.
where: ϕfused is fused heading angle (degree); ϕcam is heading angle from the camera (degree); ϕIMU is Ts heading angle from the IMU (degree); = e ; Ts is sampling time; τ is low-pass filter time constant which is chosen as 1s in this research. Figure 15 presents a comparison of the IMU response, vision detection, and the fusion results. Figure 15a highlights the IMU response obtained from Magdwick filter, showing significant detection errors due to accumulation effects and environmental uncertainties. In contrast, Figure 15b displays the vision detection result from the Aruco marker. This result is shown under the worst conditions, where light variability and pixel errors heavily impact detection accuracy. Finally, Figure 15c illustrates the outcome of the complementary filter. This fusion response demonstrates the advantages of the fusion method, effectively reducing the drawbacks of both the IMU and vision sensors. The enhanced heading angle from the fusion can be applied to experimental control in further sections.
1 z 1 (1 ) z 1 ( z) cam ( z ) (9) 1 IMU 1 z 1 z 1 Consequently, the final heading angle can be computed from the IMU and camera as equation (10).
To evaluate the performance of the proposed control system and sensor fusion approach, extensive simulations and real-world experiments are conducted. This section presents and analyzes the results obtained from both the simulation studies and the experimental validation.
fused ( s )
fused ( z )
1 s IMU ( s ) cam ( s ) (8) s 1 s 1
fused (k ) fused (k 1) IMU (k ) IMU (k 1) 1 cam (k 1) (10)
5. Results and Discussion
Figure 14. Block diagram of heading angle fusion using complementary filter
Figure 15. Heading angle detection using IMU, Vision- ArUco marker and Complementary filter fusion 70
N° 3
Journal of Automation, Mobile Robotics and Intelligent Systems
5.1. Simulation Results Simulation results using Matlab/Simulink environment are performed to compare the tracking performance of three control strategies: PID, Fuzzy, and Fuzzy-PID. The simulations were carried out under various scenarios, including desired position, rectangular reference trajectory and eight-shaped reference trajectory. For the desired position scenario, the response and error plots are illustrated in Figure 16 and Figure 17, respectively. Analysis of the response plots demonstrates that the Fuzzy controller achieves the fastest response in reaching the desired x and y positions. In comparison, the PID controller responds slightly slower and exhibits some oscillation before settling while both Fuzzy and Fuzzy-PID controllers show quick and smooth convergence to the desired yaw angle. Despite the differences in response times, all three controllers successfully converge to the target position and heading angle with minimal steady-state error.
VOLUME 20,
N° 3
2026
Next, the rectangular path trajectory scenario control results are presented in Figure 18 to Figure 20, respectively. From the robot’s path in the x-y plane, it is obvious that the Fuzzy-PID controller performs the best, attaining the smoothest and most accurate tracking (Figure 18). The Fuzzy-PID controller maintains the lowest error values in all dimensions (x, y, and heading angle), as shown by the error plots (Figure 20), which validate its improved tracking accuracy. Though having somewhat greater deviations and error magnitudes than the Fuzzy-PID controller, the fuzzy controller also shows good tracking capabilities. On the other hand, the tracking performance of the PID controller is less accurate, exhibiting more pronounced oscillations, undershoots, and overshoots, especially at the rectangle’s corners. Besides, an eight-shaped trajectory was utilized to assess the controllers’ performance further in Figure 21 to Figure 23. This difficult maneuver tests the robot’s capacity to manage abrupt direction changes and sharp turns. All three controllers are able to follow the intended
Figure 16. Position responses with respect to multi-step reference in simulation
Figure 17. Position control errors with respect to multi-step reference in simulation 71
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
Figure 18. Rectangular path trajectory tracking control responses in simulation
Figure 19. Control responses with respect to rectangular path trajectory tracking control in simulation
Figure 20. Control errors with respect to rectangular path trajectory tracking control in simulation 72
N° 3
2026
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 21. 8-shaped path trajectory tracking control results in simulation
Figure 22. Control responses with respect to 8-shape path trajectory tracking control in simulation
Figure 23. Control response error with respect to 8-shape path trajectory tracking control in simulation 73
Journal of Automation, Mobile Robotics and Intelligent Systems
8-shaped path, as seen by the trajectory tracking plot (Figure 21), with the Fuzzy-PID controller adhering to the reference trajectory that is closest. The Fuzzy-PID controller outperforms the others, as seen by the error plots (Figure 23), which show the lowest x, y, and heading angle error magnitudes. Whereas the PID controller displays greater error magnitudes, the fuzzy controller maintains very low error values.
Table 4. Comparison among the x responses of each controllers Setpoint 5
10
15
20
VOLUME 20,
PID
Fuzzy
Fuzzy-PID
Rise time (s)
0.31
0.24
0.37
Steady state error Settling time (s) Overshoot (%) Rise time (s)
Steady state error Settling time (s) Overshoot (%) Rise time (s)
Steady state error Settling time (s) Overshoot (%) Rise time (s)
Steady state error Settling time (s)
8.02
0.0273 0.63 21.8 0.26
0.0318 1.01
24.07 0.31
0.0585 1.19
26.05 0.35
0.0722 1.36
5.07
0.0089 0.5
1.43 0.34
0.006 0.55 2.09 0.22
0.0007 0.66 0.8
0.48
0.0002 0.81
0 0
0.77 0
0.44 0
0.87 0
0.47 0
0.74 2.05 0.51
0.0002 0.79
Table 5. Comparison among the yaw responses of each controllers Setpoint 90
180
270
360
74
Parameters
PID
Fuzzy
Fuzzy-PID
Rise time (s)
0.39
0.25
0.33
Overshoot (%)
Steady state error Settling time (s) Overshoot (%) Rise time (s)
Steady state error Settling time (s) Overshoot (%) Rise time (s)
Steady state error Settling time (s) Overshoot (%) Rise time (s)
Steady state error Settling time (s)
2026
Based on simulation results, the control performances of three controllers are evaluated. Table 4 and table 5 present the comparison results of these controllers. These tables display each controller’s x-position and yaw responses across four criteria: Overshoot, rise time, steady-state error, and settling time. The results demonstrate that the Fuzzy-PID controller exhibits superior performance, characterized
Parameters
Overshoot (%)
N° 3
1.24
0.2911 0.93 0.55 0.4
1.167 0.78 0.4
0.52 1.48 1
1.2
0.54 2.76 0.94
1.11 1.23 0.61 0.27 0.26
0.386 0.48 0.39 0.25
0.756 0.55 0.03 0.27 1.35 0.56
0.06 0.07 0.69 0
0.34 0.32 0.72 0
0.31
0.126 0.63 0
0.37 0.57 0.54
Journal of Automation, Mobile Robotics and Intelligent Systems
by zero overshoot, minimal steady-state error, and a fast response. Overall, the Fuzzy-PID controller has the best overall performance for both the x and yaw responses, offering exceptional stability, precision, and convergence speed, according to the provided data. The Fuzzy controller is a dependable option since it provides a decent balance between stability and responsiveness. With greater overshoot, slower response times, and larger steady-state errors than the other two controllers, the PID controller performs less well than the other two. From the obtained simulation results, we proposed a novel control approach for the FMWMR model in a real-life system. The experimental results will be discussed in detail in the following section. 5.2. Experimental Results The superior performance of the Fuzzy PID controller is demonstrated by the simulation results. In this subsection, this controller will be employed
VOLUME 20,
N° 3
2026
and experimental results will be presented under various conditions. The performance of the control system will be investigated in different scenarios, including desired position, elliptical trajectory, rectangular trajectory, circular trajectory, and eight-shaped path trajectory. Furthermore, the robustness and tracking quality of the controller will be verified under a loaded condition of 2.5 kg. The examination of the control system’s tracking accuracy and path-following capability will be prioritized, with the unloaded condition applied specifically to the evaluation of the circular trajectory and eight-shaped path trajectory scenarios. 5.2.1. Desired position control scenario In this experiment, the robot is set to move to the target position with coordinates x = 30cm, y = 30cm, and a yaw angle of 60 degrees. The control results are illustrated in Figure 24 to Figure 27. Figure 24 depicts the robot’s positional response in terms of x
Figure 24. Responses of robot with respect to target setpoint
Figure 25. Errors of robot with respect to target setpoint
75
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 26. PID Gains tuning results with respect target setpoint in x-dimension
Figure 27. PID Gains tuning results with respect to target setpoint y-dimension and y coordinates as well as yaw angle. It is evident that the robot reaches the target position within three seconds without any overshoot, while the yaw angle takes six seconds to stabilize. The steady-state error is nearly zero. The gains of the PID controllers are adjusted online and visualized in Figure 26 and Figure 27. Additionally, the effect of a 2.5kg load is examined in this scenario, revealing that the designed controller can adapt effectively to changes in load. The control performance remains consistent with that of the unloaded condition. 5.2.2. Elliptical trajectory tracking control scenario The designed controller is then put to the test to guide the robot in tracking an elliptical trajectory. The elliptical path is configured to be 50cm long and 30cm wide along the x-y directions, with a heading angle set at -45 degrees. The results of this control exper76
iment are presented in Figure 28 through Figure 32. Figure 28 illustrates the tracking performance within the x-y plane, showcasing effective trajectory tracking with minimal error. Detailed control responses for the x, y positions, and yaw angle are depicted in Figure 30, demonstrating consistent performance comparable to the initial scenario. Additionally, the tuning results for the PID controllers are visualized in Figure 31 and Figure 32. 5.2.3. Rectangular trajectory scenario The final case examined under two loading conditions involves tracking a rectangular trajectory. In this experiment, the robot is configured to follow a rectangular path with dimensions of 80 cm in length along the x-direction and 60 cm in width along the y-direction, with a desired yaw angle set at 20 degrees. The control results are presented in
18 19 20 21 22 23 24 25 26
27 28 29 30 31 32 33 34
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 28. Trajectory tracking responses with respect to elliptical trajectory in x-y plant
Figure 29. Robot position responses with respect to elliptical trajectory
Figure 30. Robot position tracking errors with respect to elliptical trajectory 77
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 31. PID gains tuning results with respect to elliptical trajectory in x dimension
Figure 32. PID gains tuning results with respect to elliptical trajectory in y dimension
78
Figure 33 rough Figure 35. These results demonstrate that the controller effectively guides the robot to track the desired trajectory, exhibiting minimal steady-state error and low overshoot. It’s notable that in this scenario, there is some overshoot due to sudden changes in the desired x-y direction, causing the robot, with its inertia, to continue moving in a certain direction. In summary, to quantitatively assess the performance of the control system, the root mean square error (RMSE) values in the x, y, and yaw dimensions are calculated. Table 6 displays the RMSE values for different scenarios under both unload and load conditions. From the RMSE table, it is observed that the RMSE values for x and y dimensions in rectangular trajectory tracking are the largest, attributed to sudden changes in the desired trajectory. However, the RMSE values under load conditions do not show significant differences. This indicates that the designed controller effectively handles load conditions.
5.2.4. Circular trajectory and eight-shape trajectory scenarios Moreover, the designed controller is evaluated for its ability to guide the robot along circular and eight-shaped trajectories. The control outcomes are illustrated in Figure 36 through Figure 41. Figure 36 to Figure 38 depict the tracking control results for the circular trajectory, while Figure 39 to Figure 41 display the results for the eight-shaped trajectory. Consistent with previous scenarios, the robot adeptly follows the circular path with commendable accuracy, maintaining a consistent shape throughout the trajectory. Overall, across various scenarios, the controller exhibits consistent performance and robustness, demonstrating its ability to adapt to different trajectories and load conditions.
5.3. Discussion The accuracy and robustness of the Fuzzy-PID controller for four wheels mecanum robot trajectory
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 33. Trajectory tracking responses with respect to rectangular trajectory in x-y plant
Figure 34. Robot position responses with respect to rectangular trajectory
Figure 35. Robot position tracking errors with respect to rectangular trajectory 79
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
Table 6. RMSE values of tracking error
RMSE value ex RMS (cm ) e y RMS (cm ) e yawRMS (degree)
Target Setpoint
Elliptical Trajectory Without Load
With Load
Rectangular Trajectory Without Load
With Load
Without Load
With Load
0.5859
0.7273
0.9443
0.8152
1.4152
1.421
0.3483
0.5684
0.5421
0.5366
1.1642
1.1629
1.583
1.8294
1.163
2.0261
0.4476
0.7072
Figure 36. Trajectory tracking response with respect to circular trajectory in x-y plant
Figure 37. x-y positions and yaw angle tracking control responses with respect to circular trajectory
80
2026
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 38. Tracking errors with respect to circular trajectory
Figure 39. Trajectory tracking response with respect to eight-shape trajectory in x-y plant
Figure 40. x-y positions and yaw angle tracking control with respect to eight-shape trajectory 81
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 41. Tracking errors with respect to eight-shape trajectory tracking under numerous scenarios and loading conditions are shown by the simulation and experimental results reported in this paper. The results of the simulation demonstrate that when it comes to providing accurate and smooth tracking of complex trajectories, the fuzzy-PID controller outperforms the PID and fuzzy controllers. The ability of the Fuzzy-PID controller to combine the advantages of fuzzy logic with PID control accounts for its higher tracking accuracy when compared to the standalone Fuzzy and PID controllers. The proposed control system’s performance in different trajectories, including desired position, elliptical, rectangular, circular, and eight-shaped path, is further validated by the experimental results. Excellent tracking precision, resilience, and adaptability are demonstrated by the control system in both loading conditions. Plots for trajectory tracking, response, error, and PID gain reveal the system’s capacity to maintain accurate tracking, stable performance, and smooth motion even when additional payloads are present. A quantitative assessment of the tracking accuracy of the control system and the effect of the load on its operation is given by the RMSE analysis. The findings demonstrate that, despite a slight decrease in tracking accuracy brought about by the increased load, the control system is still able to maintain comparatively low RMSE values, guaranteeing accurate trajectory tracking and positional control. The load has varying effects on the different dimensions; the y-dimension is most affected, followed by the x- and yaw-dimensions. For mobile robots to be implemented in a variety of industrial and service sectors where payload weight frequently varies, the control system must be able to maintain consistent performance under changing load situations.
6. Conclusion
82
In this paper, a novel control approach of combining a Fuzzy-PID controller for position control and a Fuzzy controller for heading control has been
proposed and investigated for the FMWMR. Moreover, a sensor fusion technique, which combines data from the camera and IMU, is also implemented for enhancing heading angle estimation. The control system’s outstanding performance in providing precise and smooth tracking of various trajectories under different loading cases has been proved by the results of simulations and experiments. Because of its versatility and resilience while managing changes in payload, the system has great promise for the real-world implementation of mobile robots in industrial and service applications. Our research’s finding advances the development of mobile robot control algorithms by leveraging the benefits provided by multi-sensor data fusion technique. List of abbreviations FMWMR: PID: GPS: IMU: COG: Declarations
Four Mecanum-Wheeled Mobile Robots Proportional Integral Derivative Global Positioning System Inertial Measurement Unit Center Of Gravity
Availability of data and materials The authors confirm that the data supporting the findings of this study are available within the article and its Supplementary material. Raw data that support findings of this study are available from the corresponding author, upon reasonable request. Competing interests
Not applicable
Funding This research is funded by Vietnam National University Ho Chi Minh City (VNU-HCM) under grant number C2021-20-15. Authors’ contributions Trong-Tai Nguyen conceived of the presented idea, developed the theory and performed the computations, prepared facilities for experimental, verified the analytical methods and results. Quang-Tho Le, Quang-Phuoc Pham,
Journal of Automation, Mobile Robotics and Intelligent Systems
Tran-Long Le conducted experimental results. All authors discussed the results and contributed to the final manuscript.
ACKNOWLEDGEMENTS
We acknowledge the support of time and facilities from Ho Chi Minh City University of Technology (HCMUT), VNU-HCM for this study. AUTHORS
Trong-Tai Nguyen* – Faculty of Electrical and Electronics Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam National University Ho Chi Minh City, Ho Chi Minh City, Vietnam, email: nttai@hcmut.edu.vn. Quang-Tho Le – Faculty of Electrical and Electronics Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam National University Ho Chi Minh City, Ho Chi Minh City, Vietnam, email: tholequang480@gmail.com. Quang-Phuoc Pham – Faculty of Electrical and Electronics Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam National University Ho Chi Minh City, Ho Chi Minh City, Vietnam, email: soattag@gmail.com. Tran-Long Le – Faculty of Electrical and Electronics Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam National University Ho Chi Minh City, Ho Chi Minh City, Vietnam, email: tranlong792001@gmail.com
*Corresponding author
References
[1] R. Galati, G. Mantriota, and G. Reina, “Adaptive Heading Correction for an Industrial HeavyDuty Omnidirectional Robot,” Sci Rep, vol. 12, no. 1, 2022, p. 19608. [2] D. Eck et al., “A Evaluation Test Bed for Outdoor Localization Algorithms Using a High-Precision Positioning System,” IFAC-PapersOnLine, vol. 48, no. 10, 2015, pp. 34-40. [3] I. Doroftei and B.-d.D. Mangeron, “Practical Applications for Mobile Robots Based on Mecanum Wheels - A Systematic Survey,” 2011. [4] J.P. Trevelyan, S.-C. Kang, and W. R. Hamel, “Robotics in Hazardous Applications,” Springer Handbook of Robotics, B. Siciliano and O. Khatib, eds., Springer Berlin Heidelberg, 2008, pp. 1101–1126. [5] N. N, B.U. Balappa, and S. Ravichandran, “Design and Development of Vision Based Mobile Robot for Warehouse Application,” 2023 IEEE North Karnataka Subsection Flagship International Conference (NKCon), 2023, pp. 1–6. [6] O. Buckmann, M. Krö� mker, and U. Berger, “An Application Platform for the Development and Experimental Validation of Mobile Robots for Health Care Purposes,” Journal of Intelligent and Robotic Systems, vol. 22, no. 3, 1998, pp. 331–350. [7] F. Adă� scă� liţei and I.J.T.R.R.P.M. Doroftei, “Practical Applications for Mobile Robots based on Mecanum Wheels-A Systematic Survey,” vol. 40, 2011, pp. 21–29.
VOLUME 20,
N° 3
2026
[8] O. Setiadilaga, A. Cahyadi, and A. Ataka, “Mecanum-Wheeled Robot Control based on Deep Reinforcement Learning,” 2023 15th International Conference on Information Technology and Electrical Engineering (ICITEE), 2023, pp. 25–30. [9] J. Leng et al., “Design, Modeling, and Control of a New Multi-Motion Mobile Robot based on Spoked Mecanum Wheels,” Biomimetics, vol. 8, no. 2, 2023, p. 183. [10] C.-H. Kuo, “Trajectory and Heading Tracking of a Mecanum Wheeled Robot using Fuzzy Logic Control,” 2016 International Conference on Instrumentation, Control and Automation (ICA), 2016, pp. 54–59. [11] G. Bayar and S. Ozturk, “Investigation of the Effects of Contact Forces Acting on Rollers of a Mecanum Wheeled Robot,” Mechatronics, vol. 72, 2020, p. 102467. [12] T.D. Nguyen, D.P. Dinh, and T.H. Tran, “Kinematic Modeling and Stable Control Law Designing for Four Mecanum Wheeled Mobile Robot Platform Based on Lyapunov Stability Criterion,” Journal of Technical Education Science, vol. 18, no. 5, 2023, pp. 57–65. [13] S. Mellah et al., “Trajectory Reconfiguration for Time Delay Reduction in the Case of Unexpected Obstacles: Application to 4-Mecanum Wheeled Mobile Robots (4-MWMR) for Industrial Purposes,” IFAC-PapersOnLine, vol. 53, no. 2, 2020, pp. 15653–15658. [14] P. Manzl, M. Sereinig, and J. Gerstmayr, “A Mecanum Wheel Model based on Orthotropic Friction with Experimental Validation,” Mechanism and Machine Theory, vol. 193, 2024, p. 105548. [15] N.H. Thai, T.T.K. Ly, and L.Q. Duzung, “Trajectory Tracking Control for Mecanum Wheel Mobile Robot by Time-Varying Parameter PID Controller,” Bulletin of Electrical Engineerign and Informatics, vol. 11, no. 4, 2022, pp. 1902– 1910, 2022. [16] G. Cao et al., “Fuzzy Adaptive PID Control Method for Multi-Mecanum-Wheeled Mobile Robot,” Journal of Mechanical Science and Technology, vol. 36, no. 4, 2022, pp. 2019–2029. [17] B. Wen et al., “Research on Trajectory Tracking Control Method of Omnidirectional Moving Platform based on Backstepping,” Journal of Physics: Conference Series, vol. 1828, no. 1, 2021, p. 012141. [18] A. Bessas, A. Benalia, and F. Boudjema, “Integral Sliding Mode Control for Trajectory Tracking of Wheeled Mobile Robot in Presence of Uncertainties,” J Control Sci Eng, 2016. [19] M. Szeremeta and M.J.A.S. Szuster, “Neural Tracking Control of a Four-Wheeled Mobile Robot with Mecanum Wheels,” vol. 12, no. 11, 2022, p. 5322.
83
Journal of Automation, Mobile Robotics and Intelligent Systems
84
[20] T.T.K. Ly et al., “A Neural Network Controller Design for the Mecanum Wheel Mobile Robot,” Engineering, Technology & Applied Science Research, vol. 13, no. 2, 2023, pp. 10541–10547. [21] S.A. Ahmed and M.G. Petrov, “Trajectory Control of Mobile Robots using Type-2 Fuzzy-Neural PID Controller,” IFAC-PapersOnLine, vol. 48, no. 24, 2015, pp. 138-143. [22] Q. Jia et al., “Motion Control of Omnidirectional Mobile Robot Based on Fuzzy PID,” 2019 Chinese Control And Decision Conference (CCDC), 2019, pp. 5149-5154. [23] I. Sajad Ahmad Wani et al., “Motion Control of Autonomous Mobility System Using Fuzzy-PID Controller,” Tuijin Jishu/Journal of Propulsion Technology, vol. 45, no. 2, 2024, pp. 5402–5416. [24] H. Chen, “Fuzzy PID-based Temperature Control Method for Power Transformer Coils,” International Journal of Energy Technology and Policy, vol. 19, no. 1-2, 2024, pp. 86–104. [25] G. Ji et al., “Research on Variable Universe Fuzzy PID Control for Semi-Active Suspension with CDC Dampers Based on Dynamic Adjustment Functions,” Scientific Reports, vol. 14, no. 1, 2024, p. 3442. [26] B. Zhang et al., “Fuzzy PID Control of Permanent Magnet Synchronous Motor Electric Steering Engine by Improved Beetle Antennae Search Algorithm,” Scientific Reports, vol. 14, no. 1, 2024, p. 2898. [27] Y.-H. Chen and Y.-Y. Chen, “Nonlinear Adaptive Fuzzy Control Design for Wheeled Mobile Robots with using the Skew Symmetrical Property,” Symmetry, vol. 15, no. 1; doi: 10.3390/ sym15010221 [28] M.B. Alatise and G.P.J.S. Hancke, “Pose Estimation of a Mobile Robot based on Fusion of IMU Data and Vision Data using an Extended Kalman Filter,” vol. 17, no. 10, 2017, p. 2164. [29] A. Noordin, M. Basri, and Z.J.T. Mohamed, “Sensor Fusion Algorithm by Complementary Filter for Attitude Estimation of Quadrotor with Low-Cost IMU,” vol. 16, no. 2, 2018, pp. 868–875. [30] F. Guangrui and W. Geng, “Vision-based Autonomous Docking and Re-charging System for Mobile Robot in Warehouse Environment,” 2017 2nd International Conference on Robotics and Automation Engineering (ICRAE), 2017, pp. 79-83. [31] J. Huang et al., “Indoor Positioning Systems of Mobile Robots: A Review,” Robotics, vol. 12, no. 2; doi: 10.3390/robotics12020047 [32] H. Di and Z. Chu, “Design of Indoor Mobile Robot based on ROS and lidar,” Proceedings of the 2022 2nd International Conference on Robotics and Control Engineering, Nanjing, China, 2022; doi:10.1145/3529261.3529272 [33] R. Raj and A. Kos, “A Comprehensive Study of Mobile Robot: History, Developments,
VOLUME 20,
N° 3
2026
Applications, and Future Research Perspectives,” Applied Sciences, vol. 12, no. 14; doi: 10.3390/ app12146951 [34] P. Tripicchio, S. D’Avella, and M. Unetti, “Efficient Localization in Warehouse Logistics: A Comparison of LMS Approaches for 3D Multilateration of Passive UHF RFID tags,” The International Journal of Advanced Manufacturing Technology, vol. 120, no. 7, 2022, pp. 4977–4988. [35] Y. Liu et al., “A Review of Sensing Technologies for Indoor Autonomous Mobile Robots,”, Sensors (Basel), vol. 24, no. 4, 2024. [36] B. Alsadik, “Chapter 10 - Kalman Filter,” Adjustment Models in 3D Geomatics and Computational Geophysics, B. Alsadik, ed., Elsevier, vol. 4, 2019, pp. 299–326. [37] M.K. Hoang et al., “A Hidden Markov Model for Indoor User Tracking based on WiFi Fingerprinting and Step Detection,” 21st European Signal Processing Conference (EUSIPCO 2013), 2013, pp. 1–5. [38] X. Ren et al., “An Improved Hidden Markov Model for Indoor Positioning,” Communications and Networking, Springer Natrure, 2023, pp. 403–420. [39] A.I. Sudianto, M.A. Muslim, and M. Rusli, “Sensor Fusion using Model Predictive Control for Differential Dual Wheeled Robot,” Kinetik: Game Technology, Information System, Computer Network, Computing, Electronics, and Control, vol. 8, no. 1, 2023, pp. 461–472. [40] V.L. Popov et al., “Detection and Following of Moving Target by an Indoor Mobile Robot using Multi-sensor Information,” IFAC-PapersOnLine, vol. 54, no. 13, 2021, pp. 357–362. [41] A.T. Erdem and A.Ö� .J.I.T.O.I.P. Ercan, “Fusing Inertial Sensor Data in an Extended Kalman Filter for 3D Camera Tracking,” vol. 24, no. 2, 2014, pp. 538–548. [42] A. Cavallo et al., “Experimental Comparison of Sensor Fusion Algorithms for Attitude Estimation,” vol. 47, no. 3, 2014, pp. 7585–7591. [43] S.O. Madgwick, A.J. Harrison, and R. Vaidyanathan, “Estimation of IMU and MARG Orientation using a Gradient Descent Algorithm,” 2011 IEEE international conference on rehabilitation robotics, 2011, pp. 1–7. [44] D. Parikh, S. Vohra, and M. Kaveshgar, “Comparison of Attitude Estimation Algorithms with IMU Under External Acceleration,” 2021 IEEE International Symposium on Smart Electronic Systems (iSES), 2021, pp. 123–126. [45] R. Mahony, T. Hamel, and J.-M.J.I.T.o.a.c. Pflimlin, “Nonlinear Complementary Filters on the Special Orthogonal Group,” vol. 53, no. 5, 2008, pp. 1203–1218. [46] J. Zheng et al., “Visual Localization of Inspection Robot using Extended Kalman Filter and Aruco
Journal of Automation, Mobile Robotics and Intelligent Systems
Markers,” 2018 IEEE International Conference on Robotics and Biomimetics (ROBIO), 2018, pp. 742–747. [47] S. Roos-Hoefgeest, I.A. Garcia, and R.C. Gonzalez, “Mobile Robot Localization in Industrial Environments using a Ring of Cameras and ArUco Markers,” IECON 2021 – 47th Annual
VOLUME 20,
N° 3
2026
Conference of the IEEE Industrial Electronics Society, 2021, pp. 1–6. [48] Z. Yuan et al., “Trajectory Tracking Control of a Four Mecanum Wheeled Mobile Platform: An Extended State Observer-based Sliding Mode Approach,” IET Control Theory & Applications, vol. 14, no. 3, 2020, pp. 415–426.
85
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
VISION-BASED GESTURE-CONTROLLED MOBILE ROBOT FOR HUMAN–ROBOT INTERACTION USING ROS 2 Submitted: 24th March 2025; accepted: 9th May 2025
Kaveh Hooshmandi, Mahdi Molavi DOI: 10.14313/jamris-2026-040 Abstract: This paper presents the design and implementation of a gesture-controlled differential-drive mobile robot for human–robot interaction (HRI), using the Robot Operating System(ROS 2) framework. The proposed system introduces a novel, vision-based gesture recognition approach using a Raspberry Pi and an onboard camera, combined with a custom spatial AI algorithm for realtime interpretation of human hand gestures. Unlike conventional systems that rely on physical interfaces, the robot is capable of recognizing intuitive gestures—such as start, stop, and directional commands—to control its movement without any contact-based input. The gesture recognition module integrates lightweight machine learning models optimized for embedded deployment, ensuring accurate and low-latency classification of hand signals in dynamic environments. ROS 2 serves as the middleware for seamless integration of sensory data and control of the differential-drive mechanics, enabling the robot to perform core mobility tasks such as forward, backward, turning, and halting. Experimental evaluations in real-world settings demonstrate the system’s high responsiveness, robustness, and adaptability, underscoring its potential for natural, non-invasive interaction in service robotics, assistive applications, and collaborative human–robot environments. Keywords: Differential-drive Mobile Robot, HumanRobot Interaction, ROS 2, Gesture Recognition, Collaborative Robotics
1. Introduction The integration of robotics into everyday human activities has seen significant advancements in recent years, particularly in the domain of HRI [1, 2]. HRI focuses on the collaborative interactions between humans and robots, emphasizing the importance of intuitive communication methods. As robots are increasingly deployed in service roles, ranging from healthcare [3] to domestic assistance, the need for natural and efficient control mechanisms has become paramount. Gesture recognition, as a form of nonverbal communication, offers a promising solution, allowing users to control robots through simple hand movements without the need for physical interfaces [4]. 86
In the context of gesture recognition, various methodologies have been explored. Traditional approaches often rely on depth sensors and complex algorithms, which can be cumbersome and limited in dynamic environments. Recent advancements in machine learning and computer vision have facilitated the development of more sophisticated gesture recognition systems. Notably, the MediaPipe library has emerged as a powerful tool for real-time hand tracking and gesture recognition, enabling the extraction of key points from hand movements with high accuracy [5]. Deep learning techniques have been used to recognize hand gestures in real time for controlling wheeled robots [8]. Artificial neural networks have also been applied to detect hand gestures for smart home functions [9]. Machine learning methods, particularly MediaPipe, have been used for gesture-based computer mouse control [10]. Recent advances in hand gesture recognition have led to the development of more accurate and efficient models across diverse data modalities. A realtime online gesture recognition system using Continual Graph Transformers was proposed, achieving robust recognition with spatial and temporal modeling through S-GCN and a Transformer-based Graph Encoder, evaluated on SHREC’21 [15]. A comprehensive survey covering the period from 2014 to 2024 reviewed RGB, skeleton, depth, EMG, EEG, and multimodal data, identifying trends and challenges in continuous gesture recognition [16]. A hybrid architecture combining BiLSTM, metaheuristic optimization, and a U-Net-MobileNetV2 encoder was introduced for sEMG-based classification, achieving over 90% accuracy across six datasets [17]. Deep learning with EMG signals has also been used to enhance hand gesture recognition performance, highlighting its potential in prosthetics and assistive technologies [18]. Moreover, the TCNN-KAN model, integrating Kolmogorov-Arnold Networks and pruning techniques, was proposed to optimize CNN-based gesture recognition using semg, significantly improving accuracy and computational efficiency [19]. Collectively, these works demonstrate the potential of hybrid deep learning models and advanced optimization techniques to improve gesture recognition performance across modalities and real-world applications. Furthermore, the implementation of machine learning techniques, particularly Support Vector
Open Access. © 2026 Kaveh Hooshmandi and Mahdi Molavi, published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
Journal of Automation, Mobile Robotics and Intelligent Systems
Machines (SVM), has proven effective for gesture classification tasks. SVM models excel in highdimensional spaces, making them well-suited for applications involving complex feature sets derived from hand gestures [6, 7]. The use of these techniques not only enhances recognition accuracy but also contributes to the seamless integration of HRI systems within robotic frameworks. The Robot Operating System (ROS) has significantly influenced the development of robotics applications by providing a flexible framework for robot software development [11]. The introduction of ROS 2 has further improved capabilities, offering enhanced communication protocols and support for real-time systems. This has facilitated the integration of various sensors and control algorithms, thus streamlining the development of intelligent robotic systems capable of real-time interaction with humans [12]. The incorporation of ROS 2 in the design of robotic systems has been thoroughly explored in the literature, emphasizing its advantages in modularity and real-time communication [13]. Studies have demonstrated how ROS 2 can facilitate the integration of machine learning models with robotic control systems, enabling autonomous decision-making based on real-time data inputs [14]. This paper presents the design and implementation of a gesture-controlled differential-drive mobile robot specifically developed for gesture-based HRI, utilizing the advanced capabilities of the ROS 2 framework. At the core of this system is a novel, robust, and optimized vision-based gesture recognition module, which distinguishes this work from conventional implementations. As robots increasingly integrate into service environments, healthcare, and collaborative workspaces, there is a growing demand for interaction methods that are intuitive, natural, and non-invasive. Traditional control mechanisms—such as remote controllers or touch interfaces—often require users to engage with physical devices, which can be limiting or impractical in dynamic and hands-busy scenarios. To address this challenge, the present study explores a vision-based gesture control approach that enables seamless, contactless communication between humans and robots. The system utilizes a Raspberry Pi as the embedded processing unit, interfaced with a camera that continuously captures visual data. A custom-designed spatial artificial intelligence algorithm processes these inputs in real time to identify a set of predefined hand gestures—including commands such as start, stop, forward, backward, left, and right. Unlike many prior works that rely on heavy models or controlled environments, the proposed method incorporates lightweight machine learning models specifically optimized for real-time inference on resource-constrained embedded hardware. These models are trained and fine-tuned to maintain high classification accuracy while ensuring low computational overhead, enabling robust gesture recognition in diverse, unpredictable indoor environments.
VOLUME 20,
N∘ 3
2026
The robustness of the system is achieved through several key innovations. Firstly, the gesture recognition pipeline is designed to be resilient to variable lighting conditions and background clutter, making it suitable for deployment in realworld, uncontrolled settings. Secondly, the algorithm integrates temporal filtering and noise-reduction techniques to prevent false positives and enhance consistency in gesture detection. Thirdly, the modularity of the recognition engine allows easy adaptation to new gestures or environments with minimal retraining. ROS 2 plays a central role in this architecture by acting as a middleware platform that ensures reliable communication between the perception module, decision-making logic, and motor control systems. Its support for real-time communication, distributed computing, and modular development enables seamless integration of the vision system with the differential-drive control, allowing the robot to execute basic motion commands such as moving forward, reversing, turning, and stopping with high precision. Comprehensive testing was conducted in indoor settings featuring dynamic lighting and environmental conditions to evaluate the system’s performance. Results demonstrated strong accuracy, responsiveness, and adaptability, confirming the effectiveness of the gesture-based control method for real-world HRI scenarios. The successful fusion of embedded AI, computer vision, and robotics middleware underscores the feasibility of creating low-cost, contactless, and user-friendly robotic systems. By synthesizing these advancements, this research contributes a practical and scalable solution to the field of intuitive human–robot interaction, laying the groundwork for future developments in gesture-controlled service robotics, assistive technologies, and collaborative autonomous systems.
2.
Hardware Architecture
To design the robot’s chassis, the various components were first sketched on paper. After making some adjustments, the final design was simulated using SOLIDWORKS, preparing it for laser cutting or 3D printing. The assembly process involved drilling and cutting in certain sections of the chassis. For example, the chassis did not initially account for space to accommodate the DC geared motor with an encoder (YGY6138-R528C, 12V, 65RPM). Therefore, a section was cut using a grinder to fit the motor. The motor has a voltage of 12V DC, operates at a current of 3A, features a 6mm shaft, provides a torque of 5.5Kg/cm, and has a no-load speed of 65RPM. The chassis of the robot is composed of four parts. The first part is the main chassis where all components are mounted. The second part relates to the suspension system, which supports the wheels and ensures smooth movement. The third part comprises the outer body surrounding the robot, and the fourth part is the top section, which holds the load, the camera, and the LiDAR sensor. 87
Journal of Automation, Mobile Robotics and Intelligent Systems
In this paper, a Raspberry Pi 4 board is used. This board features a Broadcom BCM2711 chip and a quadcore Cortex-A72 processor running at 1.5 GHz, with 4 GB of memory. Additionally, an 8-megapixel V2 camera module is used, which supports both Raspberry Pi and Jetson Nano boards. This module, equipped with a Sony IMX219 image sensor, is lightweight, compact, and easy to install, making it ideal for many applications. Two wheels with a diameter of 7 cm and a thickness of 0.9 cm are used for the robot’s movement. To supply the necessary voltage for the stepper motor, two 6V, 4.5Ah UPS batteries (EURONET model EUR456) are utilized. A 10,000 mAh power bank is also used to power the Raspberry Pi. A DC buck converter module is employed to reduce the voltage and control the current, with a maximum output of 5A. One potentiometer adjusts the output voltage, while the other regulates the current. Figure 1 shows the robot components, and Figure 2 illustrates the fully assembled system. The robot is equipped with two independent motors, enabling it to achieve linear velocity 𝑣 and angular velocity 𝜔 around its axis. These velocities are influenced by the wheel diameter and the distance between the wheels. By utilizing the angular velocities of the left and right wheels, 𝜔𝐿 and 𝜔𝑅 , the equations of motion are derived to control the robot’s movements. This ensures that the robot receives only the necessary speed inputs to perform its movements accurately. The relationship between the velocities is
Figure 1. Robot components used in the proposed mobile robot platform
VOLUME 20,
N∘ 3
2026
given by the following equations: 2𝑣 + 𝜔𝐿 , 𝑣𝑅 2𝑟 2𝑣 − 𝜔𝐿 𝜔𝐿 = , 𝑣𝐿 2𝑟
𝜔𝑅 =
= 𝑟𝜔𝑅 (1) = 𝑟𝜔𝐿
where 𝑣𝑅 and 𝑣𝐿 are the linear velocities of the right and left wheels respectively, 𝜔𝑅 and 𝜔𝐿 are the angular velocities of the right and left wheels respectively, 𝑟 is the radius of the wheels, and 𝐿 is the distance between the left and right wheels.
3.
Data Collection and Processing
In this section, a webcam was used to capture images for hand gesture recognition. Hand movement data was extracted using the MediaPipe library. For the predefined gestures shown in Figure 3, such as forward, backward, left, right, and stop, images and the corresponding hand key-point features were collected.Each gesture included a specific number of samples (in this case, 150 samples) recorded using the webcam and hand detection methods. After data collection, the data processing stage begins. The collected data consists of the coordinates of hand key points extracted by MediaPipe, along with labels corresponding to each gesture. These labels categorize the data into their respective classes, helping to define the meaning of each sample. The data is then transformed into a DataFrame, where each row represents a sample of hand movements and its associated features. This transformation is necessary to structure the data in a format that can be understood by the machine learning algorithm. During this stage, data quality is also assessed, and the accuracy of the labels is verified, as any errors in the data will directly lead to inaccuracies in the learning model. To evaluate the data, a scatter plot is used. By placing data points in a two-dimensional space, the scatter plot allows researchers to observe the distribution of data and examine relationships between variables. It is an efficient tool for visualizing potential patterns, identifying anomalies, and detecting relationships between features and labels. In machine learning projects, examining data features through scatter plots helps researchers understand the distribution and correlation of features before proceeding with model development. This initial understanding of the data is crucial for building accurate and effective models. Figure 4 shows the scatter plot of the collected data.
4.
Modeling and Training
After creating the Data Frame, the data is divided into three main sets: - Training Set: Used to train the model. - Validation Set: Used for tuning the model’s hyperparameters and preventing overfitting.
Figure 2. Fully assembled differential-drive mobile robot system 88
- Test Set: Used for the final evaluation of the model after training and validation. The dataset was divided into training, validation, and test sets using a 60:20:20 ratio. It’s important to note
Journal of Automation, Mobile Robotics and Intelligent Systems
(a) Right
(b) Left
VOLUME 20,
(d) Forward
Figure 3. Hand gestures for controlling the system: Right, Left, Stop, and Forward
2026
that typically, data is split into two sets, training and test. However, this method may not provide an accurate assessment of the model’s performance. A separate validation set, containing data the model hasn’t seen before, offers a better evaluation of its accuracy. In this project, the train_test_split function from the sklearn library is used to split the data into specific proportions. This ensures that the label distribution is consistent across all three sets. The model used in this project is a Support Vector Machine (SVM), designed for hand movement classification. This model is trained using features extracted from the data. The training process includes utilizing the fit method for the SVM model with training data. The model’s hyperparameters, such as the kernel and the C parameter, are also tuned during this stage. SVM has gained popularity due to its powerful data separation capabilities and effective performance, especially in high-dimensional datasets. SVM operates based on the concept of the maximum margin. The goal of this algorithm is to create a line (in two dimensions) or a hyperplane (in higher dimensions) that best separates the different data points. This hyperplane is selected in such a way that the distance between it and the closest data points, known as support vectors, is maximized. This feature enables SVM to generalize well to new data and prevents overfitting. SVM utilizes a kernel function to process data in higher dimensions. During the training phase, SVM computes the optimal hyperplane using labeled data. In the prediction phase, this hyperplane is used to classify new data. In addition to classification, SVM can also be applied to regression problems where the goal is to predict continuous values. In the present project, SVM is used to recognize hand movements and convert them into control commands for a robot. This algorithm accurately analyzes the data collected from hand movements, delivering precise results that are applied to control the robot in various environments.
5. (c) Stop
N∘ 3
Validation and Implementation
After training the model, its performance is evaluated on the validation set using the predict method. This method generates predicted labels for the validation data. For an initial assessment, the accuracy metric is used, which indicates the percentage of correct predictions. Accuracy is calculated using the accuracy score function. Although accuracy is an important metric, it is not sufficient on its own, especially when dealing with imbalanced data. To improve the evaluation of the model and to gain a better understanding of its performance across different categories, more advanced metrics such as the confusion matrix and F1Score are used. The confusion matrix is a key tool for evaluating classification models, allowing for a more precise examination of the model’s performance. This matrix is a square table where the rows represent the actual classes and the columns represent the predicted classes by the model. The confusion matrix includes four key pieces of information: 89
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 4. Scatter plot of the data collected • True Positive (TP): Samples correctly classified. • False Positive (FP): The number of samples incorrectly assigned to a class. • False Negative (FN): Samples wrongly classified into another category. • True Negative (TN): Samples correctly identified as belonging to other classes. This matrix is particularly useful for analyzing the model’s errors. High FP or FN values in specific areas highlight the model’s weaknesses in classification. The confusion matrix is also used to compute important evaluation metrics such as Precision, Recall, and F1Score, which provide a more detailed analysis of the
model’s performance. For instance, if the model performs well in recognizing gestures like moving forward but struggles with identifying gestures such as turning left, these issues will be visible as FP or FN errors in the matrix. This powerful tool helps improve the model by allowing us to analyze its mistakes. In Figure 5, the confusion matrix of the collected data model is presented. 5.1.
Model evaluation
The F1-score is a fundamental metric for evaluating the performance of classification models, particularly when the data is imbalanced or when both precision and recall are equally important. The F1Score is the harmonic mean of Precision and Recall, ensuring that if either metric is low, the F1-Score also decreases. Precision indicates the percentage of instances correctly assigned to a specific class (of the predicted samples in that class, how many truly belong to it). Recall shows how many of the actual instances of a class were correctly identified by the model. These metrics together provide a comprehensive view of Table 1. Classification performance of the SVM on the held-out test set
Figure 5. Confusion matrix of the model 90
Class Backward Forward Left Right Stop Accuracy Macro avg. Weighted avg.
Precision 0.98 0.83 0.95 0.94 1.00 – 0.94 0.93
Recall 0.68 1.00 0.92 0.96 0.97 – 0.91 0.93
F1-score 0.81 0.91 0.94 0.95 0.98 0.93 0.92 0.93
Support 95 180 171 179 131 756 756 756
Journal of Automation, Mobile Robotics and Intelligent Systems
the model’s performance and help identify areas for improvement, particularly in imbalanced datasets. The F1-Score is calculated using the following formula: 𝐹1 = 2 ×
Precision × Recall Precision + Recall
(2)
The F1-Score is particularly valuable in situations where both False Positive (FP) and False Negative (FN) errors are equally important. For instance, in a gesture recognition system, if the model incorrectly classifies the gesture ”turning left” as ”moving forward” (False Positive), or fails to correctly identify a ”turning left” gesture (False Negative), both errors can negatively affect the overall system performance. The F1-Score helps evaluate this balance, ensuring that the model not only has high precision but also maintains good recall. Therefore, the F1-Score is especially useful in scenarios where there is an imbalance in class distribution, or when both types of errors (FP and FN) are equally significant. This metric offers a more comprehensive view of the model’s performance, particularly when the focus is on balancing precision and recall, helping to identify and improve the model’s weaknesses. Table 1 presents the F1-score of the model. 5.2.
Implementation and discussion
In the final stage, a ROS 2 node was developed to classify hand gestures and convert them into motion commands for the robot. This node uses the trained model to interpret user commands and execute the corresponding actions. The node uses the trained machine-learning model to classify hand gestures and enables the robot to respond to the corresponding commands. This process integrates image processing and machine learning technologies, combining precise gesture recognition with robot control. The node, referred to as ”Hand Gesture”, is developed in ROS 2 and its main function is to predict hand gestures through visual data captured by a webcam. It uses the MediaPipe library to process images and detect key points on the hands. MediaPipe automatically identifies critical hand landmarks (such as fingertips), and these points are passed as input features to the machine learning model. Once the data is processed, the pre-trained machine learning model makes predictions based on the detected gestures. These predictions correspond to five main commands: ”Forward”, ”Back”, ”Left”, ”Right”, and ”Stop”. In the image processing workflow, video frames captured by the camera are first converted to RGB format and then analyzed. For each hand detected in the image, MediaPipe provides a set of key points, represented as (x, y) coordinates in the image space. Since hand gestures involve specific movements that alter the position of these hand points, this information is used effectively to identify different movement commands. This node simultaneously tracks both hands, crucial for simulating human-like natural robot movement (such as steering with two hands). After identifying the fingertips, the distance between them is calculated,
VOLUME 20,
N∘ 3
2026
and a large circle is drawn on the image as a ”steering wheel,” visually showing the user that the robot is receiving motion commands. The machine learning model used is SVM, which acts as a classifier for predicting hand gestures based on extracted hand features. The training data used to build the model consists of features representing the position of various hand points, predicting whether a specific gesture like ”Forward” or ”Stop” is being performed. After training, the model is saved as a pickle file and loaded into the node. The model achieved a training accuracy of 0.9497, with both validation and test accuracies equal to 0.9286. It is important to note that since an SVC was used, the model performs global optimization over the entire dataset rather than using iterative updates as in neural networks. As a result, the number of epochs is not applicable in this context. Moreover, the internal loss function of the SVC is not accessible or stored during training, and therefore, it is not possible to report a loss curve or stage-wise average loss. As soon as new hand features are received, the model makes predictions that are converted into commands for the robot. The model automatically predicts one of the five commands, and based on these predictions, the robot’s linear and angular velocity is updated. For instance, the ”Forward” gesture increases the robot’s linear speed, while the ”Stop” gesture results in a full stop. After the gestures are predicted, they are translated into motion commands for the robot. For example, the ”Forward” prediction increases the robot’s linear velocity, while the ”Left” or ”Right” predictions adjust its angular velocity. These commands are sent from the node through Twist messages, which set the robot’s linear (linear.x) and angular (angular.z) velocity values essential for movement. Finally, the robot can respond in real-time, moving according to the user’s hand movements. The node operates continuously, processing new frames and data and sending appropriate commands to the robot. Under the tested indoor lighting conditions, the system recognized the defined gestures at distances of up to 1.8 m. This means it works well in typical room lighting conditions and does not require special lighting or very close proximity to function effectively.
6.
Conclusion
This paper has presented the design and implementation of a gesture-controlled differentialdrive mobile robot utilizing the ROS 2 framework for gesture-based control in HRI. The integration of machine learning techniques, specifically through gesture recognition using the MediaPipe library, has enabled intuitive, real-time control of the robot without the need for physical interfaces. Through extensive testing, the proposed system has demonstrated its potential to enhance natural, non-invasive interactions in service robotics and collaborative environments. The use of ROS 2 not only provided a flexible and modular framework for developing the robotic system but also facilitated 91
Journal of Automation, Mobile Robotics and Intelligent Systems
seamless real-time communication, crucial for effective HRI. This research contributes to the growing field of intuitive robot control systems, providing a robust foundation for future work. Potential areas for further development include optimizing the system for use in dynamic and complex environments, improving gesture recognition accuracy, and expanding its application to other HRI scenarios such as healthcare and industrial settings. The current implementation uses discrete gesture categories mapped to fixed motion commands, including predefined speed and turning angles. However, this could be extended in future work to incorporate gesture intensity or hand movement dynamics for variable-speed control. Additionally, a suggestion is made to integrate a steering wheel interface to improve ergonomics and realism, providing users with a more familiar and physically engaging control option.
AUTHORS Kaveh Hooshmandi∗ – Department of Electrical Engineering, Arak, Iran, e-mail: k.hooshmandi@arakut.ac.ir, Arak University of Technology. Mahdi Molavi – Department of Electrical Engineering, Arak, Iran, e-mail: mahdimolavi01@gmail.com, Arak University of Technology. ∗
Corresponding author
References [1] T. B. Sheridan, “Human–robot interaction: status and challenges,” Human Factors, vol. 58, no. 4, pp. 525–532, 2016. [2] C. Bartneck, T. Belpaeme, F. Eyssel, T. Kanda, M. Keijsers, and S. Š abanović , Human-robot interaction: An introduction. Cambridge University Press, 2024. [3] T. Turja, T. Rantanen, and A. Oksanen, “Robot use self-efficacy in healthcare work (RUSH): development and validation of a new measure,” AI & Society, vol. 34, no. 1, pp. 137–143, 2019. [4] H. Liu and L. Wang, “Gesture recognition for human-robot collaboration: A review,” International Journal of Industrial Ergonomics, vol. 68, pp. 355–367, 2018. [5] S. Sreenath, D. I. Daniels, A. S. D. Ganesh, Y. S. Kuruganti, and R. G. Chittawadigi, “Monocular tracking of human hand on a smart phone camera using mediapipe and its application in robotics,” in 2021 IEEE 9th Region 10 Humanitarian Technology Conference (R10-HTC), 2021, pp. 1–6. [6] M. Soori, B. Arezoo, and R. Dastres, “Artificial intelligence, machine learning and deep learning in advanced robotics, a review,” Cognitive Robotics, vol. 3, pp. 54–70, 2023. 92
VOLUME 20,
N∘ 3
2026
[7] D. Kim, S.-H. Kim, T. Kim, B. B. Kang, M. Lee, W. Park, S. Ku, D. Kim, J. Kwon, H. Lee, et al., “Review of machine learning methods in soft robotics,” PLOS ONE, vol. 16, no. 2, p. e0246102, 2021. [8] T. B. Waskito, S. Sumaryo, and C. Setianingsih, “Wheeled robot control with hand gesture based on image processing,” in 2020 IEEE International Conference on Industry 4.0, Artificial Intelligence, and Communications Technology (IAICT), 2020, pp. 48–54. [9] P. N. Huu, Q. T. Minh, et al., “An ANN-based gesture recognition algorithm for smart-home applications,” KSII Transactions on Internet and Information Systems (TIIS), vol. 14, no. 5, pp. 1967–1983, 2020. [10] B. J. Boruah, A. K. Talukdar, and K. K. Sarma, “Development of a learning-aid tool using hand gesture based human computer interaction system,” in 2021 Advanced Communication Technologies and Signal Processing (ACTS), 2021, pp. 1–5. [11] S. Macenski, T. Foote, B. Gerkey, C. Lalancette, and W. Woodall, “Robot operating system 2: Design, architecture, and uses in the wild,” Science Robotics, vol. 7, no. 66, p. eabm6074, 2022. [12] S. Macenski, T. Moore, D. V. Lu, A. Merzlyakov, and M. Ferguson, “From the desks of ROS maintainers: A survey of modern & capable mobile robotics algorithms in the robot operating system 2,” Robotics and Autonomous Systems, vol. 168, p. 104493, 2023. [13] A. Bonci, F. Gaudeni, M. C. Giannini, and S. Longhi, “Robot Operating System 2 (ROS2)-based frameworks for increasing robot autonomy: A survey,” Applied Sciences, vol. 13, no. 23, p. 12796, 2023. [14] Y. Ye, Z. Nie, X. Liu, F. Xie, Z. Li, and P. Li, “ROS2 real-time performance optimization and evaluation,” Chinese Journal of Mechanical Engineering, vol. 36, no. 1, p. 144, 2023. [15] Slama, Rim, Wael Rabah, and Hazem Wannous. ”Online hand gesture recognition using Continual Graph Transformers.” arXiv preprint arXiv:2502.14939 (2025). [16] Shin, J., Miah, A. S. M., Kabir, M. H., Rahim, M. A., and Al Shiam, A. (2024). ”A methodological and structural review of hand gesture recognition across diverse data modalities.” IEEE Access. [17] Rezaee, Khosro, et al. ”Hand gestures classification of sEMG signals based on BiLSTM-metaheuristic optimization and hybrid U-Net-MobileNetV2 encoder architecture.” Scientific Reports, vol. 14, no. 1, 2024, p. 31257. [18] Abdelaziz, Mai H., Wael A. Mohamed, and Ayman S. Selmy. ”Hand gesture recognition based on electromyography signals and deep learning techniques.” Journal of Advances in Information Technology, vol. 15, no. 2, 2024.
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
[19] Al-Qaness, Mohammed A. A., and Sike Ni. ”TCNNKAN: Optimized CNN by Kolmogorov-Arnold network and pruning techniques for semg gesture recognition.” IEEE Journal of Biomedical and Health Informatics, 2024.
93
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
EFFICIENT COVERAGE PATH PLANNING VIA GRADIENT-BASED RECTANGULAR SEGMENTATION Submitted: 5th August 2025; accepted: 2nd March 2026
Hubert Baraniak, Konrad Cop, Morteza Haghbeigi DOI: 10.14313/jamris-2026-041 Abstract: Coverage Path Planning (CPP) is the task of finding a route that covers every point in a region or volume (i.e., all reachable cells) while avoiding obstacles. A typical CPP solution proceeds in two main stages: partitioning the environment into subregions, and planning a path within and between these regions. Effective segmentation is critical for smooth and efficient operation. For example, dividing the area into rectangular cells allows a simple back-and-forth (boustrophedon) sweep in each cell, exhaustively covering it without overlap. In this paper, we propose an adaptive, gradient-based rectangular decomposition for CPP. The algorithm analyzes the occupancy map to split free space into oriented rectangles that conform to obstacle boundaries. Within each rectangle, a straight-line sweep path is generated. Compared to conventional methods, our approach produces coverage paths that are smoother and more concise. We demonstrate on real-world maps that the proposed method achieves good coverage with fewer turns, yielding improvements in overall CPP efficiency. Keywords: Coverage path planning, Mobile robots, Room segmentation
1. Introduction Coverage Path Planning (CPP) is a fundamental problem in mobile robotics. CPP aims to compute a route that passes over all points in a given area or volume while avoiding obstacles [1]. This task is integral to many robotic applications such as agricultural field operations, floor cleaning, and automated lawn mowing [2]. The key objective in CPP is to minimize the total traversal time and energy of the robot while ensuring full coverage of the target area. Because robots often have differential-drive or car-like dynamics, smooth paths with few sharp turns are preferred to reduce wear and energy use. A common strategy is to decompose the environment into simpler subregions before path planning. For example, classical methods partition free space into cells (trapezoids or boustrophedon cells) or use grid-based spanning-tree coverage. In each subregion, a systematic sweep pattern (such as boustrophedon or spiral) is used [3–5]. This approach has been summarized as three steps [6]: (1) divide the environment into smaller regions (cells) for coverage, (2) compute a coverage path for each region and an order to visit 94
them, and (3) smooth the overall path. Our work follows this outline but enforces oriented rectangular regions. Rectangular segmentation has distinct advantages: as shown in prior work, an axis-aligned backand-forth sweep will exhaustively cover a rectangular area without any redundant overlap [3]. This yields coverage paths that consist of long straight segments and only the necessary turns, improving smoothness and trackability. This paper focuses on CPP in real-world operational maps. We assume a known occupancy grid of the environment, such as an office or warehouse floor plan. We emphasize efficiency in practical settings: the robot must cover all accessible areas while reducing idle motion. The research described in this paper originates from the practical requirements of an industrial cleaning robot, which aimed to find the best coverage path of any real-life environment. A typical characteristic of maps recorded by robots in real-life conditions is the presence of noise, which prevents decomposing such maps into regular geometric primitives. Assumption of regular shapes is however, a bottom line for most State-of-the-art algorithms. Throughout the research, we found that these algorithms solve very well the CPP problem on artificially created maps composed of regular polygons. When however applied to reallife maps, these algorithms generate unlikely paths as shown in Figure 5. To address this practical challenge we propose to adapt a decomposition strategy similar to the one a human operator would follow and divide the map into segments approximated by rectangles. This allows our algorithm to create paths that avoid unnecessary crossings of already cleaned floor. The key novelty of this work is the iterative fitting process: a gradient-based optimization loop in which rectangle parameters (position, size, and orientation) are treated as continuous variables and refined via a differentiable rectangle function. Unlike classical decomposition methods that rely on exact geometric primitives, our differentiable formulation provides non-zero gradients across the image domain, enabling smooth, stable convergence even on noisy occupancy grids. Crucially, the gradient descent steps are interleaved with structural operations—merging overlapping rectangles, deleting insignificant ones, and splitting rectangles that span disconnected regions—which reshape the decomposition topology and allow the optimizer to escape local minima. This
Open Access. © 2026 Hubert Baraniak et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
Journal of Automation, Mobile Robotics and Intelligent Systems
combination of continuous parameter optimization with discrete structural adaptation is what enables robust segmentation of real-world maps. In summary, we introduce a gradient-based adaptive algorithm that partitions a given map into rectangular zones and generates the corresponding coverage path. The main contributions are: 1. A differentiable rectangle function that enables gradient-based optimization of rectangle parameters on occupancy grids, providing non-zero gradients where classical binary representations would yield zero or undefined derivatives. 2. An iterative fitting process that combines continuous gradient descent with discrete structural operations (merge, delete, split), enabling robust rectangular segmentation of noisy, realworld maps where conventional decomposition methods fail. 3. An empirical evaluation on large-scale real-world maps demonstrating that the proposed method yields smoother coverage paths with significantly reduced curvature and execution time compared to baseline methods.
2. Related Work Complete Coverage Path Planning has been extensively studied in the robotics literature. Multiple reviews have been published that classify and analyze existing approaches. [7] provides algorithmic frameworks organizing work into heuristic and randomized methods, as well as provably complete approaches based on cellular decompositions, including approximate, semi-approximate, and exact decompositions. [1] surveys how different algorithms decompose and represent environments to generate complete coverage paths, reviewing exact cellular decompositions such as trapezoidal and boustrophedon methods, Morse-based approaches that handle complex obstacle geometries, and grid-based methods using wavefronts, spanning trees, or neural networks. [6] focuses on the main methodological components of environment decomposition, coverage execution, backtracking sequence planning, and path smoothing, surveying decomposition strategies such as cellbased, grid-based, sampling-based, and spanning-tree approaches. More recent comprehensive reviews include [8], which presents a comparative performance analysis of several coverage path planning algorithms, and [9], which explores various CPP approaches for Unmanned Aerial Vehicles, dividing methods into three main groups: no decomposition (simple patterns), exact cellular decomposition, and approximate cellular decomposition. [10] reviews classical and heuristic algorithms with emphasis on optimization characteristics, while [11] focuses specifically on complete coverage path planning for precision agriculture with emphasis on exact cellular decomposition methods. Finally, [12] provides a comprehensive
VOLUME 20,
N∘ 3
2026
survey for static and dynamic environments, classifying methods into grid-based, graph-based, samplingbased, and optimization-driven approaches. Based on these comprehensive reviews, complete coverage path planning can be understood as a threestep process: (1) segmentation or decomposition of the environment into manageable regions, (2) coverage path planning within each region, and (3) sequence optimization to determine the optimal order to visit regions. The decomposition method fundamentally influences the quality of generated coverage paths. 2.1.
Cell Grid-Based Methods
Grid-based methods divide the environment into a regular grid of cells and generate a coverage path that visits all cells. These methods represent an approximate cellular decomposition approach, where no sophisticated decomposition structure is imposed beyond the regular grid structure itself. Spanning tree coverage methods. The spanning tree coverage approach was introduced by [13], which decomposes the free space into mega cells and constructs a spanning tree that traverses all cells; the robot covers the entire area by following the tree’s boundaries and visits four smaller cells within each mega cell. This graph-based representation can guarantee complete coverage in structured environments with few obstacles, but it generates a single, continuous path without decomposing the environment into operational zones and can be computationally expensive for large or highly complex maps. Recent extensions by [14] and [15] improve efficiency through cell reduction, yet still produce monolithic coverage paths. Neural network-based coverage path planning. [16] proposed a neural dynamics-based approach for complete grid coverage using neural networks, leveraging activation patterns to guide navigation across grid-structured environments. While this provides distributed decision-making and can handle complex grid patterns, it is only effective on simple maps and requires a very dense grid, leading to prohibitive memory and computational demands for large industrial environments such as warehouses. 2.2.
Exact Cellular Decomposition Methods
Exact cellular decomposition methods break the free space into simple, non-overlapping regions called cells. These cells are typically defined based on the geometric and topological properties of the environment. Once decomposition is complete, a simple motion pattern (such as back-and-forth motions) is applied within each cell to ensure complete coverage. Boustrophedon decomposition. The boustrophedon cellular decomposition method was introduced by [3]. It is based on trapezoidal decomposition, but improves efficiency by identifying critical points that separate distinct connectivity changes and thus produces fewer cells. The method divides free space into trapezoid-shaped cells via sweep lines from polygonal obstacle 95
Journal of Automation, Mobile Robotics and Intelligent Systems
vertices and then applies simple back-andforth motions within each cell. This yields exact coverage guarantees with an easily traversable cell structure, while still relying on polygonal, wellstructured environments. Morse-based cellular decomposition. [17] generalized decomposition using critical points of Morse functions, enabling decomposition based on distancefunction properties rather than polygonal obstacle definitions. This can handle non-polygonal obstacles and complex geometries with theoretical coverage guarantees, but extensions such as [18] remain targeted at relatively structured environments with clear boundaries and do not directly scale to large, unstructured maps which can include multiple rooms, narrow corridors, and irregularly shaped spaces. 2.3.
Bio-Inspired and Optimization-Based Methods
Bio-inspired and metaheuristic algorithms have been increasingly applied to coverage path planning, either for directly generating coverage paths or for optimizing the sequence and routing of pre-computed coverage sub-regions. Genetic algorithm-based approaches. Genetic algorithms (GA) simulate natural evolution to explore the solution space. [19] applied GA to coverage path planning by encoding paths as chromosomes and using selection, crossover, and mutation, with later applications in agricultural coverage [20] and irregular regions [21]. These methods can optimize multiple criteria such as path length and coverage quality without strict geometric assumptions, but they are computationally expensive at scale and do not inherently decompose the workspace, often producing paths that jump between distant regions rather than supporting zone-by-zone industrial operations. Ant colony optimization. Ant Colony Optimization (ACO) uses swarm intelligence, where virtual ants deposit pheromone trails that guide subsequent search toward promising solutions. [22] developed a fast-spanning ACO method for mobile robot coverage, and [23] combined simulated annealing with ACO for interior cleaning applications. ACO can find near-optimal solutions and handle constraints, but it explores a large search space, tends to converge slowly in large industrial maps, and does not impose a decomposition structure, limiting its ability to produce meaningful operational zones. 2.4.
Recent Advances and Specialized Applications
Recent work in coverage path planning has addressed specialized application domains and attempted to tackle limitations of earlier approaches. [24] proposed a calculation-based shortest path search algorithm (CbSPSA) for marine survey applications, partitioning irregular polygon areas into convex sub-regions for efficient coverage. [25] developed an improved spanning tree coverage method with optimized backtracking for large-scale unstructured social environments (cleaning tasks). [26] integrated Building Information Modeling (BIM) with coverage path planning for indoor robot 96
VOLUME 20,
N∘ 3
2026
applications, incorporating semantic enrichment of maps to improve coverage efficiency. For agricultural applications, [27] proposed hybrid path planning combining nested and spiral methods for harvesting operations in convex polygonal fields. [28] developed a Voronoi-based decomposition framework for large outdoor sweeping operations, demonstrating that Voronoi decomposition outperforms grid-based decomposition. [29] addressed coverage planning in semi-structured outdoor environments by combining map preprocessing with coverage path planning to maximize coverage efficiency. Finally, [30] introduced Fields2Benchmark, a modular benchmark for standardizing evaluation of agricultural coverage path planning, decomposing the problem into field decomposition, headland generation, swath generation, route planning, and path planning stages. While these methods address specific application domains and achieve improvements in efficiency, they typically still rely on either: (1) decomposition methods that assume structured or convex environments, (2) optimization techniques that do not inherently segment the workspace into manageable operational zones, or (3) specialized approaches tailored to particular problem structures that lack generalizability to arbitrary environments. 2.5.
Analysis and Motivation
Existing CPP methods either generate monolithic coverage paths on grids, assume polygonal structure for exact decomposition, or optimize paths without enforcing meaningful segmentation, which makes them ill-suited to large, unstructured maps that require zone-by-zone operation. Most existing methods either assume idealized map conditions or require extensive preprocessing to be effective, and many fail to adapt efficiently to the irregular, noisy maps encountered in real-world warehouse environments. There is thus a clear gap in segmentation methods that can scale to industrial-sized environments while producing operationally practical, separable coverage zones. We therefore introduce our method to decompose complex layouts into efficient rectangular zones that can be covered independently. By leveraging geometric regularities and gradient information to simplify map structure while adapting to obstacle distribution, our approach enables robust, efficient, modular, and scalable path planning in large-scale cleaning tasks.
3.
Problem Description
We address the problem of generating a complete coverage path over a known 2D occupancy grid 𝑀 of size 𝑊 × 𝐻 with resolution 𝑟 (pixels per meter). Free cells (value 1) represent traversable areas, and occupied cells (value 0) represent obstacles. The objective is to compute a route that covers all accessible areas with minimal path cost, defined in terms of length, curvature, and execution time. Formally, given the grid map 𝑀 and resolution 𝑟, we seek a partition of free space into the smallest and
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
least numerous set of rectangular areas {𝑅1 , … , 𝑅𝑛 }. Each 𝑅𝑖 is an axis-aligned rectangle specified by its center, width, height, and orientation. A coverage path is then planned so that each zone is fully swept. To evaluate path quality, we use the following metrics: • Length per square meter (𝐴𝐿 ): 𝐴𝐿 =
𝐿route , 𝑁covered
(1)
where 𝐿route is the total path length (m), and 𝑁covered is the total area (m2 ) effectively covered by the robot. This metric reflects the efficiency of movement—lower values indicate less distance traveled per unit area cleaned. • Coverage ratio (𝐶): 𝐶=
𝑁covered , 𝑁free
(2)
where 𝑁free is the total free area in the map. This metric quantifies completeness of coverage—ideally close to 1, indicating minimal missed regions. • Curvature per square meter (𝜅): 𝜅=
∑ |Δ𝜃| , 𝐿route
(3)
where Δ𝜃 are turn angles along the path. High curvature implies frequent direction changes, which may slow execution or introduce mechanical wear, especially in large or cluttered environments. • Time per square meter (𝑇): 𝑇total , 𝑇= 𝑁covered
(4)
where 𝑇total includes estimated execution time, incorporating robot kinematics such as maximum velocity, acceleration limits, and time spent stopping and turning. This reflects real-world performance, prioritizing motion plans that are not just short but also smooth and feasible. 3.1.
Figure 1. Coverage paths generated by the Gradient-Based Rectangular Segmentation method. Red, green, and blue lines represent wall-following, boustrophedon, and A* paths, respectively. Grayscale intensity indicates coverage frequency, with darker regions representing areas covered multiple times where each component is calculated as: • 𝑇linear ∶ Time to traverse straight paths with acceleration/deceleration phases
Cleaning Time Estimation
The total cleaning time is estimated by modeling the robot’s motion constraints, including maximum ang lin linear and angular velocities (𝑉max , 𝑉max ) and accelerang lin ations (𝐴max , 𝐴max ). The path is segmented into linear and turning segments. For each linear segment, the time is computed by considering acceleration and deceleration phases to and from cruising speed, ensuring the robot respects acceleration limits. For turns, the time is estimated based on the angle turned, accounting for maximum angular velocity and acceleration, assuming the robot stops to rotate in place. Formally, the total time is the sum of: 𝑗
𝑖 𝑇total = � 𝑇linear + � 𝑇turn , 𝑖
𝑗
(5)
• 𝑇turn ∶ Time to perform stationary rotations given angle and angular acceleration constraints This approach realistically estimates the robot’s cleaning time considering kinematic constraints and motion phases.
4.
Methodology
Our CPP pipeline has three main stages: (1) segment the map into rectangular zones, (2) generate a local coverage path in each zone, and (3) sequence the zones into a global path. The algorithm is implemented in Python and takes a binary occupancy grid as input. The goal of the segmentation step is to take an input map and generate the smallest and least numerous set of rectangles that cover the entire open space. 97
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 2. Comparison of initial (left) and regressed (right) rectangles
Figure 3. Progression of rectangle representation during optimization steps after 3 steps (left), 30 steps (middle), 300 steps (right) To achieve this, we employ an optimization approach based on gradient descent, which minimizes the error between the real map and a model of the map represented as the sum of several rectangular regions. Below, we outline the key components of our segmentation process.
• 𝑅 is the matrix of rectangle parameters, with each row representing a different rectangle.
The initial set of rectangles should approximately cover the area of interest, providing a starting point for the optimization process. For this task, we utilize a Voronoi decomposition of the map to divide it into regions that approximate the open spaces. We then apply the probabilistic Hough transform to identify the directions and centers of most rooms and corridors. Using these center points, we define initial rectangles by assigning each position and length of the corresponding line, with a predefined width.
• 𝑟(𝑅𝑗 ) is the rectangle function, which takes the parameters of a rectangle 𝑅𝑗 (position 𝑥, 𝑦, width 𝑤, height ℎ, and rotation angle 𝑟) and returns an image of the rectangle. This parameterization was specifically chosen to ensure parameter space continuity during Gradient Descent. By defining the rectangle via its center (𝑥, 𝑦) and a continuous rotation angle 𝜃, we avoid the ‘jump’ discontinuities that occur when using cornerbased coordinates. This ensures that infinitesimal updates to the parameters result in smooth, predictable transformations of the rectangle geometry.
4.2.
• 𝑀 is the input map.
4.1.
Initial Rectangle Generation
Rectangle Fitting via Gradient Descent
Once the initial set of rectangles is defined, we apply a gradient descent algorithm to refine the fit of these rectangles to the open spaces in the map. The goal is to adjust the position and dimensions of each rectangle to minimize the difference between the actual map and the sum of the rectangles. The objective function to be minimized is defined as: 𝑘
2
𝑂(𝑅) = �� 𝑟(𝑅𝑗 ) − 𝑀� + 𝐽(𝑅), 𝑗=1
98
where:
• 𝐽(𝑅) is a regularization term that penalizes overly small rectangles, encouraging the replacement of smaller rectangles with larger ones. Here w(R) and h(R) refers to width and height of the given rectangle respectively 𝑘
1 𝐽(𝑅) = � ��𝑤(𝑅𝑗 ) + �ℎ(𝑅𝑗 )� 𝑘
(7)
𝑗=1
(6) The main objective is to ensure that the sum of the rectangle images closely matches the input map, while minimizing the number of unnecessary or redundant rectangles.
Journal of Automation, Mobile Robotics and Intelligent Systems
4.3.
Coverage Enhancement through Weighted Objective Function
While the basic objective function (Eq. 6) treats all pixel-level errors equally, in practice uncovered free space is more detrimental to cleaning performance than minor overlaps or overestimation. To address this asymmetry, we introduce a per-pixel weighting scheme that amplifies the penalty for pixels where the map indicates free space but the current rectangle configuration fails to cover it. 𝑘 Let 𝑆 = ∑𝑗=1 𝑟(𝑅𝑗 ) denote the sum of all rectangle functions and 𝑒𝑖,𝑗 = 𝑆𝑖,𝑗 − 𝑀𝑖,𝑗 the per-pixel residual. We define a weighting function 𝑤(𝑒𝑖,𝑗 ) that assigns different importance to each pixel based on the sign of its residual: 𝛼 𝑤(𝑒𝑖,𝑗 ) = � 1
VOLUME 20,
if 𝑒𝑖,𝑗 < 0 (uncovered free space) otherwise (8)
1 2 � �𝑤(𝑒𝑖,𝑗 ) ⋅ 𝑒𝑖,𝑗 � + 𝐽(𝑅), 𝑁
(9)
𝑖,𝑗
where 𝑁 is the total number of pixels and 𝐽(𝑅) is the regularization term as defined previously. The intuition behind this formulation is straightforward: when the residual 𝑒𝑖,𝑗 < 0, the map contains free space at pixel (𝑖, 𝑗) that no rectangle currently covers. By scaling such residuals by 𝛼 > 1, the gradient signal from uncovered regions is amplified, causing the optimizer to preferentially expand or shift rectangles toward those gaps. In contrast, pixels where rectangles extend slightly beyond the free space (𝑒𝑖,𝑗 > 0) receive a standard penalty, allowing the optimizer to tolerate minor overestimation when it helps achieve better overall coverage. This weighting mechanism also helps preserve coverage when rectangles are merged or deleted during the iterative fitting process: by penalizing uncovered free space more heavily, rectangles are discouraged from collapsing prematurely. The effect of 𝛼 on coverage and execution time is evaluated in Section 5.3.. 4.4.
Rectangle Function
The rectangle function 𝑟(𝑅) is a key element in our optimization process. It takes the parameters of a rectangle (𝑥, 𝑦, 𝑤, ℎ, 𝑟) and returns an image where the pixels inside the rectangle are white (indicating open space), and the background pixels are black (indicating obstacles or areas not requiring cleaning). The dimensions of the output image match the dimensions of the input map. However, simply drawing rectangles directly onto a binary image is not sufficient for our optimization. The gradient of the objective function would either
2026
be zero or undefined for most pixel values, making it impossible to adjust the rectangle parameters during gradient descent. To resolve this, we define a differentiable rectangle function where the partial derivatives of pixel values with respect to the rectangle parameters are nonzero across a significant portion of the image. This ensures that the gradient descent algorithm can effectively update the rectangle parameters to minimize the objective function.
where 𝛼 > 1 is a weighting parameter that controls how strongly the optimizer prioritizes covering free space over reducing other errors. The modified objective function is then:
𝑂𝑤 (𝑅) =
N∘ 3
𝑟𝑖,𝑗 = max(min(𝑔𝑖,𝑗 , ℎ𝑖,𝑗 , 1), 0)
(10)
ℎ𝑖,𝑗 = 𝐶 ⋅ �−|𝑥 ′ | +
𝑤 � + 0.5 2
(11)
𝑔𝑖,𝑗 = 𝐶 ⋅ �−|𝑦 ′ | +
ℎ � + 0.5 2
(12)
𝑥 ′ = (𝑖 − 𝑥) cos(𝑟) − (𝑗 − 𝑦) sin(𝑟)
(13)
𝑦 ′ = (𝑖 − 𝑥) sin(𝑟) + (𝑗 − 𝑦) cos(𝑟)
(14)
The result of this differentiable rectangle function is a blurred image of a rectangle blurred around the edges. The rectangle function takes the smaller value of its border functions 𝑔 and ℎ, clipped to the range [0, 1]. Each border function attains its maximum at the center of the rectangle and a value of 0.5 at the edge of the rectangle. Values below 0.5 lie outside the rectangle. The coordinates 𝑥 ′ and 𝑦 ′ are in the rectangle’s frame of reference. The constant 𝐶 stands for clarity. As 𝐶 approaches infinity, all rectangle function values become either 0 or 1. Smaller values of 𝐶 produce more blur around the edges of the rectangle. 4.5.
Iterative Fitting Process
The iterative fitting process is the core contribution of this work. It involves iterative updates of all rectangle parameters based on the gradient of the objective function, interleaved with structural operations that modify the decomposition topology. This combination of continuous optimization with discrete restructuring is what distinguishes our approach from classical decomposition methods and enables robust segmentation of noisy, real-world maps. To improve both convergence speed and the quality of the solution, we employ three key operations: merging, deleting, and splitting of rectangles. Rectangle parameter update At each iteration, the parameters of every rectangle are updated using gradient descent. The objective function measures the difference between the map and the current set of rectangles, and its gradient dictates how to adjust each rectangle. The updated parameters ensure that the rectangles better fit the open space of the map. 99
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 4. Comparison of rectangle function with different sharpness parameters. Both plots show a rectangle with dimensions 6.0 × 2.5 and rotation angle 30°. Left: Lower clarity (C=0.5) produces smoother transitions. Right: Higher clarity (C=2.0) produces sharper edges. The heatmaps display the continuous function values 𝑟𝑖,𝑗 , with the red outline showing the corresponding rectangle boundaries. The boundaries of the corresponding rectangle lie in the middle of the transition region Merging operation The merging operation is applied when two rectangles overlap or cover areas that can be more efficiently represented by a single rectangle. We proceed as follows: • Compute the average orientation 𝜃avg of the two rectangles, 𝜃1 and 𝜃2 . • Create a new rectangle, oriented along 𝜃avg , that covers all vertices of the two input rectangles. • Compare the area of the merged rectangle 𝐴merged with the sum of the areas of the two input rectangles 𝐴1 and 𝐴2 . If 𝐴merged < (𝐴1 + 𝐴2 ) ⋅ 𝑐,
(15)
where 𝑐 is a constant slightly larger than 1 (e.g., 𝑐 = 1.05), then merge the two input rectangles into the newly calculated rectangle. Deleting operation The deleting operation removes rectangles whose area is below a predefined threshold 𝐴min . This prevents small, insignificant rectangles from cluttering the representation. Splitting operation In some cases, poor initialization can cause a single rectangle to cover multiple disconnected open spaces, especially in the presence of narrow obstacles like walls. When this occurs, the rectangle should be split into multiple smaller rectangles, each corresponding to a separate open space. We proceed as follows: • Identify if the rectangle spans across multiple regions separated by obstacles by using a regiongrowing algorithm to find connected components within the rectangle. • For each connected region, apply Principal Component Analysis (PCA) to the set of points within the region to find the main orientation of the space. 100
• Use the principal PCA vector to orient a new rectangle that covers the region. • Replace the original rectangle with the newly calculated set of smaller rectangles. 4.6.
Path Planning
Inside each rectangle, we use a simple boustrophedon path along the longer side of the rectangle. If the path encounters an obstacle, we use A* to find the shortest path that avoids that obstacle. Finally, we sequence the set of rectangles by solving a Traveling Salesman Problem (TSP) to find a near-optimal solution that visits all regions. Sweeping along the longer axis of each rectangle is a deliberate design choice that prioritizes predictability and geometric efficiency: 1. Predictability in industrial settings: Boustrophedon paths along consistent axes produce deterministic, easily predictable trajectories. This is critical for industrial cleaning robots operating in shared human environments, where path predictability enhances safety and operational reliability. 2. Geometric efficiency of rectangular decomposition: Our segmentation algorithm naturally decomposes free space into rectangles (corridors, hallways). For such geometries, sweeping along the longer axis minimizes turning points and maximizes straight-line coverage segments, reducing execution time and energy consumption. 4.7.
Implementation Details
The entire pipeline is implemented in Python 3. Input maps can be loaded from standard formats (e.g., image files or ROS-style occupancy grids). We use NumPy for array processing and OpenCV for imagebased operations. For each map, the segmentation
Journal of Automation, Mobile Robotics and Intelligent Systems
outputs a list of rectangles, which we store in a JSON file. The final coverage path is stored as an ordered list of (𝑥, 𝑦) waypoints in the same JSON. We also generate plots of the map, rectangles, and path using Matplotlib for verification. This modular design makes it easy to modify or extend each step. The optimization algorithm was implemented using the Keras library to facilitate automatic gradient calculation and parameter updates. We use the Adam optimizer [31] to improve convergence speed and stability of the segmentation process. The fitting process is set to run for 300 iterations, with merge, delete, and split operations performed every 30 iterations. The choice of these hyperparameters reflects a balanced trade-off between two competing objectives. First, the Adam optimizer leverages adaptive learning rates and momentum estimates to accelerate convergence and navigate the non-convex optimization landscape effectively. However, excessive iterations without structural modification can lead to suboptimal local minima, as the gradient-based updates alone cannot escape geometrically inefficient rectangle configurations. Second, the merge, delete, and split operations are crucial for escaping such local minima by restructuring the solution space. Performing these operations too frequently (e.g., every iteration) introduces instability, as the optimizer has insufficient time to refine parameters before the topology changes. Conversely, performing them too infrequently reduces their effectiveness in improving the overall segmentation quality.
VOLUME 20,
5.1.
Benchmark Algorithms
To evaluate the advantages of the proposed method, results are compared to three benchmark planning methods: • Boustrophedon: A classical coverage path algorithm that generates zigzag paths. [3]
The implementations of the benchmark algorithms were sourced from the open-access library https://gi thub.com/ipa320/ipa_coverage_planning [8]. 5.2.
• Energy functional: The coverage path planning method directs the robot to systematically explore a room by following a grid of observation points, selecting each next point based
Performance Metrics
For each method we measured average values of the metrics described above. The values are presented in Table 1. For every algorithm, the following parameters were used. Parameters were selected based on real cleaning robot data. Same set of parameters are used on our proposed Efficient Coverage Path Planning via Gradient-Based Rectangular Segmentation method and three benchmark methods. • Robot radius: 0.4 m • Coverage radius: 0.32 m • Time calculation parameters: – Maximum linear velocity: 0.5 m/s – Maximum angular velocity: 1.0 rad/s – Maximum linear acceleration: 0.3 m/s2 – Maximum 0.5 rad/s2
angular
acceleration:
– Maximum linear deceleration: 0.3 m/s2 – Maximum 0.5 rad/s2
angular
deceleration:
Overall, the proposed rectangular segmentation yielded smoother and more efficient coverage paths. The path visuals (see Figure 5) show clean back-andforth sweeps covering each region. Our method outperformed the benchmarks in average path length per 𝑚2 giving overall shorter paths, significantly smaller curvature per 𝑚2 meaning paths are simpler and the robot can achieve higher average velocity. Time per 𝑚2 is also significantly smaller than other algorithms due to the shorter and simpler paths. The algorithm maintained good coverage (comparable to the boustrophedon benchmark). 5.3.
• Neural network: The method uses a topologically organized neural network, governed by shunting dynamics based on Hodgkin and Huxley’s equation, to autonomously generate collision-free, complete coverage paths for cleaning robots in dynamic environments through local neural interactions and previous robot positions. [32]
2026
on an energy function that balances travel efficiency, smooth navigation, and area coverage while adapting to dynamic obstacles. [33]
5. Simulation and results We evaluated our method on multiple test environments recorded in warehouses, shopping malls and production facilities. Test environments have different dimensions ranging from 2109 × 1221 px to 6159 × 5079 px. All experiments were run on a standard desktop (Intel Core i7 CPU, 16 GB RAM). Our software environment was Python 3.11.
N∘ 3
Sensitivity to Coverage Weighting Parameter
To quantify the effect of the weighting parameter 𝛼 introduced in the weighted objective function, we evaluated three representative values on our benchmark maps: • 𝛼 = 1: Coverage of 86.64% with 4.59 s/m2 execution time • 𝛼 = 1.5: Coverage of 89.98% with 4.69 s/m2 execution time • 𝛼 = 3.0: Coverage of 92.97% with 5.22 s/m2 execution time 101
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 1. Summary of results for different methods Method Boustrophedon [3] Neural Network [32] Energy Functional [33] Gradient-Based Segmentation
avg. Length per 𝑚2
avg. Coverage
avg. Curvature per 𝑚2
avg. Time per 𝑚2
2.861 5.110 2.563 2.094
0.8713 0.8713 0.9831 0.8664
1.001 1.880 2.520 0.474
6.893 12.907 10.863 4.594
Figure 5. Visual comparison of coverage paths generated by different algorithms. Boustrophedon (top left), Energy Functional (top right), Neural Network (bottom left), Gradient-Based Rectangular Segmentation (bottom right) • 𝛼 = 10.0: Coverage of 93.31% with 5.50 s/m2 execution time These results demonstrate a clear trade-off between coverage completeness and execution efficiency. Higher values of 𝛼 lead to better coverage but increase path complexity and execution time, as rectangles are pushed to cover marginal areas at the cost of geometric regularity. For the experiments reported in this paper, we used 𝛼 = 1, which provides a balanced compromise between coverage quality and efficiency. However, applications requiring maximum coverage (e.g., critical cleaning tasks 102
in sterile environments) may benefit from higher values despite the increased execution time. Further improvements, such as reinitializing new rectangles in uncovered regions at a later stage of the optimization, remain future work.
6.
Conclusion
We have presented a new coverage path planning framework that uses gradient-based rectangular segmentation of the environment. By partitioning the map into axis-aligned rectangles adapted to
Journal of Automation, Mobile Robotics and Intelligent Systems
obstacle boundaries, our method enables simple backand-forth sweep coverage in each zone. This leads to coverage paths that are smooth, easy to track, and efficient in terms of travel distance. The experimental results demonstrate that our approach improves the efficiency of the Boustrophedon path planning algorithm. Notably, the proposed method achieves shorter paths with fewer turns while maintaining high coverage of the target area. This highlights the robustness and scalability of our solution, making it particularly useful in industrial and commercial robotics applications. Segmentation of the map enables easy manual modifications of the paths. Another advantage is that path of the robot is easier to predict, which is important in environments with a human traffic. Despite these promising results, several areas for future work remain. Firstly, the computational cost of the segmentation process, particularly during gradient descent, could be further reduced. By employing more sophisticated optimization techniques, such as metaheuristic approaches, may enhance the computational performance. The low coverage of the method can be improved by creating planning paths in areas not covered by the algorithm. Additionally, rectangle function can be modified to avoid local minima. Another avenue for future research involves exploring more sophisticated strategies for path planning within segmented rectangles. While the current approach employs simple Boustrophedon paths, more advanced path-planning methods could further improve the efficiency of the robot’s movement, especially in cluttered environments. Our method provides a significant step forward in the field of Coverage Path Planning by combining map segmentation, optimization, and efficient path generation. Future research should focus on enhancing computational efficiency, robustness to real-world conditions to improve system performance further. We believe that these developments will contribute to the broader adoption of autonomous robotic solutions in various sectors, including manufacturing, logistics, and cleaning services.
AUTHORS Hubert Baraniak∗ – Hubert Baraniak, Software Solutions Wiś niowa 1, 95-063 Rogó w, e-mail: hbsoftware@protonmail.com. Konrad Cop – United RobotsPrymasa Tysiąclecia 46, 01-242 Warszawa, e-mail: konrad.cop@unitedrobots.co. Morteza Haghbeigi – Warsaw University of TechnologyPl. Politechniki 1, 00-661 Warsaw, e-mail: morteza.haghbeigi.dokt@pw.edu.pl. ∗
Corresponding author
References [1] E. Galceran and M. Carreras, “A survey on coverage path planning for robotics”, Robotics and Autonomous Systems,
VOLUME 20,
N∘ 3
2026
vol. 61, no. 12, 2013, 1258–1276, https://doi.org/10.1016/j.robot.2013.09.004. [2] T. Oksanen and A. Visala, “Coverage path planning algorithms for agricultural field machines”, Journal of Field Robotics, vol. 26, no. 8, 2009, 651–668, https://doi.org/10.1002/rob.20300. [3] H. Choset and P. Pignon, “Coverage path planning: The boustrophedon cellular decomposition”, Proc. Field and Service Robotics, 1998, 203–209. [4] X. Miao, J. Lee, and B.-Y. Kang, “Scalable coverage path planning for cleaning robots using rectangular map decomposition on large environments”, IEEE Access, vol. 6, 2018, 38200–38215. [5] S. Bochkarev and S. L. Smith, “On minimizing turns in robot coverage path planning”, Proc. 2016 IEEE International Conference on Automation Science and Engineering (CASE), 2016, 1237–1242. [6] A. Khan, I. Noreen, and Z. Habib, “On complete coverage path planning algorithms for non-holonomic mobile robots: Survey and challenges.”, Journal of Information Science and Engineering, vol. 33, no. 1, 2017, 101–121. [7] H. Choset, “Coverage for robotics – a survey of recent results”, Annals of Mathematics and Artificial Intelligence, vol. 31, 2001, 113–126, 10.1023/A:1016639210559. [8] R. Bormann, F. Jordan, J. Hampp, and M. Hä gele, “Indoor coverage path planning: Survey, implementation, analysis”, Proc. 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018, 1718–1725, 10.1109/ICRA.2018.8460566. [9] T. Cabreira, L. Brisolara, and P. R. F. Jr., “Survey on coverage path planning with unmanned aerial vehicles”, Drones, vol. 3, 2019, 4, 10.3390/drones3010004. [10] C. S. Tan, R. Mohd-Mokhtar, and M. R. Arshad, “A comprehensive review of coverage path planning in robotics using classical and heuristic algorithms”, IEEE Access, vol. 9, 2021, 119310–119342, 10.1109/ACCESS.2021.3108177. [11] M. Hö ffmann, S. Patel, and C. Bü skens, “Optimal guidance track generation for precision agriculture: A review of coverage path planning techniques”, Journal of Field Robotics, vol. 41, 2024, 823–844, 10.1002/rob.22286. [12] K. P. Jayalakshmi, V. G. Nair, and D. Sathish, “A comprehensive survey on coverage path planning for mobile robots in dynamic environments”, IEEE Access, vol. 13, 2025, 60158–60185, 10.1109/ACCESS.2025.3556446. 103
Journal of Automation, Mobile Robotics and Intelligent Systems
[13] Y. Gabriely and E. Rimon, “Spanning-tree based coverage of continuous areas by a mobile robot”, Annals of Mathematics and Artificial Intelligence, vol. 31, 2001, 77–98, 10.1023/A:1016610507833. [14] H. V. Pham, P. Moore, and D. X. Truong, “Proposed smooth-stc algorithm for enhanced coverage path planning performance in mobile robot applications”, Robotics, vol. 8, 2019, 10.3390/ROBOTICS8020044. [15] K. R. Guruprasad and T. D. Ranjitha, “Cpc algorithm: Exact area coverage by a mobile robot using approximate cellular decomposition”, Robotica, vol. 39, 2021, 10.1017/S026357472000096X. [16] A. Singha, A. K. Ray, and A. B. Samaddar, “Neural dynamics based complete grid coverage by single and multiple mobile robots”, SN Applied Sciences 2021 3:5, vol. 3, 2021, 543–, 10.1007/s42452-021-04508-5. [17] H. Choset, E. Acar, A. A. Rizzi, and J. Luntz, “Exact cellular decompositions in terms of critical points of morse functions”, Proc. 2000 IEEE International Conference on Robotics and Automation (ICRA), Millennium Conference, vol. 3, 2000, 2270–2277. [18] S. Karakaya and M. Z. Konyar, “Hybrid boustrophedon and direction-biased region transitions for mobile robot coverage path planning: A region-based multi-cost framework”, Applied Sciences (Switzerland), vol. 15, 2025, 10.3390/app152312666. [19] Z. Wang and Z. Bo, “Coverage path planning for mobile robot based on genetic algorithm”, Proc. 2014 IEEE Workshop on Electronics, Computer and Applications (IWECA), 2014, 732–735, 10.1109/IWECA.2014.6845726. [20] X. Wu, J. Bai, F. Hao, G. Cheng, Y. Tang, and X. Li, “Field complete coverage path planning based on improved genetic algorithm for transplanting robot”, Machines, vol. 11, 2023, 10.3390/machines11060659.
N∘ 3
2026
robot based on improved simulated annealing algorithm and ant colony algorithm”, Signal, Image and Video Processing, vol. 18, 2024, 10.1007/s11760-023-02989-y. [24] J.-H. Li, H. Kang, M.-G. Kim, H. Jin, M.-J. Lee, G. R. Cho, and C. Bae, “Full coverage of confined irregular polygon area for marine survey”, IEEE Access, vol. 11, 2023, 92200–92208, 10.1109/ACCESS.2023.3308145. [25] C. Wang, W. Dong, R. Li, H. Dong, H. Liu, and Y. Gao, “An improved stc-based full coverage path planning algorithm for cleaning tasks in large-scale unstructured social environments”, Sensors, vol. 24, 2024, 10.3390/s24247885. [26] Z. Chen, H. Wang, K. Chen, C. Song, X. Zhang, B. Wang, and J. C. Cheng, “Improved coverage path planning for indoor robots based on bim and robotic configurations”, Automation in Construction, vol. 158, 2024, 10.1016/j.autcon.2023.105160. [27] N. Wang, Z. Jin, T. Wang, J. Xiao, Z. Zhang, H. Wang, M. Zhang, and H. Li, “Hybrid path planning methods for complete coverage in harvesting operation scenarios”, Computers and Electronics in Agriculture, vol. 231, 2025, 10.1016/j.compag.2025.109946. [28] B. F. Gó mez, A. Jayadeep, M. A. V. J. Muthugala, and M. R. Elara, “A framework for coverage path planning of outdoor sweeping robots deployed in large environments”, Mathematics, vol. 13, 2025, 2238, 10.3390/math13142238. [29] K. Wu, Z. Wu, S. Lu, and W. Li, “Full coverage path planning strategy for cleaning robots in semi-structured outdoor environments”, Robotics and Autonomous Systems, vol. 192, 2025, 10.1016/j.robot.2025.105050. [30] G. Mier, A. M. C. Faulı́, J. Valente, and S. de Bruin, “Fields2benchmark: An open-source benchmark for coverage path planning methods in agriculture”, Smart Agricultural Technology, vol. 12, 2025, 10.1016/j.atech.2025.101156. [31] D. P. Kingma and J. Ba. “Adam: A method for stochastic optimization”, 2017.
[21] G. Chen, Y. Du, X. Xi, K. Zhang, J. Yang, L. Xu, and C. Ren, “Improved genetic algorithm based on bilevel co-evolution for coverage path planning in irregular region”, Scientific Reports, vol. 15, 2025, 10.1038/s41598-025-93492-6.
[32] S. Yang and C. Luo, “A neural network approach to complete coverage path planning”, IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), vol. 34, no. 1, 2004, 718–724, 10.1109/TSMCB.2003.811769.
[22] C. Carr and P. Wang, “Fast-spanning ant colony optimisation for mobile robot coverage path planning”, Proc. UK Workshop on Computational Intelligence, 2022, 463–474.
[33] R. Bormann, J. Hampp, and M. Hagele, “New brooms sweep clean - an autonomous robotic cleaning assistant for professional office cleaning”, Proc. IEEE International Conference on Robotics and Automation, 2015, 4470–4477, 10.1109/ICRA.2015.7139818.
[23] K. Shi, W. Wu, Z. Wu, B. Jiang, and H. R. Karimi, “Coverage path planning for cleaning 104
VOLUME 20,
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
AN EFFICIENT SWARM CONTROL ALGORITHM FOR COVERING AN AREA WITH COMMUNICATION NETWORK Submitted: 29th December 2023; accepted: 29th January 2024
Dariusz Miedziński, Mariusz Jacewicz, Kacper Kaczmarek, Sebastian Topczewski, Robert Głębocki, Antoni Kopyt DOI: 10.14313/jamris-2026-042 Abstract: There is an increase in the popularity of Unmanned Aerial Vehicles that can fly in autonomous swarms. Many applications, such as surveillance, disaster response, or enabling wireless access networks can be found for such drones. In this work, an efficient swarm control algorithm for establishing a WiFi network over a specified area in emergency situations was prepared and tested in a developed swarm simulator. The test results show that the area can be effectively and fully covered by the drone formations, without in-flight collisions, and keeping the necessary distances between the drones. Keywords: UAV, swarm, simulation
1. Introduction Unmanned Aerial Vehicles (UAVs) are becoming increasingly popular in a vast number of applications, such as surveillance, remote sensing, or small payload delivery, to name a few. Drones are relatively cheap and easy to use in commercial and military applications compared to other types of flying vehicles. Human- or Pilot-in-the-Loop is the usual mode of operation, in which the drone is remotely controlled by the operator, driven by the visual feedback from the drone’s onboard camera. The next step for the UAV’s technology development is the simultaneous flight of the drone swarm. It is necessary to combine path planning, routing, formation, and overall drone swarm control to achieve that. The challenges of the suitable interaction between the swarm and humans are discussed in [1], where two types of control are specified. Another critical aspect of the UAV swarm is the communication between the drones, which allows for an increase in the level of swarm autonomy and flight safety. The proposition of using 5G mobile network systems is discussed in [2]. Recently, an increase in the development of the various swarm control algorithms that allow to alleviate human engagement and attention in the mission operation of the swarm can be observed. A comprehensive review of various approaches can be found in [3]. A very popular approach is to use flocking behavior [4, 5], where the swarm’s overall movement results from a few basic rules like alignment, separation, and cohesion. A digital pheromone field approach, in which the swarm is attracted to the places of interest and repelled from the danger zones specified by the
operator, for swarm control is presented in [6–8]. A combination of flocking and stigmergic (attraction to pheromones) behavior for swarm control is investigated in [9, 10]. A different approach, using a visionbased formation control of the group of mobile robots, is explored in [11]. The authors in [12] proposed a decentralized hybrid control mechanism for swarm formation control, allowing for collision avoidance. One of the popular methods for flying in a swarm is to use drone formations. Swarm formation control using evolutionary or genetic algorithms is explored in [13–15]. A combination of genetic algorithm and PSO was introduced in [16]. The use of convex optimization techniques for formation control was investigated in [17–19], and a model predictive control approach was presented in [20]. Another important aspect of swarm mission planning is efficient and optimized routing and path planning. The use of evolutionary algorithms for this purpose is presented in [21]. Particle Swarm Optimization (PSO) was used in [22, 23]. A comparison between Genetic Algorithms and PSO can be found in [24], while a convex optimization approach was discussed in [25]. To test the prepared swarm control algorithms, there is a need to perform high-fidelity simulations that account for the environmental conditions, such as wind and drones’ dynamic behavior. Examples of available swarm simulators are Gazebo [21], ARGos [26], and SwarmLab [27]. The abovementioned tools are not suitable for the considered application due to various reasons: low computational efficiency, limited customization options, or lack of some functionalities. It was a practical necessity to develop an advanced solution that can be used for flight test planning purposes. The purpose of this paper is to present a full six degrees of freedom (6DOF) drone swarm simulator running an efficient swarm control algorithm that allows covering a specified area with a grid of drones, taking into account environmental conditions, terrain, and battery supply, for the purpose of establishing a WiFi network in emergency situations. The remainder of the article is structured as follows. Section 2 describes the overall simulator’s architecture, and sections 3 and 4 present the models of the drones’ dynamics and swarm control algorithm, respectively. Section 5 describes the proposed mission of the swarm, section 6 presents the results of the
Open Access. © 2026 Mariusz Jacewicz, published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
105
Journal of Automation, Mobile Robotics and Intelligent Systems
simulation, and section 7 sums up the paper with conclusions.
2. Simulator’s Architecture The developed swarm simulator was written entirely in MATLAB/Simulink R2023a. It is divided into two major components, the drone model and swarm control system, presented schematically in Figure 1. Two drone models were prepared. One is based purely on kinematics and simulates the three degrees of freedom (3DOF) movement of the point mass. This model, due to its fast computational time, can be used for a quick evaluation of the feasibility of the prepared mission plan. In this model, the commanded drone velocity, calculated for each of the drones by the swarm control algorithm, is directly equal to the drone’s actual velocity. The position is calculated by integrating this velocity. The second model utilizes the full 6DOF rigid body movement. It includes a four-channel autopilot with control mixer, propellers model, and aerodynamic drag as well as atmosphere model including wind with constant speed and direction. This model can be used to accurately evaluate drones’ performance during mission execution. The swarm control algorithm enables to flying of a swarm comprised of two types of drones: Leaders and Followers. There can be many Leaders flying simultaneously and each one of them can have multiple Followers arranged in a formation around them. The Leader drone is responsible mainly for tracking the specified trajectory given in the form of waypoints. The Follower’s mission is to simply follow the Leader, sticking to its predefined formation. Additionally, a formation reconfiguration possibility is implemented which allows for swapping the Followers between formation and base, e.g. after a low battery condition is detected. Drone’s 6DOF dynamics and swarm control algorithm are described in detail in the following sections.
VOLUME 20,
N∘ 3
2026
Table 1. Leaders’ parameters Parameter
Value
Mass [kg]
2.182
Moment of inertia 𝐼𝑥𝑥 �kg m2 �
0.055744
2
Moment of inertia 𝐼𝑦𝑦 �kg m �
0.058427
Moment of inertia 𝐼𝑧𝑧 �kg m2 �
0.068685
1-st rotor location r1 [m]
[0.195, 0.135, 0]]
2-nd rotor location r2 [m]
[-0.195, 0.135, 0]
3-rd rotor location r3 [m]
[-0.195, -0.135, 0]
4-th rotor location r4 [m]
[0.195, -0.135, 0]
Propeller diameter [m]
0.3048
3.2.
Drones description
Both Leader’s (Figure 3) and Follower’s (Figure 4) are quadrotors in X-configuration propelled by brushless direct-current (BLDC) motors, although they differ in size. Leaders are bigger because a quite sophisticated autopilot was installed on the board to ensure appropriate computational resources. The directions of the propeller rotations are presented in Figure 2. Rotors 1 and 3 rotate counterclockwise (when looking from the top), and rotors 2 and 4 rotate clockwise. The basic data of the Leaders are presented in Table 1 and Followers in Table 2, respectively. Model parameters were obtained experimentally. Moments of inertia of the drones were measured using trifilar pendulum. The characteristics of the propellers were collected by static tests. The maneuverability of the drones was limited programmatically in the autopilot settings. For the Leader, the roll and pitch angles fell into the range of ±30∘ . Ascent speed was 3 m/s, descent 3.4 m/s. For the Follower the roll and pitch angles are limited to ±35∘ , ascent/descent speed is ±3 m/s. 3.3.
Drones simulation
The realistic dynamic model is crucial for understanding the swarming drone behavior. It was Table 2. Followers’ parameters
3. Drone’s Dynamics
3.1.
Coordinate systems
A set of coordinate systems was introduced to describe the quadrotor motion (Figure 2). The origin of the body-fixed frame 𝑂𝑏 𝑥𝑏 𝑦𝑏 𝑧𝑏 was attached to the quadrotor center of mass, 𝑂𝑏 𝑥𝑏 axis is pointed forward, 𝑂𝑏 𝑦𝑏 to the right, and 𝑂𝑏 𝑧𝑏 down. North-EastDown (NED) coordinate system 𝑂𝑛 𝑥𝑛 𝑦𝑛 𝑧𝑛 was used to define the position of the drone in space. Additionally, gravity frame 𝑂𝑔 𝑥𝑔 𝑦𝑔 𝑧𝑔 was used to obtain the loads from the gravity force. 106
Parameter
Value
Mass [kg]
0.274
Moment of inertia 𝐼𝑥𝑥 �kg m2 �
0.000466
Moment of inertia 𝐼𝑦𝑦 �kg m2 �
0.000476
2
Moment of inertia 𝐼𝑧𝑧 �kg m �
0.002024
1-st rotor location r1 [m]
[0.0775, 0.0775, 0]
2-nd rotor location r2 [m]
[-0.0775, 0.0775, 0]
3-rd rotor location r3 [m]
[-0.0775, -0.0775, 0]
4-th rotor location r4 [m]
[0.0775, -0.0775, 0]
Propeller diameter [m]
0.1016
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 1. Top-level block diagram of the swarm simulator
Figure 2. Coordinate systems used in the simulation
Figure 4. Follower
𝐼𝑥𝑥 𝑃̇ − (𝐼𝑦𝑦 − 𝐼𝑧𝑧 )𝑅𝑄 = 𝐿𝑏
(4)
𝐼𝑦𝑦 𝑄̇ − (𝐼𝑧𝑧 − 𝐼𝑥𝑥 )𝑃𝑅 = 𝑀𝑏
(5)
𝐼𝑧𝑧 𝑅̇ − (𝐼𝑥𝑥 − 𝐼𝑦𝑦 )𝑃𝑄 = 𝑁𝑏
(6)
where 𝐿𝑏 , 𝑀𝑏 , 𝑁𝑏 are torques expressed in bodyfixed axes. The net forces and moments were calculated as a sum of gravity, propulsion, and aerodynamic loads. The gravity forces are:
Figure 3. Leader (with propellers removed) assumed that the sensors could perfectly measure the actual drone flight parameters. Also, the modeling of drone communication was omitted to simplify the analysis. The equations of motion of the quadrotor are: 𝑚(𝑈̇ + 𝑊𝑄 − 𝑉𝑅) = 𝑋𝑏
(1)
𝑚(𝑉̇ + 𝑈𝑅 − 𝑊𝑃) = 𝑌𝑏
(2)
𝑚(𝑊̇ + 𝑉𝑃 − 𝑈𝑄) = 𝑍𝑏
(3)
where 𝑚 is drone mass, 𝑈, 𝑉, 𝑊 are linear velocities, 𝑃, 𝑄, 𝑅 are angular rates, 𝑋𝑏 , 𝑌𝑏 , 𝑍𝑏 are forces in body-fixed frame 𝑂𝑏 𝑥𝑏 𝑦𝑏 𝑧𝑏 . For rotational dynamics:
𝑋𝑔 = −𝑚𝑔 sin Θ
(7)
𝑌𝑔 = 𝑚𝑔 sin Φ cos Θ
(8)
𝑍𝑔 = 𝑚𝑔 cos Φ cos Θ
(9)
Since the origin 𝑂𝑏 of the body-fixed frame coincides with the center of mass, then the resulting torque from gravity is zero. Thrust force and torque generated by propellers were calculated as [28]: 𝑇𝑗 = 𝜌𝑆𝑝 𝑅𝑝2 Ω2𝑗 𝑘𝑓
(10)
𝑀𝑗 = 𝜌𝑆𝑝 𝑅𝑝3 Ω2𝑗 𝑘𝑚
(11) 107
Journal of Automation, Mobile Robotics and Intelligent Systems
where 𝜌 is the air density, 𝑆𝑝 is the rotor area, 𝑅𝑝 is the propeller radius, Ω𝑗 is the angular rate of the 𝑗th propeller, 𝑘𝑓 is the thrust coefficient, and 𝑘𝑚 is the torque coefficient. Motor dynamics was modeled as a first-order system. The maximum possible angular rate of the propellers was limited using saturation. The autopilot structure consists of four control channels (roll, pitch, yaw, and altitude). Proportionalintegral-derivative (PID) controllers (with antiwindup protection) were used in each control channel. The PID controllers’ settings were tuned manually in flight tests of the actual objects and then used in the model. A detailed description of the mathematical model and the controller structure can be found in [29]. The implementation of the drone’s dynamics is presented in Figure 5. The drone dynamic was implemented using ”For Each Subsystem” block. In that way, it was possible to realize the parallel simulations of multiple drones. Aerodynamic characteristics were tabulated as functions of angles of attack and sideslip. Finally, the 6DOF drone dynamic model was optimized to achieve high computational efficiency. A built-in Simulink option called ”Accelerator mode” was used to reduce the time of numerical simulations. All unnecessary blocks were removed because they can significantly affect the simulation efficiency. This issue is crucial for the overall performance of the swarm simulator. Next, the developed simulation module was validated using data from real experiments in outdoor conditions. A series of flight tests of the individual drones took place to obtain the data logs and identify the coefficients of the model. Two kinds of data logs were available. The data were logged on the onboard memory card and additionally sent by the telemetry to the ground control station. These files were analyzed using Mission Planner software and the UAV Log Viewer online service. Next, the gathered data were used to validate the developed models of Leader and Follower. Details about the model validation methodology can be found in [30].
4. Swarm Control Algorithm The swarm control algorithm is divided into two parts – one for controlling the Leaders’ movement and the other for controlling the Followers’ movement. The swarm is divided into formations, usually based on a circle, consisting of one Leader and several Followers. The control algorithms’ outputs are commanded velocities, which are calculated for each drone as a weighted sum of several components depending on a few simple rules presented in Table 3 and Table 4 (see also Figure 6). The weights used to calculate the output velocity depend on the drone’s state and are tuned manually based on the simulation results. This velocity is then inputted to the drone’s autopilot or directly integrated when using the kinematic model. 108
VOLUME 20,
N∘ 3
2026
Table 3. Leaders’ control rules Rule
Description
Alignment
It ensures that the Leaders fly at an average velocity of all Leaders.
Cohesion
It ensures that the Leaders fly in a flock.
Separation
It ensures that the Leaders do not collide with each other.
Avoidance
It ensures that the Leaders do not collide with known, static obstacles.
Targeting
It ensures that the Leaders fly through their designated waypoints.
Table 4. Followers’ control rules Rule
Description
Alignment
It ensures that the Followers fly with the velocity of their Leader.
Cohesion
It ensures that the Followers tend to move to their Leader’s position.
Separation
It ensures that the Followers do not collide with each other within their formation.
Avoidance
It ensures that the Followers do not collide with known, static obstacles.
Formation
It ensures that the Followers keep the designated distance from each other within their formation.
In the following sections, the details about the drones’ states and the transition between them, as well as the control rules for each of the states, will be provided, as it is a core of the developed algorithm that allows the efficient execution of the proposed mission. 4.1.
Leaders’ states and transition conditions
Leaders can be assigned one of the six states, which are described in Table 5. The transition conditions between them are presented in the following paragraphs. Every Leader starts the mission in the Takeoff state. Transition from “1 – Takeoff” to “0 – Cruise” While in the “Takeoff” state, the Leader awaits in the base for the ascend command for the predefined takeoff delay time, which allows the Leaders to start in a sequence. Then the Leader ascends to a set, intermediate altitude and awaits for the Followers to join and form a formation. After that, the Leader starts the flight through its designated waypoints. Transition from “0 – Cruise” to “2 – Landing”
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 5. Drone’s dynamic model developed in Simulink Table 5. Leaders’ states No.
Name
Description
0
Cruise
Main state in which the Leader flies through its designated waypoints. The Leader is allowed to wait in each of them for a specified amount of time, before flying to the next one.
1
Takeoff
State that is applied to the Leaders at the beginning of the mission. It indicates that the Leader is waiting in the base for the time when it is supposed to ascend or that it is ascending and waiting for the Followers to join it.
2
Landing
State in which the Leader, after reaching the last waypoint, is descending until it touches the ground.
3
Malfunction
State in which all communication with the Leader was lost and it is excluded from further calculations.
4
Standby
State in which the Leader awaits in the base for further instructions.
6
Return
State in which the Leader returns to the base.
After the Leader reaches the last waypoint and the predefined waiting time for this waypoint passes, it transitions to the “Landing” state.
Other states No transitions have been implemented between any other states yet. 4.2.
Transition from “2 – Landing” to “4 – Standby” While in the “Landing” state, the Leader descends until it reaches a threshold altitude above the ground, where it waits for the Followers to achieve the Landing formation. Then the Leader starts descending further until touching the ground when it transitions to “Standby” state.
Followers’ states and transition conditions
Followers can be assigned one of the seven states presented in Table 6. The transition conditions are presented in the following paragraphs. The Followers assigned to any of the formations at the beginning of the mission are starting in the “Takeoff” state. The Followers that will be replacing drones during reconfiguration are waiting in the base in the “Standby” state. 109
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 6. Followers’ states No.
Name
Description
0
Cruise
Main state in which the Follower flies in a formation and follows its Leader.
1
Takeoff
State that is applied to the Followers at the beginning of the mission. It indicates that the Follower is waiting in the base for the time of ascend or that it is ascending vertically before joining its Leader and forming a formation.
2
Landing
State in which the Follower is descending until it touches the ground.
3
Malfunction
State in which all communication with the Follower was lost and it is excluded from further calculations.
4
Standby
State in which the Follower awaits in the base for further instructions.
5
Flight to the Leader
State in which the Follower performs an independent flight to the Leader after takeoff from the base when one of the formations needs reconfiguration.
6
Return
State in which the Follower returns to the base after one of the formations needed reconfiguration.
Reconfiguration transitions – from “0 – Cruise” to: “2 – Landing”, “3 – Malfunction”, or “6 – Return” There are three types of formation reconfiguration implemented that can happen during the mission. The condition that triggers these transitions in the simulation is the predefined time at which the event is supposed to happen (which can be changed to any other type of event, for example low battery condition during the real mission). In the case of the reconfiguration, the formation can be adjusted depending on the mission scenario. It can either reshape into a different circle-based figure, taking into account the decreased number of Followers, or stay the same by keeping one or more of its vertices empty. The Follower leaving the formation can either return to base, land directly below or become malfunctioned by transitioning to the appropriate state. Figure 6. Leaders’ control rules applied in Simulink Transition from “6 – Return” to “2 – Landing” Transition from “1 – Takeoff” to “0 – Cruise” The Follower that is assigned to one of the formations at the beginning of the mission awaits in the base for its Leader to start the ascend. Then, it ascends vertically to the intermediate altitude before transitioning to “Cruise” state. Transition from “0 – Cruise” to “2 – Landing” The Follower transitions from “Cruise” to “Landing” when its Leader transitions to “Landing” state, which happens when the Leader reaches it’s last waypoint. Transition from “2 – Landing” to “4 – Standby” During the landing phase, the Follower reaches a threshold altitude above ground while forming a landing formation. Then it descends until it touches the ground and transitions to “Standby” state. 110
If the Follower transitioned into the “Return” state it flew back to base. After reaching the base landing spot coordinates with appropriate accuracy, it transitions to “Landing” state. Transition from “4 – Standby” to “1 – Takeoff” When one of the Followers starts the reconfiguration and detaches from it’s formation, the other Follower that is assigned to fill a spot of a missing member of one of the formations gets the signal to ascend, it transitions to “Takeoff” to reach a designated altitude. Transition from “1 – Takeoff” to “5 – Flight to the Leader” After reaching the designated altitude during the ascend, the Follower transitions to “Flight to the Leader” state and flies towards the Leader.
Journal of Automation, Mobile Robotics and Intelligent Systems
Transition from “5 – Flight to the Leader” to “0 – Cruise”
After the Follower reaches the Leader, it transitions to “Cruise”, attaches to the formation, and begins its flight following the Leader.
VOLUME 20,
𝑁
v𝑆 = 𝑤𝑆 �
There are no transitions implemented between any other states yet. Leader’s trajectory control
The Leader’s control algorithm is comprised of a set of rules from Table 3 and differs depending on the state in which the Leader currently is. Before going into the autopilot is it caped at the predefined maximum allowable speed to prevent controller’s misbehavior, see also Figure 6. The control rules are straightforward due to the fact that they will need to be executed on small microcontrollers inside the drone, with very limited memory and computation capabilities. Cruise state In the Cruise state, which is a nominal state for a Leader, the control algorithm consists of five rules: Targeting, Alignment, Cohesion, Separation, and Avoidance. The resulting commanded velocity is a weighted sum of all five rules. The Targeting rule’s purpose is to steer the drone into the direction of the next waypoint. The commanded velocity for this rule is calculated as: r𝑇 − r v𝑇 = 𝑤𝑇 |r𝑇 − r|
(12)
where 𝑤𝑇 is the weight, r𝑇 is a waypoint position, and r is the position of the drone. When the drone is closer to the waypoint than the predefined threshold value, the target is switched to the next waypoint. The drone can also wait at the waypoint for a specified amount of time. The Alignment rule makes the drone fly with the average velocity of all of the Leaders in the average direction. The commanded velocity is calculated as: v𝐴 = 𝑤𝐴 �
∑𝑁 𝑖=1 v𝑖 � 𝑁
v𝐶 = 𝑤𝐶 �
∑𝑁 𝑖=1 r𝑖 − r� 𝑁
𝑖=1
where 𝑤𝐶 is the weighting factor, 𝑟𝑖 is the position of the 𝑖-th Leader, 𝑁 is the number of Leaders, and r is the position of the drone.
(15)
where 𝑤𝑅 is the weighting coefficient, which is adaptive based on the distance between the drone and the obstacle, 𝑟𝑃𝑖 is the position of the 𝑖-th obstacle, 𝑁 is the number of obstacles, and r is the position of the drone. Takeoff state In the Takeoff state, which is an initial state for a Leader, the control algorithm consists of two rules: Targeting and Separation. The resulting commanded velocity is a weighted sum of the rules. In this state, the Leader is waiting in a base for a Takeoff command and then ascending to the intermediate height to weight for the Followers to join and form a formation. The Separation condition is identical to the one described in the Cruise state paragraph. The Targeting rule’s purpose is to ascend the drone to an intermediate height. The commanded velocity for this rule is calculated as: v𝑇 = �0
(14)
1 r − r𝑖
where 𝑤𝑆 is the weighting factor, which is adaptive based on the distance between drones, 𝑟𝑖 is the position of the 𝑖-th Leader, 𝑁 is the number of Leaders, and r is the position of the drone. This rule has additional conditions. It is switched off whenever the drone is further away from any other Leaders than the specified threshold or if there is a sufficient difference in height at which two Leaders are flying. The last rule is the Avoidance. This rule prevents the Leader from colliding with stationary obstacles of a known position. The commanded velocity is calculated as: 𝑁 r − r𝑃𝑖 v𝑅 = 𝑤𝑅 � (16) 2 𝑖=1 �r − r𝑃𝑖 �
(13)
where 𝑤𝐴 is the weight, 𝑣𝑖 is the velocity of the 𝑖-th Leader, and 𝑁 is the number of Leaders. The Cohesion rule forces the drone to fly to the geometric center of all of the Leaders’ positions. The overall purpose of this rule is to keep the drones in a group or flock. The commanded velocity is given by:
2026
The Separation rule forces the drone to avoid collisions between any other Leaders. Also, due to the fact that every Leader is surrounded by the formation of Followers, this rule avoids the collisions between individual formations keeping the Leaders separated by a distance greater than the sum of their formation radii. Its commanded velocity is given as:
Other states
4.3.
N∘ 3
0
ℎ −𝑟𝑧 𝑇 � 𝑇 −𝑟𝑧 |
𝑤𝑇 |ℎ𝑇
(17)
where, 𝑤𝑇 is a weighting coefficient, ℎ𝑇 is the intermediate height, and 𝑟𝑧 is the vertical component of the position of the drone. Landing state In the Landing state, the control algorithm consists of only one rule: Targeting. In this state, the Leader is located at its last specified waypoint. It then descends to the intermediate altitude and waits for the Followers to achieve the landing formation and then descends further until reaching the ground. 111
Journal of Automation, Mobile Robotics and Intelligent Systems
The commanded velocity for the Targeting rule in this stat is calculated as the one described in the Takeoff paragraph for the first part of the maneuver and for the second part with the equation: v𝑇 = 𝑤𝑇
−r |r|
(18)
where, 𝑤𝑇 is the weight and r is the drone’s position. Return state In this state, the Leader is flying back to the base. The control algorithm consists of three rules: Targeting, Separation, and Avoidance. All of the rules are almost identical to the ones described in the paragraph regarding the Cruise state. The only difference is in the Targeting rule, where the waypoint’s position is replaced by the base landing coordinates. The Leader is flying directly above the Landing place and then switches to the Landing state where it performs the Landing maneuver. Other states
VOLUME 20,
Follower’s trajectory control
The control algorithm of a Follower is comprised of a set of rules from Table 4, very similar to the rules for a Leader, and differs depending on the state in which the Follower currently is. Before going into the autopilot it is also capped at the predefined maximum allowable speed to prevent the controller’s misbehavior.
2026
where 𝑀𝑥 , 𝑀𝑦 , 𝑀𝑧 are the mutual distance matrices that hold the distances between each Follower in 𝑥, 𝑦, and 𝑧 position coordinates that are calculated using the equations: 𝑀𝑥 (𝑖, 𝑗)
=
𝑅 �sin 𝛼𝑗 − sin 𝛼𝑖 �
(21)
𝑀𝑦 (𝑖, 𝑗)
=
𝑅 �cos 𝛼𝑗 − cos 𝛼𝑖 �
(22)
𝑀𝑧 (𝑖, 𝑗)
=
0
(23)
where 𝑅 is the radius of the formation and 𝛼 is the angular position of the Follower in that formation given as: 2𝜋 (𝑖 − 1) + 𝜃 𝛼𝑖 = (24) 𝑁 where 𝜃 is the formation angular offset. The Alignment rule makes the Follower fly with the velocity vector of its Leader. The commanded velocity is calculated as: v𝐴 = 𝑤𝐴 v𝐿 (25) where 𝑤𝐴 is the weight, and 𝑣𝐿 is the velocity of the Leader. The Cohesion rule is forcing the Follower to fly to the position of the Leader. The overall purpose of this rule is to keep the formation around the Leader during the flight. The commanded velocity is given by: v𝐶 = 𝑤𝐶 �r𝐿 −
There are no algorithms implemented in “Malfunction” and “Standby” states. 4.4.
N∘ 3
∑𝑁 𝑖=1 r𝑖 � 𝑁
(26)
where 𝑤𝐶 is the weight, 𝑟𝐿 is the position of the formation Leader, 𝑁 is the number of Followers in the formation, and r𝑖 is the position of the 𝑖-th Follower. The Separation rule forces the drone to avoid collision with any other Followers inside the formation. Its commanded velocity is given by the equation analogous to the Leader’s rule. The Avoidance rule prevents the Follower from colliding with stationary obstacles of a known position. The commanded velocity is calculated in the same way as for the Leader.
Cruise state Takeoff state In the Cruise state, which is a nominal state for a Follower, the control algorithm consists of five rules: Formation, Alignment, Cohesion, Separation, and Avoidance. The resulting commanded velocity is a weighted sum of all five rules. The Formation rule is responsible for keeping all of the Followers of a particular Leader in a formation. It makes the followers keep the desired distances between them during the flight. The equation for the commanded velocity is given as: 𝑁
v𝐹 = 𝑤𝐹 � �𝑘𝑝 (r𝑖 − r − M) + 𝑘𝑑 (v𝑖 − v)�
(19)
𝑖=1
where 𝑤𝐹 , 𝑘𝑝 , and 𝑘𝑑 are weighting factors, 𝑁 is the number of Followers in the formation, r𝑖 , v𝑖 are the position and the velocity of every Follower in the formation, and r, v are the position and velocity of the current Follower. The M is given by: M = �𝑀𝑥 (𝑖, 𝑗) 112
𝑀𝑦 (𝑖, 𝑗)
𝑇
𝑀𝑧 (𝑖, 𝑗)�
(20)
This state is separated between two types of Followers. The first type is the Follower which starts as a member of a formation. This type of Follower in this state is firstly waiting for the Takeoff of its Leader and then is ascending vertically for a few seconds before transitioning to the Cruise state. The second type of the Follower is a part of the reconfiguration procedure. This type obeys two rules in this state: Targeting and Separation. The latter one is almost the same as for the Cruise state, with the difference that the Follower is avoiding the collision between other Followers in the Reconfiguration condition. The Targeting rule’s purpose is to make the Follower achieve the appropriate altitude at which it will be continuing the flight toward the Leader. The commanded velocity for this rule is calculated as: v𝑇 = �0
0
ℎ −𝑟𝑧 𝑇 � 𝑇 −𝑟𝑧 |
𝑤𝑇 |ℎ𝑇
(27)
where, 𝑤𝑇 is a weight, ℎ𝑇 is the commanded altitude and 𝑟𝑧 is the vertical position of the drone.
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Landing state This state is the same as the one described for the Leader. Flight to the Leader state This state is a part of the Reconfiguration process. In this state, the Follower is flying towards the Leader in order to join its formation. It obeys three rules: Targeting, Separation, and Avoidance. The Avoidance rule is the same as for the Leader, and the Separation is the same as in the Takeoff state for the Follower. The Targeting’s rule purpose is to steer the drone in the direction of a Leader. The commanded velocity for this rule is calculated as: v𝑇 = 𝑤𝑇
r𝐿 − r |r𝐿 − r|
(28)
where, 𝑤𝑇 is a weight, r𝐿 is a Leader position with the offset on the vertical position in order to avoid collisions with other flying formations, and r is the position of the Follower. When the drone is closer to the Leader than the predefined threshold it joins the formation and switches to the Cruise state. Return state This state is a part of the Reconfiguration process. In this state, the Follower is flying towards the base. It obeys three rules: Targeting, Separation, and Avoidance. All of the rules are the same as for the Flight to the Leader state. In the Targeting rule the target is the base landing point. The Follower is flying at a different altitude than all of the formations and Followers in the Flight to the Leader state in order to avoid collisions. The Targeting rule’s purpose is to steer the drone in the direction of a Leader. The commanded velocity for this rule is calculated as: v𝑇 = 𝑤𝑇
r𝐿 − r |r𝐿 − r|
(29)
where, 𝑤𝑇 is a weight, r𝐿 is a Leader position with the offset on the vertical position in order to avoid collisions with other flying formations, and r is the position of the Follower. When the drone is closer to the Leader than the predefined threshold it joins the formation and switches to the Cruise state. Other states There are no algorithms implemented in “Malfunction” and “Standby” states.
5. Mission Description The purpose of the prepared swarm control algorithm is to cover a specified area with a grid of drones for the purpose of establishing a WiFi network in emergency situations. In order to take into account the battery life of the drones the grid is filled using a “snake” pattern, as shown in Figure 7.
Figure 7. The snake pattern for a grid coverage
There are fifteen cells in the prepared grid which will need fifteen Leaders to fully cover it. Each of the Leaders has six Followers in the formation which gives 105 drones in total. Each of the formations will start sequentially from the same place marked with a red dot in Figure 7, and will follow the pattern until it reaches the last way-point shown. In all of the way-points, the Leaders will wait for a specified amount of time, dependent on the battery life and the overall required length of the mission. To avoid collisions between the formations, especially in the turns, and to preserve the necessary distances between the drones for communication and WiFi coverage purposes, every other formation will be flying at a different altitude, even-numbered formations higher and odd-numbered formations lower than the way-points’ defined height. After reaching the last way-point shown (second to last of the mission), the formations will proceed to the landing zone obeying the rules of the formation landing sequence. In the next section, the detailed simulation results will be presented. In this particular scenario, each formation starts with a delay of 90 seconds with respect to the previous one and waits at each of the way-points for 25 seconds. Every formation has a radius of 50 meters during the flight, which shrinks to 5 meters during the landing phase. The distance between every way-point is equal to 132 meters and between the second to last and the last to 561 meters. The way-points are located at the height of 120 meters, even formations are flying 5 meters above them, and the odd ones 5 meters below, to avoid collisions. The formations are moving at a speed of 3 meters per second.
6. Simulation Results The swarm control algorithm was run for the mission described in the previous section. The computations were performed on the standard desktop computer (equipped with an Intel Core i7 CPU and 32 GB RAM). The formations were starting sequentially from the same starting position, as shown in Figure 8. The Leaders are marked purple, 113
Journal of Automation, Mobile Robotics and Intelligent Systems
N∘ 3
2026
Figure 8. Beginning of the mission
Figure 10. All of the drones are in the air
and the Followers are light blue, the way-points are shown as green circles. The trajectory is made of 15 way-points with one additional, different for every Leader, in the top left corner reserved for landing. All of the formations follow the same path, only the landing points are unique for each of the formations. Additionally, at the top of the graph, the actual time (in seconds) of the mission is presented. An important aspect of the algorithm is the precise tuning of all of the algorithm’s parameters. To cover up the whole grid at the same time, it is necessary to specify the start time delay for all of the Leaders and their way-points wait time. As can be seen in Figure 9 the drones are starting to fill out the space, keeping the necessary distances required for effective communication and collision avoidance. When all of the drones are in the air, see in Figure 10, all of the way-points are fully covered. After completing the main part of the trajectory, the formations proceed to the last waypoint, where they perform the landing procedure, as shown in Figure 11. The formations are reducing their radii, and the whole group is landing in the set coordinates. The whole mission took 50 minutes to complete. Because of all of the control rules and the spread of altitudes between the
respective formations, there are no collisions during the flight.
Figure 9. Drones filling up the grid 114
VOLUME 20,
Figure 11. Return to the landing zone
7. Conclusions In the presented study, a full 6DOF drone swarm simulator created in MATLAB/Simulink was presented and described. An efficient swarm control algorithm was developed that allows it to cover a specified area with a grid of drones. The simulator allows taking into account environmental conditions like wind, terrain heights, and drone battery supply for the purpose of establishing a WiFi network in emergency situations. Several contributions in the field of an autonomous swarm of drones might be mentioned. The results of the prepared simulation show that the area can be effectively and fully covered by the drone formations, without in-flight collisions, and keeping the necessary distances between the drones. In the considered scenario, the Leaders moved along a predefined flight path. It was found that regular hexagons are the most suitable formation shapes to cover the whole region. The developed simulator has a modular structure and is highly customizable. The data of other quadrotors can be easily implemented. Moreover, the developed software is scalable and allows for performing
Journal of Automation, Mobile Robotics and Intelligent Systems
simulations for a large number of drones. The software user can define various simulation scenarios. Calculations can be easily performed on the standard desktop computer without the need to use sophisticated computational resources (like workstations or clusters). This is a significant advantage when compared to existing simulators (for example, Gazebo). However, one of the most challenging issues is to set the appropriate parameters of the proposed control algorithm. In the presented study, these parameters were tuned manually by trial and error method, which is quite a time-consuming task. When some significant changes in the system configuration are introduced, then these weights must be tuned again to meet the desired performance. In the future, automatic tuning methods should be used to make this process more effective. It is possible to achieve complicated swarm behavior using a set of relatively simple mathematical rules. This feature is important from a practical point of view because of the limited onboard computational resources. Simple algorithms can be easily implemented in the autopilot software. The mission was completed in about 50 minutes. In the real system, the drone must be equipped with an onboard battery with the appropriate capacity to realize the whole mission. This is a significant problem when small drones are used as Followers because they cannot carry heavy batteries and must be replaced by drones with fully charged batteries during the mission, complicating the task. Future plans include modeling the drone battery usage and using that information to drive the reconfiguration process during the simulation. The developed simulator can be practically used for flight test planning purposes. In the near future, the drones’ dynamic model and proposed swarm control algorithm will be validated using the data obtained experimentally during flight tests in outdoor conditions. The influence of drone failure on the performance of the whole formation should be studied in detail. The effects of communication delays and signal losses must also be investigated.
AUTHORS Dariusz Miedziński – Institute of Aeronautics and Applied Mechanics, Warsaw University of TechnologyNowowiejska 24, 00-665 Warsaw, e-mail: dariusz.miedzinski@pw.edu.pl. Mariusz Jacewicz∗ – Institute of Aeronautics and Applied Mechanics, Warsaw University of TechnologyNowowiejska 24, 00-665 Warsaw, e-mail: mariusz.jacewicz@pw.edu.pl. Kacper Kaczmarek – Faculty of Power and Aeronautical Engineering, Warsaw University of TechnologyNowowiejska 24, 00-665 Warsaw, e-mail: kacper.kaczmarek3.stud@pw.edu.pl.
VOLUME 20,
N∘ 3
2026
Sebastian Topczewski – Institute of Aeronautics and Applied Mechanics, Warsaw University of TechnologyNowowiejska 24, 00-665 Warsaw, e-mail: sebastian.topczewski@pw.edu.pl. Robert Głębocki – Institute of Aeronautics and Applied Mechanics, Warsaw University of TechnologyNowowiejska 24, 00-665 Warsaw, e-mail: robert.glebocki@pw.edu.pl. Antoni Kopyt – Institute of Aeronautics and Applied Mechanics, Warsaw University of TechnologyNowowiejska 24, 00-665 Warsaw, e-mail: antoni.kopyt@pw.edu.pl. ∗
Corresponding author
ACKNOWLEDGEMENTS This work was supported by the National Centre of Research and Development under POIR programme under research project agreement POIR.01.01.01-000040/22
References [1] F. Saffre, H. Hildmann, and H. Karvonen, “The design challenges of drone swarm control”. In: Engineering Psychology and Cognitive Ergonomics: 18th International Conference (EPCE 2021), 2021, 408–426, 10.1007/978-3030-77932-0_32. [2] M. Campion, P. Ranganathan, and S. Faruque, “UAV swarm communication and control architectures: a review”, Journal of Unmanned Vehicle Systems, vol. 7, no. 2, 2019, 93–106, 10.1139/juvs-2018-0009. [3] S. Cohen and N. Agmon, “Recent advances in formations of multiple robots”, Current Robotics Reports, vol. 2, 2021, 1–17, 10.1007/s43154021-00049-2. [4] S. Oweis, S. Ganesan, and K. C. Cheok, “Server based control flocking for aerialsystems”. In: IEEE International Conference on Electro/Information Technology, vol. 1, 2014, 314–319, 10.1109/EIT.2014.6871783. [5] R. Olfati-Saber, “Flocking for multi-agent dynamic systems: algorithms and theory”, IEEE Transactions on Automatic Control, vol. 51, no. 3, 2006, 401–420, 10.1109/TAC.2005.864190. [6] B. Walter, A. Sannier, D. Reiners, and J. Oliver, “UAV swarm control: Calculating digital pheromone fields with the GPU”, The Journal of Defense Modeling and Simulation: Applications, Methodology, Technology, vol. 3, 2006, 10.1177/154851290600300304. [7] V. Parunak, M. Purcell, and R. O’Connell, “Digital pheromones for autonomous coordination of swarming UAV’s”. In: Proceedings of the 1st AIAA Unmanned Aerospace Vehicles, Systems, Technologies, and Operations Conference, vol. 1, 2002, 10.2514/6.2002-3446. 115
Journal of Automation, Mobile Robotics and Intelligent Systems
[8] J. A. Sauter, R. Matthews, H. Van Dyke Parunak, and S. A. Brueckner, “Performance of digital pheromones for swarming vehicle control”. In: Proceedings of the Fourth International Joint Conference on Autonomous Agents and Multiagent Systems, 2005, 903–910, 10.1145/1082473.1082610. [9] A. Antonio Luca, M. G. C. Cimino, N. De Francesco, A. Lazzeri, M. Lega, and G. Vaglini, “Swarm coordination of mini-UAVs for target search using imperfect sensors”, Intelligent Decision Technologies, vol. 12, 2018, 1–14, 10.3233/IDT-170317. [10] M. G. C. A. Cimino, A. Lazzeri, and G. Vaglini, “Combining stigmergic and flocking behaviors to coordinate swarms of drones performing target search”. In: 2015 6th International Conference on Information, Intelligence, Systems and Applications (IISA), vol. 1, 2015, 1–6, 10.1109/IISA.2015.7387990. [11] A. Das, R. Fierro, V. Kumar, J. Ostrowski, J. Spletzer, and C. Taylor, “A vision-based formation control framework”, IEEE Transactions on Robotics and Automation, vol. 18, no. 5, 2002, 813–825, 10.1109/TRA.2002.803463. [12] M. A. Toksö z, S. Oğ uz, and V. Gazi, “Decentralized formation control of a swarm of quadrotor helicopters”. In: 2019 IEEE 15th International Conference on Control and Automation (ICCA), 2019, 1006–1013, 10.1109/ICCA.2019.8899628. [13] G. Vá sá rhelyi, C. Virá gh, G. Somorjai, T. Nepusz, A. E. Eiben, and T. Vicsek, “Optimized flocking of autonomous drones in confined environments”, Science Robotics, vol. 3, no. 20, 2018, eaat3536, 10.1126/scirobotics.aat3536. [14] D. H. Stolfi and G. Danoy, “An evolutionary algorithm to optimise a distributed UAV swarm formation system”, Applied Sciences, vol. 12, no. 20, 2022, 10.3390/app122010218. [15] A. Ló pez-Gonzá lez, J. Meda Campañ a, E. Herná ndez Martı́nez, and P. P. Contro, “Multi robot distance based formation using parallel genetic algorithm”, Applied Soft Computing, vol. 86, 2020, 105929, 10.1016/j.asoc.2019.105929. [16] H. Duan, Q. Luo, Y. Shi, and G. Ma, “Hybrid particle swarm optimization and genetic algorithm for multi-UAV formation reconfiguration”, IEEE Computational Intelligence Magazine, vol. 8, no. 3, 2013, 16–27, 10.1109/MCI.2013.2264577. [17] J. C. Derenick and J. R. Spletzer, “Convex optimization strategies for coordinating largescale robot formations”, IEEE Transactions on Robotics, vol. 23, no. 6, 2007, 1252–1259, 10.1109/TRO.2007.909833. [18] J. Alonso-Mora, E. Montijano, T. Nä geli, O. Hilliges, M. Schwager, and D. Rus, “Distributed 116
VOLUME 20,
N∘ 3
2026
multi-robot formation control in dynamic environments”, Autonomous Robots, vol. 43, 2019, 10.1007/s10514-018-9783-9. [19] J. Alonso-Mora, S. Baker, and D. Rus, “Multirobot navigation in formation via sequential convex programming”. In: 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), vol. 1, 2015, 4634–4641, 10.1109/IROS.2015.7354037. [20] W. Dunbar and R. Murray, “Model predictive control of coordinated multi-vehicle formations”. In: Proceedings of the 41st IEEE Conference on Decision and Control, 2002., vol. 4, 2002, 4631–4636, 10.1109/CDC.2002.1185108. [21] G. B. Lamont, J. N. Slear, and K. Melendez, “UAV swarm mission planning and routing using multi-objective evolutionary algorithms”. In: 2007 IEEE Symposium on Computational Intelligence in Multi-Criteria Decision-Making, vol. 1, 2007, 10–20, 10.1109/MCDM.2007.369410. [22] L. Pyke and C. Stark, “Dynamic pathfinding for a swarm intelligence based UAV control model using particle swarm optimisation”, Frontiers in Applied Mathematics and Statistics, vol. 7, 2021, 10.3389/fams.2021.744955. [23] Y. Fu, M. Ding, and C. Zhou, “Phase angleencoded and quantum-behaved particle swarm optimization applied to three-dimensional route planning for UAV”, IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans, vol. 42, no. 2, 2012, 511–526, 10.1109/TSMCA.2011.2159586. [24] V. Roberge, M. Tarbouchi, and G. Labonte, “Comparison of parallel genetic algorithm and particle swarm optimization for real-time UAV path planning”, IEEE Transactions on Industrial Informatics, vol. 9, no. 1, 2013, 132–141, 10.1109/TII.2012.2198665. [25] F. Augugliaro, A. P. Schoellig, and R. D’Andrea, “Generation of collision-free trajectories for a quadrocopter fleet: A sequential convex programming approach”. In: 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, vol. 1, 2012, 1917–1922, 10.1109/IROS.2012.6385823. [26] C. Pinciroli, V. Trianni, R. O’Grady, G. Pini, A. Brutschy, M. Brambilla, N. Mathews, E. Ferrante, G. Di Caro, F. Ducatelle, M. Birattari, L. M. Gambardella, and M. Dorigo, “ARGoS: a modular, parallel, multi-engine simulator for multirobot systems”, Swarm Intelligence, vol. 6, 2012, 271–295, 10.1007/s11721-012-0072-5. [27] E. Soria, F. Schiano, and D. Floreano, “SwarmLab: a Matlab drone swarm simulator”. In: 2020 IEEE/RSJ International Conference on Intelligent Robots and
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Systems (IROS), vol. 1, 2020, 8005–8011, 10.1109/IROS45743.2020.9340854. [28] P. Pounds, R. Mahony, and P. Corke, “Modelling and control of a large quadrotor robot”, Control Engineering Practice, vol. 18, no. 7, 2010, 691–699, 10.1016/j.conengprac.2010.02.008. [29] R. Głębocki, M. Ż ugaj, and M. Jacewicz, “Validation of the energy consumption model for a quadrotor using monte-carlo simulation”, Archive of Mechanical Engineering, vol. 70, no. 1, 2023, 151–178, 10.24425/ame.2022.144075. [30] M. Jacewicz, M. Ż ugaj, R. Głębocki, and P. Bibik, “Quadrotor model for energy consumption analysis”, Energies, vol. 15, no. 19, 2022, 10.3390/en15197136.
117
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
FAST AERIAL MANIPULATION OF A VARYING PAYLOAD Submitted: 21st November 2025; accepted: 2nd February 2026
Tariel Simonyan, Oleg Gasparyan DOI: 10.14313/jamris-2026-043 Abstract: This paper addresses the robust trajectory tracking problem of an Unmanned Aerial Vehicle (UAV) equipped with a 2-DOF manipulator, which is designed for fast aerial manipulation of varying payloads. To overcome the high computational cost and adaptability limitations of traditional model-based controllers, this work introduces a novel hybrid gain-scheduling framework that shifts the computational complexity to the pre-flight phase. The approach utilizes an approximate inverse dynamics linearization based on fixed nominal models, which transforms the complex nonlinear system into a simple linear plant with bounded, structured uncertainties. The entire configuration space, including manipulator states and a range of payload properties, is partitioned into dynamically similar regions using K-Means clustering. For each local region, a dedicated robust PD controller is designed using a multi-objective Genetic Algorithm (GA). This framework also successfully implements a gain interpolation technique to mitigate the potential for abrupt control actions. Simulation results validate the controller’s ability to maintain high-precision tracking during fast maneuvers and payload switching, confirming the robustness and adaptability of the offline-tuned design. Keywords: gain-scheduling, unmanned aerial vehicles, aerial manipulation, intelligent control system
1. Introduction The aerial manipulator is a new type of aerial system that can both fly and physically interact with the world. Aerial robots have attracted much interest for their increased mobility compared to ground robots, as they are not restricted by terrain and can navigate in hard-to-access locations [1]. These systems offer solutions in environments that can be too dangerous for a human operator, and are used in diverse fields like inspection, transportation, architecture, building, and construction, as well as in military applications. Nevertheless, the integration of a robotic arm with a floating base presents significant control difficulties. The system is subject to complex, timevarying dynamics coupling that results in strong nonlinearities. The motion of the manipulator and the interaction with varying payload continuously alter the system’s inertia tensor and shift the center of mass (CoM) during flight. Consequently, achieving robustly stable and precise control remains a complex problem. The current control literature is full of different control methods for those systems that have been 118
proposed to manage the coupled dynamics, including nonlinear, adaptive controllers and sophisticated compensation schemes. Despite these advances, existing control systems are still facing some limitations. To achieve robust stability, many designs including robust optimal methods like the classic nonlinear H∞ controller [2] — rely on explicit, continuous real-time calculation and compensation for the reaction forces induced by the robotic arm [3, 4]. Model-based solutions, derived from advanced techniques using numerical methods for dynamic compensation [5], increase computational cost and restrict the performance of those platforms [6]. Complex solutions — such as combining sliding mode control with adaptive laws and disturbance observers for simultaneous internal and external disturbance compensation [7, 8], or adaptive control methods combined with neural networks to manage high modeling error [8, 9] — are promising, but result in a large computational burden that can challenge a system’s stability during fast aerial manipulation maneuvers. Control stability is often dependent on strictlydefined mechanical parameters. For example, stability conditions often rely on critical assumptions, like keeping the mass of the manipulator sufficiently small with respect to the aerial vehicle [4] or requiring complex algorithms to estimate time-varying inertia when manipulating objects [10]. These inherent structures are not impervious to the changes introduced by grasping different payloads [8, 11]. While robust stability can be analytically proved using complex frameworks like nested saturation control (achieving input-to-state stability) [3, 4], these systems are designed with controller structures that are difficult to tune across significantly varying flight conditions. Based on preceding analysis, this study’s central research problem involves the trade-off between the high computational cost of real-time nonlinear compensation and the need for robust stability during fast maneuvers with varying payloads. Though current control methods for aerial manipulation can achieve stability, they typically rely on real-time calculations that create a significant computational burden, or lack the necessary robustness to handle rapid payload changes and fast maneuvers simultaneously. To address this problem, the research objectives are as follows:
Open Access. © 2026 Tariel Simonyan and Oleg Gasparyan, published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
angles 𝛶 = [𝜀1 , 𝜀2 ]𝑇 : 𝑇
𝑞=[𝑧, 𝜙, 𝜃, 𝜓, 𝜀1 , 𝜀2 ] ,
(1)
Unlike conventional UAVs, the dynamics of an aerial manipulator are intrinsically coupled to the arm’s motion, which displaces the CoM and alters the vehicle’s inertia tensor during the flight. The inputs driving the system are denoted as 𝜏 and combine the quadrotor’s aerodynamics with the servo torques: 𝜏=[𝑇Σ , 𝜏𝜙 , 𝜏𝜃 , 𝜏𝜓 , 𝜏𝑗1 , 𝜏𝑗2 ]𝑇 ,
Figure 1. UAV manipulator system
(2)
where 𝑇Σ represents the collective thrust of the four rotors: 4
𝑇∑ = � 𝑇𝑖 , i) To minimize the real-time computational cost associated with nonlinear dynamic compensation. ii) To enhance the adaptability of the control system to rapid shifts in payload characteristics and system inertia during fast maneuvers. iii) To guarantee high-precision trajectory tracking and robust stability throughout the whole operational envelope. The proposed solution is based on shifting the computational complexity to the pre-flight phase. By employing an inverse dynamics linearization [12] based on fixed, pre-computed nominal models, the complex nonlinear system can be transformed into a simple linear system with residual nonlinearities captured as bounded uncertainties. Those structured uncertainties are quantified using the Linear Fractional Transformation (LFT) framework [13]. The entire configuration space, including the manipulator state, the payload mass and the object dimensions, is divided into clusters using the K-Means algorithm, optimized via the Elbow Method [14]. For each local region, a PD controller is designed, and its gains are tuned using a multi-objective GA [15]. The reward function is tasked with minimizing the tracking error and the control effort, while maximizing the local robust stability margin using the Small Gain Theorem [13]. This method yields a control system that can tolerate grasping of various payloads in the defined range of object masses and dimensions.
2. System Dynamics The aerial platform consists of a quadrotor base carrying a 2-DOF manipulator with two revolute joints, as depicted in Fig. 1. The kinematic description of the system is based on two coordinate systems, {𝐼} inertial frame, and {𝐵} body-fixed frame, aligned with the principal axes of inertia. The transition from {𝐼} to {𝐵} is given by the orthogonal rotation matrix constructed from the UAV’s Euler angles in the ZYX convention. The system state is defined using a generalized coordinate vector 𝑞∈R6 , which consists of UAV’s height 𝑧 in {𝐼}, its attitude vector (roll, pitch, yaw) 𝛺=[𝜙, 𝜃, 𝜓] 𝑇 in {𝐵}, and the manipulator’s joint
(3)
𝑖=1
where each 𝑖-th rotor generates a thrust 𝑇𝑖 that is proportional to the square of angular velocity of rotors 𝜎𝑖 (i.e. 𝑇𝑖 = 𝑐𝑇 𝜎𝑖2 , 𝑐𝑇 > 0) and acts along the body-fixed axis 𝑍𝐵 [16, 17]. The 𝜏𝜙 , 𝜏𝜃 , 𝜏𝜓 are the quadrotor torques, and 𝜏𝑗1 , 𝜏𝑗2 are the torques of each joint of the manipulator. The thrust-to-torque mapping for the quadrotor can be expressed using the four-dimensional vector of thrusts 𝑇𝑖 , (𝑇𝑈𝐴𝑉 = [𝑇1 , 𝑇2 , 𝑇3 , 𝑇4 ]𝑇 ), as well as the quadrotor configuration matrix 𝐵𝑈𝐴𝑉 : �
𝑇Σ � = 𝐵𝑈𝐴𝑉 𝑇𝑈𝐴𝑉 . 𝜏𝑈𝐴𝑉
(4)
Given the needed controls 𝑇Σ and 𝜏, the equation (4) allows for the computation of the required thrusts 𝑇𝑖 (or, which is equivalent, the velocities 𝜎𝑖 ) of the rotors. The system’s torque-to-input mapping can be described as: 𝜏 = 𝐵𝑆𝑌𝑆 𝑇𝑆𝑌𝑆 ,
(5)
where 𝑇𝑆𝑌𝑆 is the combined vector of the quadrotor motor’s thrusts and the manipulator joint torques. 𝐵𝑆𝑌𝑆 for the X configuration quadrotor with a 2-DOF robotic arm will have the form: 1
⎡ ⎢ √2𝐿/2 𝐵𝑆𝑌𝑆 = ⎢ −√2𝐿/2 ⎢ −𝑘𝜓 0 ⎢ 0 ⎣
1 √2𝐿/2 √2𝐿/2 𝑘𝜓 0 0
1 −√2𝐿/2 √2𝐿/2 −𝑘𝜓 0 0
1 −√2𝐿/2 −√2𝐿/2 𝑘𝜓 0 0
0 0 0 0 1 0
0 ⎤ 0 ⎥ 0 ⎥ , 0 ⎥ 0 ⎥ 1 ⎦
where 𝐿 is the length of the quadrotor arms, and 𝑘𝜓 (𝑘𝜓 > 0) is the drag coefficient. The equations of motion of the system are derived using the Euler-Lagrange method [18], and have the standard form: 𝑀 (𝑞) 𝑞̈ + 𝐶 (𝑞, 𝑞) ̇ 𝑞̇ + 𝐺 (𝑞) = 𝜏 + 𝜏𝐷 .
(6)
where 𝑀 (𝑞) is 𝑅6𝑥6 symmetric, positive-definite inertia matrix, 𝐶 (𝑞, 𝑞) ̇ contains Coriolis and centrifugal terms, 𝐺 (𝑞) is the gravity vector, and 𝜏𝐷 is the disturbance force. 119
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 3. Block diagram of the linearized system Figure 2. Block diagram of the UAV manipulator system 𝑁̃ (𝑞, 𝑞) ̇ = 𝑁̂ (𝑞, 𝑞) ̇ − 𝑁 (𝑞, 𝑞) ̇ . To facilitate the control design, the Coriolis, centrifugal, gravitational and disturbance terms are combined into a single nonlinear vector 𝑁 (𝑞, 𝑞). ̇ Thus, equation (6) can be rewritten as 𝑀 (𝑞) 𝑞̈ + 𝑁 (𝑞, 𝑞) ̇ = 𝜏.
(7)
This equation is linear in the control 𝜏, and has full rank 𝑀 (𝑞), which can be inverted for any valid configuration. Taking the control 𝜏 as a function of the manipulator states in the form: 𝜏 = 𝑀 (𝑞) 𝑦 + 𝑁 (𝑞, 𝑞) ̇ ,
(8)
leads to the system described by 𝑞̈ = 𝑦
(9)
𝑦=𝐾𝐷 (𝑞̇ 𝑑𝑒𝑠 −𝑞) ̇ +𝐾𝑃 (𝑞𝑑𝑒𝑠 −𝑞) +𝑟
(10)
where 𝑟 is the reference component [12, 19]. This ideal linearization approach relies on the exact cancellation of nonlinearities in the system equations of motion. Its practical implementation requires consideration of various sources of uncertainty, including modeling errors, unknown payloads, and computation errors [20]. It also requires online computation of the exact time-varying 𝑀 (𝑞) and 𝑁 (𝑞, 𝑞), ̇ which is computationally expensive, impractical, and in some cases impossible for real-time aerial manipulation tasks. To overcome this, we propose ̂ using an offline computed fixed nominal model (𝑀,̂ 𝑁) in the inverse dynamics controller (Fig. 2): 𝜏 = 𝑀̂ (𝑞) 𝑦 + 𝑁̂ (𝑞, 𝑞) ̇ .
(11)
This approach does not result in perfect cancellation of nonlinearities. Instead, applying (11) to the actual plant (7) results in a perturbed linear system. The difference between the fixed nominal model and the actual varying dynamics is isolated into structured uncertainty blocks. The system dynamics become: 𝑞̈ = 𝑦 + △𝑚 𝑦 + △𝑎 ,
(12)
where △𝑚 , △𝑎 are multiplicative and additive uncertainties:
120
This formulation, shown in Fig. 3, allows us to treat the nonlinear complex aerial manipulator dynamics as a linear plant subjected to bounded perturbations that can be managed using robust control techniques. The robust control input 𝜏 for the resulting system is expressed by: ̇̌ 𝑃 𝑞�̌ +𝑁̂ (𝑞, 𝑞) 𝜏= 𝑀̂ (𝑞) �𝑞̈ 𝑑𝑒𝑠 +𝐾𝐷 𝑞+𝐾 ̇ ,
̂ △𝑚 = 𝑀−1 (𝑞)𝑀(𝑞) − 𝐼,
(13)
△𝑎 = 𝑀−1 (𝑞) 𝑁̃ (𝑞, 𝑞) ̇ ,
(14)
(16)
where 𝑞̌ and 𝑞̇̌ are the tracking error and differential error terms. The robust stability of this control law against the uncertainties △𝑚 and △𝑎 is assessed through the 𝜇analysis and Small Gain procedure, as detailed in the next section.
3.
where 𝑦 represents a virtual input vector:
(15)
Control System Design
The dynamic model derived in the previous section, particularly the feedback-linearized form in Eq. (12), isolates the system’s complex, time-varying, and payload-dependent dynamics into a structured uncertainty term: Δ = 𝑑𝑖𝑎𝑔(△𝑚 , △𝑎 ). For a global controller, the bounds on this uncertainty must encompass all possible manipulator configurations, velocities, and payload properties. Such a global bound is necessarily large, forcing a monolithic robust controller to be highly conservative. This severely degrades tracking performance and agility. In this paper, a hybrid gain-scheduling framework is introduced to minimize the uncertainty bounds. The main strategy is to partition the system’s operating space into a finite K-number of smaller, dynamicallysimilar regions. If a dedicated controller is designed for each region, it is only necessary to guarantee stability against much smaller local uncertainty bounds, enabling high-performance control in every region. 3.1.
Dynamic-Based State Space Clustering
To avoid the conservatism of a single robust controller, a comprehensive sampling space was defined that includes not only the manipulator kinematics, but also the full range of expected payload masses and object dimensions. The dataset, generated by sampling the manipulator’s kinematics and potential payload characteristics, is then subjected to unsupervised learning. The K-Means algorithm is employed to minimize the within-cluster variance of the dynamic parameters. An important design parameter is the cluster cardinality K. Rather than arbitrarily selecting this value,
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
controller 𝐾 (𝑠) = 𝑑𝑖𝑎𝑔 {𝐾1 (𝑠) , … , 𝐾6 (𝑠)} , where each 𝐾𝑘 (𝑠) = 𝐾𝑝𝑘 + 𝐾𝑑𝑘 𝑠/(T𝑠 + 1), and where T is the time constant for the derivative filter, interconnection matrix 𝐺 (𝑠) is a 12 × 12 block matrix given by:
𝐺 (𝑠) = �
𝐾(𝑠) � 𝑠 +𝐾(𝑠) 1 𝑑𝑖𝑎𝑔 � 2 � 𝑠 +𝐾(𝑠)
−𝑑𝑖𝑎𝑔 � 2
𝐾(𝑠) � 𝑠 +𝐾(𝑠) �. 1 𝑑𝑖𝑎𝑔 � 2 � 𝑠 +𝐾(𝑠)
−𝑑𝑖𝑎𝑔 � 2
(17) According to the Small Gain Theorem, robust stability for cluster 𝑘 is guaranteed if the supremum of the structured singular value 𝜇Δ (𝑃(𝑗𝑤)) is less than 1 for all frequencies 𝑤 [13]: sup 𝜇Δ𝑘 (𝑃(𝐶𝑘 , 𝑗𝑤)) < 1
(18)
𝑤
Figure 4. Elbow method for optimal K
where 𝐶𝑘 is the local controller for cluster 𝑘. This condition can be expressed as ensuring the robust stability margin, 𝛽𝑘 = 1/𝜇Δ𝑘 , is greater than 1. Calculating the optimal feedback 𝐾𝑝𝑘 and 𝐾𝑑𝑘 that simultaneously satisfy the Small Gain condition 𝛽𝑘 > 1 and minimize tracking error is a non-convex optimization problem. To navigate this complex search space, a GA is employed to evolve a population of candidate controllers for each cluster. The fitness of each candidate solution is evaluated using a composite cost function that penalizes instability while rewarding precision. The evaluation criteria 𝐽 balances 3 objectives: 2
𝐽 = − �𝛼1 𝑒 + 𝛼2 ‖𝜏‖ + 𝛼3 max(0, 𝜇Δ𝑘 (𝑃(𝑗𝑤)) − 1)� , Figure 5. LFT model of the system the distortion curve is analyzed, plotting the withincluster variance against an increasing number of centroids. By identifying the point of rate decrease — the geometric “elbow” of the curve — the optimal K number is determined, balancing model fidelity against controller complexity. This granularity ensures that within any specific cluster 𝑘, the deviation between the local nominal model and the actual dynamics remains within the predictable bounds required for robust stability synthesis. The distortion curve for the system is shown in Fig. 4, which represents the Sum of Squared Errors (SSE) against the number of clusters. 3.2.
Robust Controller Synthesis
For each of the K regions, a dedicated robust controller is designed. The system, linearized by approximate inverse dynamics, is analyzed using the LFT framework, as shown in Fig. 5. This structure isolates the nominal plant 𝑃 (𝑠) from the local, bounded uncertainty Δ𝑘 = 𝑑𝑖𝑎𝑔(△𝑘𝑚 , △𝑘𝑎 ). The interconnection matrix 𝐺 (𝑠) maps the uncertainty blocks (𝑦 from △𝑚 , 𝑤 from △𝑎 ) to their respective inputs (𝑦, 𝑧) through the nominal closed-loop system. For the 6-DOF system under consideration, with a diagonal plant 1 1 𝑃𝑝𝑙𝑎𝑛𝑡 (𝑠) = 𝑑𝑖𝑎𝑔{ 2 , … , 2 } and a diagonal PD 𝑠
𝑠
where 𝛼1 , 𝛼2 , 𝛼3 are tuning weights. The first term minimizes the Root Mean Square Error (RMSE) of the step response, and the second term penalizes excessive control effort. The third term returns zero if the stability margin 𝛽𝑘 > 1, but applies a heavy penalty if the robust stability constraint is violated. Consequently, the optimization process selects a set of controller parameters that maximize precision without compromising the system’s stability margins. As can be observed from the relationship between the linearized dynamics (9) and the control law (10), the selection of positive-definite gain matrices ensures the asymptotic convergence of the tracking error to zero. This structure guarantees that the system accurately follows the desired trajectory, even when subjected to the bounded model uncertainties defined in (12). 3.3.
Gain-Scheduling Specification
The gain-scheduling framework functions as an adaptive mechanism that bridges real-time system dynamics and a library of robust controllers. During operation, the system calculates the Euclidean distance between the current dynamic state and the pre-computed operating points, or cluster centroids, which represent the partitioned workspace. By identifying the most relevant clusters in this space, the system retrieves the corresponding control gains that were specifically optimized during the offline tuning phase. 121
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 1. UAV and manipulator characteristics Parameter UAV mass UAV arm length UAV moment of inertia (X,Y axes) UAV moment of inertia (Z-axis) Length of link 1 Length of link 2 Mass of link 1 Mass of link 2
Symbol 𝑚𝑈𝐴𝑉 𝐿𝑈𝐴𝑉 𝐼𝑋 , 𝐼𝑌
Value 4.5 0.5 0.075
Unit 𝑘𝑔 𝑚 𝑘𝑔𝑚2
𝐼𝑍
0.12
𝑘𝑔𝑚2
𝐿1 𝐿2 𝑚𝐿1 𝑚𝐿2
0.2 0.3 0.3 0.45
𝑚 𝑚 𝑘𝑔 𝑘𝑔
4. Simulation and Results A series of simulations were conducted to validate the performance and robustness of the proposed control framework. The primary goal of the simulations was to demonstrate that the offline-tuned controller could robustly manage the dynamic variations introduced by rapid manipulator motion and the need to grasp a variety of payloads. 4.1.
Simulation Setup
Simulations were conducted in the MATLAB software to validate the performance and robustness of the UAV manipulator system. The configuration space was sampled by focusing on the dominant variables: the manipulator joint angles, the payload mass, and the payload dimensions. The grid of manipulator positions was combined with a set of payload masses, along with three distinct object dimension sets (small, medium, and big). For each combination, a set of five fast velocity vectors (up to 𝜋/2 rad/s) were applied to capture the full dynamics, including Coriolis and Gravitational effects. This process generated a comprehensive dataset of ∼30,000 valid dynamic samples 𝑑𝑖 = [𝑀, 𝑁]. The KMeans algorithm, using 𝐾 = 300 (determined from the Elbow Method, Fig. 4), was used to partition this data, yielding 300 local nominal models (𝑀̂ 𝑘 , 𝑁̂ 𝑘 ). The 99th percentile 𝐻∞ of the local uncertainty Δ𝑘 = 𝑑𝑖𝑎𝑔(△𝑘𝑚 , △𝑘𝑎 ) was calculated for each cluster. This resulted in significantly small bounds for both multiplicative (𝛾𝑘𝑚 ) and additive (𝛾𝑘𝑎 ) uncertainties (Fig. 6), which were then used as constraints for the controller synthesis. For each of the K = 300 clusters, a local diagonal PD controller was tuned, using the GA as described in the previous section. The dynamic performance of the proposed system depends on the behavior of the entire setup, including the masses and dimensions of the UAV and manipulator. The physical parameters used in the simulations are summarized in Table 1. 4.2.
Robust Stability Validation
The first validation test was conducted to confirm that the GA successfully synthesized controllers that meet the robust stability conditions. The robust stability margin, 𝛽𝑘 = 1/𝜇Δ𝑘 , was calculated using the final tuned controller in each of the 300 clusters. As 122
Figure 6. Uncertainty norms in clusters
Figure 7. Robust stability margin (𝛽) on a logarithmic scale shown in Fig. 7, the stability margin for all 300 clusters is greater than the robust stability boundary of 𝛽 = 1. This result confirmed that the offline GA tuning, constrained by the local bounds, successfully generated a set of controllers that are guaranteed to be robustly stable within their respective operating regions. The second simulation was implemented to validate the adaptability and precision of the control framework when confronted with the simultaneous changes in both manipulator configuration and payload properties. This simulation was executed over
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 2. Payload characteristics Object Small Medium Big
Time 0-10 sec 10-20 sec 20-30 sec
Mass (kg) 0.1 0.5 1
Dimensions (m) 0.1, 0.02, 0.02 0.4, 0.05, 0.05 0.6, 0.08, 0.08
a 30-second span, forcing the system through critical operational states. The UAV’s translational and attitude states were commanded to hold a still position of 𝑍 = 1𝑚 and 𝜙, 𝜃, 𝜓 = 0, while manipulator joints were commanded to perform a sinusoidal tracking motion from -𝜋/2 to 𝜋/2 radians, actively switching the payload every 10 seconds to stress the controller’s adaptability. The payload characteristics and corresponding grasping times are given below in Table 2. The resulting plots are depicted in Fig. 8. The plots show that the actual trajectory remains tightly coupled with the desired trajectory for all controlled states. Only minimal tracking errors were observed during the manipulation of the biggest payload configuration, and these deviations were quantitatively negligible. Gain switching was logged along the simulations to analyze the control stability at transition points. To mitigate the potential for abrupt control actions, the gain interpolation technique was used to ensure a smooth transition between scheduled gains. Instead of selecting a single best controller, this method blends the control laws from multiple adjacent clusters. Using the k-nearest neighbors algorithm, the system identifies the four closest cluster centroids to the current dynamic state. This provides a set of indices for the four most relevant local controllers. For each of the neighboring clusters, the weight component 𝑤𝑖 is calculated based on its Euclidean distance 𝐷𝑖 from the current state. The weight is computed as the inverse of the distance, with a small 𝜀 added to avoid division by zero 𝑤𝑖 = 1/𝐷𝑖 + 𝜀. (19)
Figure 8. Trajectory tracking of the system
The weights are then normalized to sum to one, esuring a stable weighted average: 𝑤𝑖∗ =
𝑤𝑖 , 𝑘 ∑𝑗=1 𝑤𝑗
(20) Figure 9. Active cluster index over time
𝑘
𝐾𝑏𝑙𝑒𝑛𝑑 = � 𝑤𝑖∗ ⋅ 𝐾𝑖 .
(21)
𝑖=1
The resulting gain-switch plot is presented in Fig. 9. As a result, a total of 176 gain switches occurred, with 48 unique cluster gains used, indicating continuous adaptation. The integrated gain interpolation successfully mitigated the discontinuities of hard switching. This is quantitatively supported by the RMSE plot in Fig. 10, which confirms that the smooth switching mechanism helped maintain a consistently low tracking error across the entire mission, even when controlling the system under the biggest payload (Obj. 3).
5.
Conclusion
This study successfully addressed the complex challenge of robust trajectory tracking for an aerial manipulator during fast maneuvers with varying payloads. A novel gain-scheduling framework was introduced that shifted the computational burden to preflight phase, overcoming the real-time slow computations and poor adaptability of traditional model-based controllers. The core contribution is a method that first linearized the system using approximate inverse dynamics with fixed nominal models, isolating the complex, 123
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
References [1] J. Meng et al., “On Aerial Robots With Grasping and Perching Capabilities: A Comprehensive Review”, Frontiers in Robotics and AI, vol. 8, 2022, pp. 1–24. DOI: 10.3389/frobt.2021.739173. [2] J. Morais, D. Cardoso, and G.V. Raffo, “Robust Optimal Nonlinear Control Strategies for an Aerial Manipulator”, Congresso Brasileiro de Automática, vol. 2, no. 1, 2020. DOI: 10.48011/asba.v2i1.1616. [3] N. Mimmo et al., “Robust Motion Control of Aerial Manipulators”, Annual Reviews in Control, vol. 49, 2020, pp. 230-238. DOI: 10.1016/j.arcontrol.202 0.04.006. Figure 10. RMSE by state
time-varying dynamics as structured uncertainties. The entire configuration space, including payload variations, was then partitioned into 300 dynamically similar clusters using K-Means. For each local region, a dedicated robust PD controller was synthesized using a multi-objective GA, which optimized tracking performance while mathematically guaranteeing robust stability via the Small Gain Theorem. Simulation results validated this approach, confirming that all 300 local controllers were robustly stable within their operating regions. The system demonstrated high-precision trajectory tracking for all states while performing an aggressive sinusoidal motion and payload switching from 0 to 1 kg. Furthermore, a smooth gain interpolation technique based on the K-nearest neighbors algorithm successfully mitigated abrupt control motions, ensuring continuous stability and consistently low tracking error across all operational scenarios. This work confirms that an offline-tuned, gain-scheduled control system can achieve the high performance, robustness, and adaptability required for fast aerial manipulation tasks. Future work will focus on the experimental validation of the proposed hybrid gainscheduling framework on a physical UAV manipulator platform to evaluate performance under real-world constraints such as sensor noise and actuator latency. Funding The research was supported by the Higher Education and Science Committee of MESCS RA (Research project N∘ 10-4/24AA-2B048).
AUTHORS Tariel Simonyan∗ – Control Systems, National Polytechnic University of Armenia, Yerevan, 0099, Armenia, e-mail: tariel.simonyan@polytechnic.am. Oleg Gasparyan – Control Systems, National Polytechnic University of Armenia, Yerevan, e-mail: ogasparyan@polytechnic.am. ∗
124
Corresponding author
[4] R. Naldi et al., “Robust Control of an Aerial Manipulator Interacting with the Environment”, IFACPapersOnLine, vol. 51, no. 13, 2018, pp. 537–542. DOI: 10.1016/j.ifacol.2018.07.335. [5] C. Carvajal et al., “Multitask Control of Aerial Manipulator Robots with Dynamic Compensation Based on Numerical Methods”, Robotics and Autonomous Systems, vol. 173, 2024. DOI: 10.10 16/j.robot.2023.104614. [6] P. Kremer, J. Sanchez-Lopez, and H. Voos, “A Hybrid Modelling Approach for Aerial Manipulators”, Journal of Intelligent & Robotic Systems, vol. 105, no. 74, 2022. DOI: doi.org/10.1007/s1084 6-022-01640-1. [7] Y. Chen et al., “Robust Control for Unmanned Aerial Manipulator Under Disturbances,” IEEE Access, vol. 8, 2020, pp. 129869–129877. DOI: 10.1109/ACCESS.2020.3008971. [8] I. Al-Darraji et al., “Adaptive Robust Controller Design-Based RBF Neural Network for Aerial Robot Arm Model,” Electronics, vol. 10, 2021, 831. DOI: 10.3390/electronics10070831. [9] Hai Li et al., “Adaptive Neural Network Backstepping Control Method for Aerial Manipulator Based on Variable Inertia Parameter Modeling,” arXiv:2212.04250, 2022. DOI: 10.48550/arXiv. 2212.04250. [10] C. Park, A. Ramirez-Serrano, and M. Bisheban, “Estimation of Time-Varying Inertia of Aerial Manipulators Performing Manipulation of Unknown Objects,” Proceedings of the 10th International Conference of Control Systems, and Robotics (CDSR 23), 2023, 209. DOI: 10.11159/ cdsr23.209. [11] I. Imran, K. Wood, and A. Montazeri, “Adaptive Control of Unmanned Aerial Vehicles with Varying Payload and Full Parametric Uncertainties,” Electronics, vol. 13, 2024, 347. DOI: 10.3390/ele ctronics13020347. [12] B. Siciliano, Robotics: Modelling, Planning and Control, Springer, 2010. [13] K. Zhou and J. Doyle, Essentials of Robust Control, Prentice-Hall, 1998.
Journal of Automation, Mobile Robotics and Intelligent Systems
[14] E. Umargono, J. Suseno, and S.K. Gunawan, “KMeans Clustering Optimization Using the Elbow Method and Early Centroid Determination Based on Mean and Median Formula,” Proceedings of the 2nd International Seminar on Science and Technology (ISSTEC 2019), pp. 121–129. DOI: 10. 2991/assehr.k.201010.019. [15] Y. Ma, “Optimization Of Basic PID Control Algorithm Based on Genetic Algorithm and Matlab”, Proceedings of the 3rd International Conference on Computing Innovation and Applied Physics, vol. 30, 2024, pp. 178–186. DOI: 10.54254/27 53-8818/30/20241103. [16] O. Gasparyan et al, “Robustness Analysis of UAVs’ Control Systems in Case of Motors’ Partial Efficiency Degradation”, Bulletin of High Technology, vol. 25, no. 1, 2023, pp. 67–80.
VOLUME 20,
N∘ 3
2026
[17] Gasparyan O., Linear and Nonlinear Multivariable Feedback Control: A Classical Approach, John Wiley & Sons Ltd, UK, 2008, p. 374. [18] D. Li and W. Hongtao, “Dynamical Modelling and Robust Control for an Unmanned Aerial Robot Using Hexarotor with 2-DOF Manipulator,” International Journal of Aerospace Engineering, 2019, 5483073. DOI: 10.1155/2019/5483073. [19] S. Friedrich and M. Buss, “Parameterizing Robust Manipulator Controllers Under Approximate Inverse Dynamics: A Double-Youla Approach”, International Journal of Robust Nonlinear Control, vol. 29, 2019, pp. 5137–5163. DOI: 10.1002/rnc.4671. [20] Spong M., Hutchinson S., and Vidyasagar M., Robot Modeling and Control, John Wiley & Sons, 2006, p. 407.
125
VOLUME 20, N° 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
Design and Development of Vacuum-Actuated Soft Gripper for Pick and Place Different Irregular Shaped Objects Submitted: 30th September 2025; accepted 2nd February 2026
S.Senthil Raja, R.Gangadevi, M. Dharmaraj DOI: 10.14313/jamris-2026-044 Abstract: This article presents the design, fabrication and testing of a novel vacuum-based soft robotic gripper. This work included the fabrication of a cylindrical gripper via a 3D-printed mold, into which a combination of liquid silicone elastomer and curing agent was injected. The deformation behavior of the soft gripper was analyzed for different pressures using the ANSYS software. The ANSYS results revealed that the pressure, material and size of the cavity inside the gripper are the most important factor affecting the deformation of the soft gripper. Finally, the grasping power and holding time of the fabricated gripper was tested by grasping different objects, including fruits, mobile phones, plastic bottles and stainless-steel utensils. The validation findings demonstrate the gripping and holding capabilities of the proposed vacuum gripper, suggesting its suitability for various automation and industrial applications. Keywords: Soft Gripper, vacuum actuator, grasping power, liquid silicone, finite element model
1. Introduction The end effector is a crucial component of a robotic system, enabling the execution of various operations and the handling of objects throughout the process. Nevertheless, the manipulation of fragile and flexible objects is one of the most significant challenges in robotics. Following advances in soft materials, soft grippers have attracted more interest among academics recently. Silicone rubber, elastomers, and thermoplastic polyurethanes (TPU) are predominantly utilized in soft gripper applications. The production of soft grippers utilizing these materials in additive manufacturing provides enhanced operational reliability and savings. Various forms of soft grippers have been designed and developed by various researchers that have used mechanical linkages, pneumatics, or magnets. When the mechanical linkages in the soft gripper are used, the force that is applied by the gripper is increased, which assists in the handling of heavy objects with more precision. In this way, the stability of the grasp on the object is improved, and the repeatability and reliability of the grip are improved as well. While mechanically-operated soft grippers offer numerous advantages, they also present several disadvantages, including challenges in control, increased wear and tear, limited compliance, and higher initial costs. 126
To address these limitations, researchers have designed and developed a novel type of pneumatically operated soft grippers and successfully tested them in various real-time applications. Pneumatic soft grippers are a type of gripper that utilizes high-pressure air to grasp objects of both regular and irregular shapes. These grippers are very flexible, lightweight, and cost-effective compared to other types of grippers. As a result of these advantageous aspects, the use of soft grippers has increasingly been considered for agricultural [1, 2], food processing [3, 4], and medical [5, 6] applications in the last few years. Various types of pneumatically-powered soft grippers are used in different fields. Fig. 1 illustrates many kinds of grippers and their applications. An increasing number of studies have focused on the soft vacuum gripper due to its many advantages over other types of grippers, including its simplicity, low cost, and ease of use. Li et al. [7] designed a vacuum-operated soft gripper utilizing an origami magic ball and an elastic thin membrane. The experimental results demonstrated that the soft gripper generates substantial grasping force via vacuum, making it promising for various applications in the handling of heavy objects. Tawk et al. [8] created a 3D-printed, multifunctional soft gripper that utilizes a suction cup to grasp various items of differing weights, sizes, and shapes. That study’s findings indicated that a soft gripper equipped with a single suction cup generates a gripping force of about 30.35N within 94ms. By contrast, the multi-functional soft gripper produces a gripping force of 31.31N and a maximum payload-to-weight ratio of 7.06. Maggi et al. [9] introduced an innovative gripper called POLYPUS that is designed to provide exceptional load-lifting capabilities for diverse products such as cardboard, glass, sheet metal, and plastic via the integration of under-actuation and vacuum-gripping techniques. The simulation results revealed that the POLYPUS gripper provided minimal suction force for handling various items. Jang et al. [10] introduced an innovative concept, inspired by air fuses, for creating a soft valve for a soft vacuum gripper, and examined its practicality via experimental methods. The study also developed analytical models to forecast pressure conservation and operational duration for the proposed valve, and to validate the experimental findings. The suggested soft vacuum gripper with a soft valve was successfully constructed and evaluated under various operating circumstances.
Open Access. © 2026 S.Senthil Raja et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP. This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 License.
Journal of Automation, Mobile Robotics and Intelligent Systems
Figure 1. Details of different types of soft gripper Tawk et al. [11] used 3D printing to develop innovative soft vacuum actuators inspired by the sporangium of ferns. The actuators were evaluated for several applications, including locomotion and hopping robots, and the system demonstrated versatility via multiple degrees of freedom and variable length, using a set of 3D-printed vacuum hinges. Tahir et al. [12] developed the pneumatically powered silicone rubber gripper, termed PASCAV, and evaluated its grabbing efficacy on various items under diverse working circumstances. Their experimental research findings indicate that this gripper is more effective for flat, curved, and irregularly-shaped items, providing more grabbing power. The experimental results have been confirmed by finite element analysis (FEA) findings, demonstrating strong agreement. Numerous researchers have fabricated and studied different soft robotic grippers and tested them in different applications. [13–16]. Over the past several years, the field of medicine has witnessed consistent and rapid technological advancements in this regard, with one revolutionary step forward being when the Internet of Things (IoT) idea was first implemented in clinical settings [17]. The aforementioned literature clearly shows the potential of vacuum-driven soft grippers. Despite much study conducted on vacuum-operated soft grippers, there is significant potential to evaluate their effectiveness across many applications. This study thus undertook the design and fabrication of an innovative gripper made from soft material that operates using vacuum power. Its grabbing capability was evaluated by handling items of various shapes and sizes. Finite element analysis was conducted to confirm the experimental findings and examine the pressure distribution in the gripper.
2. Materials and Methods
In order to analyze the grasping capability of the soft gripper, a compact vacuum gripper was designed that was 43 mm in height and 65 mm in
VOLUME 20,
N° 3
2026
diameter, and capable of handling a maximum payload of 500 grams. The design of the proposed cylindrical bell-shaped vacuum soft robotic gripper is illustrated in Fig. 2. The required mold for the proposed cylindrical bell-shaped gripper was designed using SolidWorks software and fabricated using a 3D printer. The proposed gripper has a total height of 34 mm, consisting of a 3-mm-thick base flange and a 31-mm-tall vertical cylinder bell. The outside diameter of the base flange is maintained at 65mm to provide enough surface area for grasping the item. Likewise, the diameter of the top cylinder is presumed to be 42mm so that it can attach the gripper to the robotic manipulator effortlessly. To provide appropriate rigidity, the gripper was designed with an interior cavity diameter of 36mm and a wall thickness of 3mm. To create a vacuum, one 4-mm cylindrical hole is added to the top surface of the gripper and internal cavity. The vacuum pump sucks the air from the internal cavity through the 4-mm-diameter hole. To fabricate the vacuum soft gripper’s mold was created using a Bambu lab P1S FDM 3D printer and WOL3D PLA filament. Fig. 3 is a photographic representation of the 3D-printed mold. Liquid silicone was poured into the mold and allowed to be cured for 8 hours. Because of its low viscosity, greater flexibility and good elongation, SILOCZEST advanced Liquid Silicone Rubber was used in this study. The hardened soft vacuum gripper was then taken out from the mold for further testing. Fig. 4 shows the resulting soft vacuum gripper. A schematic diagram of the step-by-step procedure adopted in this research is presented in Fig. 5.
Figure 2. Measurements of this study’s soft gripper
Mold – Top por on
Mold – bo om por on
Figure 3. Pictorial views of the top and bottom portions of the fabricated 3D-printed mold 127
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
3. Results and Discussion Using the ANSYS software, analysis of different pressure variations in the deformation was performed. In order to obtain the most accurate results, the displacement in the outer cylinder surface was assumed to be zero. The model was created in the SolidWorks software, exported in .STEP format, and imported into the ANSYS software for further analysis. In this study, the pressure varied from 10 Pa to 1000 Pa, with increments of 10 Pa applied to the inside surface of the air chamber. Variations in the deformation with respect to pressure are depicted in Fig. 6, which indicates that the gripper surface undergoes nonlinear deflection in relation to pressure. During the first phase (up to 300 Pa), there is a significant rate of deformation with little variations in pressure. This signifies that the material has a soft nature within this pressure range. With an increase in pressure over 300 Pa, the material’s stiffening effect results in a reduced rate of deformation. To evaluate the variation in deflection of the proposed soft vacuum gripper, a simulation was conducted using ANSYS. The series of pictures shown in Fig. 7 illustrates the differences in deflection of the finite element model at different pressures, including 250 Pa, 500 Pa, 750 Pa, and 1000 Pa. The findings indicated that the largest deformation occurs at the middle of the gripper and progressively diminishes towards the edges. This is because the center section of the gripper exhibits significant flexibility relative to the areas adjacent to the vertical wall and the direct pressures exerted at this location. As a result, the maximal deformation is 0.71193 mm at 250 Pa, but it increases substantially to 9.170 mm at 1,000 Pa. When the pressure increases from 250 Pa to 500 Pa, the deformation also increases non linearly from 0.71193 mm to 6.818 mm. A small change in deformation was observed when the pressure increases from 750 Pa to 1000 Pa because of increased material stiffness. The trends of these results are in agreement with those obtained by Tahir et al. [12]
4. Applications of proposed soft gripper
128
To test the grasping capabilities of the proposed soft gripper, several experiments were conducted with differently shaped and sized objects. Fig. 8 illustrates the grasping capabilities of the designed vacuum gripper. Each image (from a to l) in these figures shows the proposed gripper’s lifting, holding and grasping abilities for the different items, including a water bottle, fruits, plastic plates and stainless-steel utensils. These figures provide the strong evidence of the gasping and holding capacity of the proposed vacuum soft gripper. In subfigure (a), the vacuum gripper securely holds a plastic bottle (which has a curved surface and elliptical cross-section) for an adequate duration without causing damage to the item. A rectangular plastic bottle weighing about 100 gms was gripped in subfigure (b); the contact surface area of this item is much greater than that of the object seen in subfigure (a), and therefore the retention duration for this item is slightly higher.
Figure 4. Fabricated vacuum-actuated soft gripper
Figure 5. Step-by-step procedure adopted to fabricate the soft gripper
Figure 6. Variations of deflection of the soft gripper surface with respect to pressure Subfigure (c) depicts the grasping of a textured item, highlighting the gripper’s ability to maintain airtightness despite surface irregularities. Subfigure (d) illustrates the manipulation of a tiny plastic cylindrical container with a flat top surface weighing about 10 gms. Subfigures (e) and (g) illustrate the gasping capabilities of the vacuum gripper for various fruits weighing 200 gms, demonstrating that the gripper securely grasps and retains the fruit using vacuum pressure without causing damage. Similarly, subfigure (f) demonstrates the gripper’s potential to grasp large, irregularly shaped water bottles, thus confirming its holding stability. Subfigure (j) illustrates the grasping of a mobile phone with flat, smooth surfaces, clearly exhibiting the gripper’s flex-
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 7. The variations of deflection of the finite element model at 250Pa, 500Pa, 750 Pa and 1000 Pa ibility even when total sealing is difficult. From these figures, it can be clearly understood that the vacuum gripper designed for this study can be used to grasp not only smaller sized objects, bigger-sized, irregularly-shaped objects as well. Hence, this kind of soft robotic gripper has a promising future in many automated industrial applications, where it could grasp objects without any physical damage.
3.
4.
5. Conclusions and Future Works
The main objective of this study was to design and develop a novel cylindrically-shaped vacuum actuated soft gripper using molding and casting fabrication methods. The fabricated gripper was tested in a realtime environment, and the major conclusions can be summarized as follows. 1. The cylindrically-shaped vacuum-actuated soft gripper was successfully designed using SolidWorks software and fabricated by pouring the liquid silicone and curing agent mixture into the mold cavity. 2. The designed gripper was analyzed using the ANSYS software, and the results confirmed its nonlinear characteristics under different pressures. In this study, the maximum deflections of about 0.71193 mm, 6.818mm and 9.170 mm were obtained for 250 Pa, 500 Pa and 1000 Pa pressure levels, respectively.
5.
xperimental tests were carried out to validate E the gripper’s grasping and holding abilities and to confirm that it can securely grasp the different irregularly-shaped objects without any physical damage. Due to its compact design and enhanced material flexibility, the proposed gripper is suitable for handling fragile and delicate items, including glass, thin plastic components, lightweight consumer goods, and food and agricultural products, as well as electronic components such as PCBs and medical and surgical instruments. However, as the grasping and holding ability of the proposed vacuum actuated gripper successfully demonstrated, integrating the sensors is one of the most challenging tasks, and identification of another suitable method will be needed. Future research may focus on improving the performance of this gripper using automated silicone mixing systems, vacuum degassing, multi material printing systems, etc.
ACKNOWLEDGEMENTS
This work was supported SRM Institute of Science and Technology, Kattankulathur, Tamil Nadu, India. We would like to thank the Selective Research Initiative SERI-”2023”, SRMIST for their financial sup-port 129
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 8. Applications of the proposed vacuum soft gripper: (a) Grasping plastic bottle; (b) Grasping rectangular plastic box; (c) Grasping jewelry box; (d) Grasping medicine bottle; (e) Grasping mosambi fruit; (f) Graspin empty water bottle; (g) Grasping orange; (h) Grasping book; (i) Grasping plastic bottle cap; (j) Grasping mobile phone; (k) Grasping stainless steel cup; (l) Grasping stainless steel tiffin box AUTHORS S. Senthil Raja* – Department of Mechatronics Engineering, SRM Institute of Science and Technology, Kattankulathur, Tamil Nadu, India – 603203, Email: senthils9@srmist.edu.in R. Gangadevi – Department of Mechatronics Engineering, SRM Institute of Science and Technology, Kattankulathur, Tamil Nadu, India – 603203, Email: gangader@srmist.edu.in M. Dharmaraj – Department of Mechancial Engineering, Kongu Engineering college, Perundurai, Tamil Nadu, India – 638060, Email: mdmech@kongu. ac.in *Corresponding author
References
130
[1] S.M.G. Vidwath, P. Rohith, R. Dikshithaa, N. Nrusimha Suraj, R.G. Chittawadigi, and M. Sambandham, “Soft Robotic Gripper for Agricultural Harvesting,” In Kumar, R., Chauhan, V.S., Talha,
M., Pathak, H. (eds), Machines, Mechanism and Robotics: Lecture Notes in Mechanical Engineering, Springer Singapore, 2022, pp. 1347–1353. https://doi.org/10.1007/978-981-16-05505_128 [2] A. Kultongkham, S. Kumnon, T. Thintawornkul, and T. Chanthasopeephan, “The Design of a Force Feedback Soft Gripper for Tomato Harvesting,” Journal of Agricultural Engineering, vol. 52, no. 1, 2021. https://doi.org/10.1016/j. robot.2020.103427 [3] Z. Wang, K. Or, and S. Hirai, “A Dual-Mode Soft Gripper for Food Packaging,” Robotics and Autonomous Systems, vol. 125, 2020, 103427. https://doi.org/10.1016/j.robot.2020.103427 [4] J.H. Low et al., “Sensorized Reconfigurable Soft Robotic Gripper System For Automated Food Handling,” IEEE/ASME Transactions On Mechatronics, vol. 27, no. 5, 2021, pp, 3232–3243. DOI: 10.1109/TMECH.2021.3110277
Journal of Automation, Mobile Robotics and Intelligent Systems
[5] I. Zhou, L. Ren, Y. Chen, S. Niu, Z. Han, and L. Ren, “Bio‐Inspired Soft Grippers Based on Impactive Gripping,” Advanced Science, vol. 8, no. 9, 2021, 2002017. https://doi.org/10.1002/ advs.202002017 [6] M.Q. Shahid, F. Amin, A.A. Haris, H. Al Faisal, A. Imran, and M.F. Qureshi, “Design and Development of a Soft Robotic Gripper for Precision Control for Biomedical Applications,” International Conference on Engineering & Computing Technologies (ICECT 24), 2024, pp. 1–6. DOI: 10.1109/ICECT61618.2024.10581172 [7] S. Li, J.J. Stampfli, H.J. Xu, E. Malkin, E.V. Diaz, D. Rus, and R.J. Wood, “A Vacuum-Driven Origami ‘Magic-Ball’ Soft Gripper,” International Conference on Robotics and Automation (ICRA 19), 2019, pp. 7401–7408. DOI: 10.1109/ ICRA.2019.8794068 [8] C. Tawk, A. Gillett, M. in het Panhuis, G.M. Spinks, and G. Alici, “A 3D-Printed Omni-Purpose Soft Gripper”, IEEE Transactions on Robotics, vol. 35, no. 5, 2019, pp. 1268–1275. DOI: 10.1109/ TRO.2019.2924386 [9] M. Maggi, G. Mantriota, and G. Reina, “Introducing POLYPUS: A Novel Adaptive Vacuum Gripper,” Mechanism and Machine Theory, vol. 167, 2022, 104483. https://doi.org/10.1016/j. mechmachtheory.2021.104483 [10] G. Jang, H.G. Shin, and W.K. Chung, “Soft Transistor Valve for Versatile and Fast Soft Vacuum Gripper”, IEEE Robotics and Automation Letters, vol. 9, no. 11, 2024, pp. 9565–9572. DOI: 10.1109/LRA.2024.3458598 [11] C. Tawk, M. in het Panhuis, G.M. Spinks, and G. Alici, “Bioinspired 3D Printable Soft Vacuum
VOLUME 20,
N° 3
2026
Actuators for Locomotion Robots, Grippers and Artificial Muscles,” Soft Robotics, vol. 5, no. 6, 2018, pp. 685–694. https://doi.org/10.1089/ soro.2018.002 [12] A.M. Tahir, M. Zoppi, and G.A. Naselli, “PASCAV Gripper: A Pneumatically Actuated Soft Cubical Vacuum Gripper,” International Conference on Reconfigurable Mechanisms and Robots (ReMAR 18), 2018, pp. 1–6. DOI: 10.1109/ REMAR.2018.8449863 [13] K. Blanco, E. Navas, L. Emmi, and R. Fernandez, “Manufacturing of 3D Printed Soft Grippers: A Review,” IEEE Access, vol. 12, 2024, pp. 30434– 30451. DOI: 10.1109/ACCESS.2024.3369493 [14] J. Cortes and C. Miranda, “Design, Control, and Applications of Granular Jamming Grippers in Soft Robotics,” Robotics, vol. 14, no. 10, 2025, 132. https://doi.org/10.3390/robotics14100132 [15] L.M. Thomas, T.A. Binoy, R. S. Sudheesh, P.P. Lalu, and P.A. Abdul Samad, “Finite Element Analysis of a Bio-inspired Soft Vacuum Actuator,” Proceedings of the International Conference on Systems, Energy & Environment (ICSEE 20), 2020. [16] Y. Wei, Y. Chen, T. Ren, Q. Chen, C. Yan, Y. Yang, and Y. Li, “A Novel, Variable Stiffness Robotic Gripper Based on Integrated Soft Actuating and Particle Jamming,” Soft Robotics, vol. 3, no. 3, 2016, pp. 134–143. https://doi.org/10.1089/ soro.2016.0027 [17] F. Mulita, G.I. Verras, C. N. Anagnostopoulos, and K. Kotis, “A Smarter Health Through the Internet of Surgical Things,” Sensors, vol. 22, no. 12, 2022, 4577. https://doi.org/10.3390/s22124577
131
VOLUME 20, N° 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
Energy-Efficient Control of Collaborative Robots through Intelligent Optimization Algorithms Submitted: 25th November 2025: accepted: 2nd February 2026
Nandkishor Marotrao Sawai, Satpalsing Devising Rajput, Minal Vilas Gade, Dipak D. Bage, Aniruddha S. Rumale DOI: 10.14313/jamris-2026-045
1. Introduction
Abstract This research introduces a comprehensive framework for attaining energy-efficient control of collaborative robots via intelligent optimization techniques, including Particle Swarm Optimization (PSO), Genetic Algorithm (GA), and Reinforcement Learning (RL). This research commences with an intricate kinematic and dynamic modeling of the UR5e collaborative robot that, employs the Denavit-Hartenberg (DH) convention to formulate precise motion relationships and dynamic equations that incorporate actuator power, inertia, and gravitational influences. It creates a complete model of energy use that accounts for included actuators, sensors, control systems, and other parts. It also makes fair guesses about how efficient they would be and how they would work. The issue formulation centers on reducing overall energy consumption during job execution while preserving trajectory precision, motion fluidity, and compliance with dynamic and safety restrictions. A multi-objective optimization strategy is suggested to reconcile energy efficiency, job completion duration, and motion smoothness through the utilization of weighted cost functions. The UR5e robot did pick-and-place and assembly tasks on the ROS-Gazebo and MATLAB/Simulink platforms. These tasks were tested in both real world environments and in simulations. The results showed that PSO reduced energy by the most, at 29%, followed by RL with 25% and GA, with 22%. None of the three methods affected the duration or accuracy of the work. RL made motion smoother by being able to learn and adapt. Sensitivity analysis showed that the size of the population, the constraints on iterations, and the learning rate all had significant effects on how well optimization worked. The main contributions of this study are the creation of a complete energy modeling framework, the development of multi-objective optimization methods, and the testing of the system in real-world robotic tasks. These findings show that intelligent optimization can enhance green manufacturing and human-centric manufacturing initiatives by allowing robots to consume less energy, operate more reliably, and exhibit greater environmental sustainability. Future studies are also suggested the future use of AI-driven optimization, predictive maintenance and digital twin integration, which will allow robots to work together to control energy in real time.
The emergence of Industry 4.0, Industry 5.0, advanced manufacturing, cognitive robotics, intelligent scheduling, and networked autonomous systems has led to significant growth in robotics research. These include energy-efficient mobile and industrial robots; advanced human-robot collaboration (HRC); cognitive and AI-driven robotic planning; motion optimization; fault diagnosis; privacy protection; vision systems; scheduling in robotic manufacturing; and swarm robotics. The purpose of all of these activities is to make future generations of robots and cyber-physical environments more self-sufficient, flexible, sustainable, smart. Current publications concentrate on significant contemporary issues, including energy consumption, inertia, autonomous coordination, communication latency, adaptive control, robust scheduling, multi-robot path planning, and cognitive modeling. Many studies present new algorithms, including metaheuristics, machine learning, reinforcement learning, optimization frameworks, hybrid modeling approaches, and intelligent evolutionary strategies. Some other studies focus on running hardware; checking real-time systems; creating datasets; modeling motion; and actually executing experiments with robotics platforms. The present study investigates the problem of communication costs in IoT devices employed in swarm robotics and finds that adaptive transmission methods, modeling, and optimization can make devices work for a long time and stay connected. The proposed system integrates real-time energy-efficient algorithms, adaptive power management, and duty-cycling methodologies. These solutions reduce energy use by 20% and keep packet loss below 5%. MAR (Mobility Aware Routing) provides better scheduling of sleep and wake times, giving the network around 20% more life [1] and enhancing collaborative beam forming (CB) in mobile robotic networks an L-based robot selection method. The method in [2] for multi-objective optimization (SDNED) minimizes energy loss and makes the network last longer. It performs computations more than 66% faster, which is better than any other algorithm. The study shows that when the forward speed (0.17–0.5 m/s) and payload (1–5 KN) are changed, the drawbar pull goes up from 122.36 N to 720.36 N, which also makes the resistance force go up from 15.26 N to 28.05 N. The speed of forward motion has a greater effect on resistance than the weight of the load.
Keywords: Collaborative robots, Energy-efficient control, Particle swarm optimization, Genetic algorithm, Reinforcement learning, Trajectory optimization
132
Open Access. © 2026 Nandkishor Marotrao Sawai et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements
PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
Journal of Automation, Mobile Robotics and Intelligent Systems
The linear regression model in [3] explains more than 99% of the variance, making it a great tool for predicting how robots will affect mobility and for designing robots that operate better in factories. The authors propose the ECACO and PFACO algorithms, which combine energy-saving heuristics with approaches that are not incompatible with each other for planning paths for several robots. The study simulates rough terrain, frictional effects and environmental barriers. Their energy-aware heuristics and pheromone updates take into account kinetic energy, traction resistance, and how much power the hardware needs. PFACO also helps avoid conflicts between nodes, jobs, obstacles and counterpoints. The findings in [4] indicate that their technique increases the likelihood of finding solutions while consuming less energy in both single-robot and multi-robot contexts, resulting in [4] a model for dual-arm robotics that is capable of planning both the placement and movement of the robots. Using GA to optimize PID gains and PSO for motion and configuration planning, it consumes 18% less energy and takes 16% less time to run than planning using only PID. In [5] Duero Robot and K-Rosette are used in a real-life setting to demonstrate how dual-arm setups use less energy than single-arm setups. The study introduces a Self-Learning Memetic Algorithm (SLMA) for Human-Robot Collaboration Scheduling in Remote Fuzzy Welding Settings. Their model minimizes the total energy consumption (TEC) and time required to produce anything. A self-learning variable neighborhood search (SLVNS), which hybridizes Q-learning and VNS, is also developed in this study to enhance exploitation capability, and a resource adjustment strategy is proposed to further optimize TEC. Additionally, to validate the effectiveness of the proposed SLMA, extensive experimental comparisons with five other optimization algorithms are conducted. The research in [6] presents an innovative methodology for scheduling in the framework of Industry 5.0. It proposes a dual-self-learning co-evolutionary algorithm (DSLCEA) for flexible job-shop scheduling utilizing processing-transportation composite robots (PTCRs). The MILP in [7] is good for elementary problems, while DSLCEA adds capabilities like hybrid initialization, chaotic mapping, and dual self-learning. The findings indicate that this method exhibits superior convergence, a broader spectrum of high-quality solutions, and outperforms the most effective evolutionary algorithms currently available. The authors propose that future investigations should focus on PTCR unit-task combinations and the optimization of investment-based robot deployment. The authors of [8] develop two energy-efficient scheduling models, RJSP-E and RJSP-EM, designed to optimize the equilibrium between productivity and energy consumption. Machines and mobile robots alter their speed using a V-scale architecture. RJSP-E decreases energy use by 15%, but this means less work is done. Meanwhile, RJSP-EM uses 10% less
VOLUME 20,
N° 3
2026
energy without adding much time to the process. Both solutions improve the performance of robotic cells. The research in [9] tackled the human-robot collaborative scheduling problem (HCWSSP) by introducing a PMA that integrates GA investigation with VNS exploitation. The experimental results demonstrate that the PMA outperformed alternative algorithms in minimizing makespan and TEC [9]. The hybrid method described in [10] is a distributed auction-based allocation framework (AOCTA) and an enhanced IBPSO, facilitating collaborative robot operation in dynamic industrial environments. It is a substantial advance over standard centralized allocation systems, using less energy, getting more work done, and making the system more scalable, The research in [11] examines the mental simulation of skilled activities, emphasizing ergonomic design and worker safety in Industry 4.0 environments. Mental simulation methods can determine an individual’s likelihood of injury and facilitate their learning of movement. Theoretical implications suggest a shift toward brain-inspired AI as an alternative to brute-force machine learning, which would promote generalizable cognitive architectures defined by multimodal communication and symbolic-subsymbolic integration [11]. A dynamic role-adaptive robot architecture is proposed in [12], employing reinforcement learning for real-time duty reallocation, integration of sustainability (reducing energy use and waste), and the ability to work with many different types of cobots. This research is mandatory for actual implementation and increase of computational overhead. A novel cross-domain co-attention network (CDCAN) is proposed in [13] as a co-observation model that combines information from the time and frequency domains to help diagnose problems in bearings and gearboxes. This method improves the classification accuracy when load, noise level and speed change. Its strength is that it works well with a variety of machinery and can be used in a variety of fields. Future work might develop CDCAN indications for temperature, current and other physical parameters [13]. A comprehensive assessment of AMR energy optimization, identifying locomotion (50%), computing (33%), and sensing (11%) as the principal energy contributors ,is performed in [14]. This is followed by a discussion about how to manage batteries, PID/ MPC/DRL control methods, hybrid modeling, path planning, scheduling, the effects of friction, the effects of load, and environmental restrictions, in a review that discusses the crucial role of BMS, energy-aware control, multi-robot scheduling, and algorithmic selection. Old approaches are combined with deep video in [15] to make colorized historical images and videos that look good and make sense over time. The study also makes available a range of datasets encompassing diverse periods, locales, and styles of clothing to assist with further study. 133
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Object detection can be improved using a hardware-assisted IoT pipeline that includes MATLAB, STM32F426zi, CC3200, and a wheeled robot [16]. Images are communicated in real time using UART, Wi-Fi, and UDP. This makes them easier to recognize in factories and other places where they are used. The technology described in this study reduces transmission delays and enables it automatically follow a line with infrared light. Leakage of users’ privacy information in location-based services has led to the development of LP-BT, which use ball trees for spatial indexing [17]. The model figures out the highest anonymity entropy and then utilizes neural network prediction to make privacy protection even better. The simulation results indicate that our strategy surpasses traditional K-anonymity techniques. Taguchi design, ANOVA, and ANN are employed in [18] to enhance a missile’s glide trajectory through the optimization of PSO parameters. Choosing the right parameters increases the gliding range by 12.77% and the flight time by 15.71%. The results indicate that the constraint satisfaction was robust in all test cases. A comprehensive analysis of energy-saving techniques is presented in [19], encompassing adaptive control, intelligent motion planning, predictive maintenance, sensor fusion, data-driven optimization, energy-efficient hardware, and model-based control. The writers emphasize the importance of optimizing multiple goals at once, such as energy, quality, and time. It also discusses the future of machine learning, energy storage, human-robot collaboration and regenerative energy systems. As smart manufacturing and Industry 5.0 environments increasingly use collaborative robots (cobots), the need for energy-efficient control solutions that do not sacrifice safety, accuracy or flexibility for people has increased. Robots work with people and require a lot of energy because they always moving, understanding and making decisions. This is especially true when workloads vary and people are doing multiple things at once. Traditional control methods sometimes ignore the complex relationships between robot dynamics, job uncertainty, and energy flow, resulting in inefficiencies and reduced system durability. Recent advances in intelligent optimization algorithms, such as evolutionary computation, swarm intelligence, and reinforcement learning, provide robust methods for building controllers that autonomously adjust parameters, adapt to real-time changes, and reduce power consumption while ensuring optimal performance. The present study examines an integrated framework that uses intelligent optimization to enhance robot energy efficiency, thus promoting more sustainable, responsive and cost-effective human-robot collaboration in modern industrial environments.
energy these robots use requires thinking about how they move and how their parts use energy. System modeling is the first step in generating energy-capacity control plans and more complex optimization methodologies.
Collaborative robots are designed to work securely with people in the same area. Figure out how much
Figure 1. Kinematic models of a UR5e robot using DH parameters
2.1 Kinematic and Dynamic Modeling Researchers can learn about how the robot moves by looking at kinematic and dynamic models. Kinematics is the study of the geometric relationships between a robot’s joints and linkages, but not its power. This lets the figure out exactly where the robot will end up and which way it will be facing. Researchers need to know this to make sure that the robot functions effectively and responds well to changes in the outside world. Adding additional algorithms to these models can make the robot more adaptable and able to deal with a wider range of scenarios. This is why the Denavit-Hartenberg (DH) convention is so popular. It provides a consistent method for indicating offsets, link lengths, and joint angles. For instance, UR5e joins robots to show how the six joints are arranged, refer Figure 1. To start and finish activities, it needs to use parameters. The dynamic modeling improves the structure by taking into account the combined acceleration, gravitational effects, and inertial characteristics of the connections, as well as the qualities of the materials used. It is crucial for actuators to manage how much electricity they consume to increase performance and cut down on energy waste, and these equations of motion make it easy to figure out how much mechanical effort is needed to run a machine. 2.2 Energy Consumption Modeling There are many ways that robots consume energy. Motors and actuators are the key energy users that turn electrical power into mechanical movement. These pieces waste energy because they aren’t particularly efficient. To make them operate better, they need to be modeled in great detail. Even if a gripper is built in with little power, it still needs power to work and to hold onto items. Over time, this power use could build up. Sensors,
2. System Modeling for Energy-Efficient Collaborative Robots
134
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Table 1. Assumptions in energy modeling of collaborative robots Sr. No. Parameter 1
Supply Voltage
4
Sensor Power Consumption
2 3 5 6 7 8 9
Motor Efficiency
Control System Efficiency
Symbol / Unit Vs (V)
ηm (%)
24 V
85%
Standard voltage for robot actuators and sensors
Ps (W)
25 W
Average power required by vision and proximity sensors
Pi (W)
5W
ηc (%)
Computing Load Power
Pc (W)
Duty Cycle
Dc (%)
Idle Power Consumption Ambient Temperature Energy Recovery Efficiency
Assumed Value Description
Ta (°C) ηr (%)
Figure 2. Energy consumption breakdown of a collaborative robot
Figure 3. Energy consumption comparisons of different robotic operations controllers, and cooling systems are all examples of assistive systems that utilize a small but important quantity of energy, refer Figure 2. A complete energy model combines these parts to give a complete picture of how much electricity is used during operation. Tests on devices like [10] indicate that both electrical and mechanical components are needed to figure out how well energy-efficient controls operate and how they can be improved.
90%
20 W 70%
25 °C 10%
Efficiency of DC or servo motors driving actuators Power conversion and control losses considered
Energy used for onboard data processing and control Baseline power consumption when robot is in standby mode Active operation percentage during task cycle Normal laboratory working conditions
Portion of kinetic energy recovered during braking or deceleration
2.3 Assumptions and Simplifications System modeling for energy-efficient collaborative machines uses some assumptions and simplifications to find the balance between accuracy and cost-effectiveness. For example, it is often assumed that actuators and motors constantly work efficiently, since their loss usually does not have a significant impact on the overall energy trend. Minor energy consumers such as sensors, controllers, and communication modules are typically neglected when their contribution is small. Simplified dynamic models and predefined motion or speed profiles are also used to limit computational cost while preserving essential system behavior. These simplifications enable effective analysis and optimization of energy consumption without significantly compromising model reliability, refer Figure 3. 2.4 Identification of Energy-Intensive Operations Not all tasks need the same amount of energy. When swift movement requires rapid fluctuations in speed, which makes things harder to move. But when robots have to lift big things, gravity causes electricity use to increase. Paths that are complicated and not straight may be less efficient, since robots need more torque and energy to change direction frequently. Finding these high-energy functions makes it easy to focus on optimization measures, including speeding up, redistributing cargo, and changing the launch configuration, to spend less energy overall. To use energy wisely, collaborative robots need to have accurate kinematic and dynamic models; be able to figure out how much energy are using jointly; and be able to choose tasks that focus on energy use. With this information, researchers can design better optimization algorithms that will reduce energy use while maintaining high operating efficiency. Table 1 lists the main parts and ideas that were used to find out how much energy the collaborative robot needs. The system runs on a standard 24V power supply. Taking into account some
135
Journal of Automation, Mobile Robotics and Intelligent Systems
electromechanical losses, it has a motor efficiency of 85% and a control efficiency of 90%. The computer and sensor consume 25 watts and 20 watts, respectively, and the passive power uses 5 watts. The robot works in a lab 70% of the time, where the temperature is 25°C. Also, a 10% energy recovery efficiency is also assumed to take into consideration the effects of regeneration during deceleration [20]. This makes it easier to have a better idea of how energy-efficient the whole system is.
3. Problem Formulation: Energy Minimization in Collaborative Robots
This study evaluates the electrical energy related to the power consumed by a robot. It calculates total input power over time by considering factors such as supply voltage, motor current, motor efficiency and control system efficiency. Sensors are generally excluded from this calculation. To connect the analytical model to the experimental data, the measured electrical power is converted into voltage and current data. This makes it possible to directly compare two sets of data and introduce changes to the energy model. 3.1 Objective Function The main goal of this study was to reduce the total amount of energy that a collaborative robot uses while following a set path q(t) ∈ Rn over a time period t ∈ [0,T]. As an integral part of the power used by the actuator, the total energy J is shown: min J
q( t ),u( t )
T
P(q(t ), q(t ), (t ), u(t )) dt 0
Where P(⋅) is instantaneous power, τ(t) is joint torques, and u(t) is control inputs. The mechanical power of electromechanical actuators can usually be measured in this way: T
..
T
T
J 0 T (q , q , q) qdt 0 Paux ,i dt i
Where the formula Paux,i represents auxiliary power for grippers and sensors. This expression follows standard formulations for energy-optimal robot trajectories. 3.2 System Constraints The optimization has to follow dynamic, kinematic, and safety rules: I. Dynamic Equilibrium (Rigid-Body Dynamics): M(q)q¨ C(q , q )q g(q) , Where M(q) is the inertia matrix, C(q,q) denotes Coriolis and centrifugal effects, and g(q) is the gravity vector.
136
II. Task Requirements The end-effector must follow a desired trajectory, x(q(t)) → xdesired(t), with minimal deviation, to ensure that the pick-and-place or assembly processes are accurate.
VOLUME 20,
N° 3
2026
III. Joint and Actuator Limits qmin q( t ) qmax,
qmin , q ( t ) qmax , i ( t ) i ,maxi
These ensure mechanical feasibility and prevent overheating and motor saturation.
IV. Payload and Safety Constraints Additional gravitational torques g(q,mp) add up the effects of the payload. To keep people and robots from communicating in harmful ways, cooperative operations put the most extensive limits on safety speed, acceleration, and jerk. 3.3 Multi-Objective Extension All components of the objective function are normalized to eliminate discrepancies between different physical units, thereby obtaining dimensionless quantities that can be directly compared. This generalization ensures that energy consumption, efficiency, and all other factors have the same effect. Weight changes reflect how important each goal is in relation to the others, giving more importance to particular parameters like energy efficiency, rather than just how long it will take to complete an activity. Many weight sensitivity tests, demonstrating their influence on the results of adaptation, underscore the strength of the procedure and support the selected weight. For many tasks, it is necessary to find a balance between the energy cost, how long it takes to finish the task, and how smoothly the movement goes. Hence, a weighted multi-objective cost function is formulated as:
min Jw wE J wT w S T0 x (t ) 2 dt ,
Where wE, wт, wS are tunable weights for energy, time, and smoothness objectives respectively. The last statement punishes a lot of jerks, which helps people feel better and makes the actuators last longer. 3.4 Discretized Optimization Form There is an issue with other limits for the independent optimization form of computing solutions and N time nodes are split up: N
min k qk t qk ,uk k 1
Subject to
M qk
qk 1 2qk qk 1 t 2
C kq k gk k
Heuristic methods — such as nonlinear programming (like IPOPT), direct collocation, particle swarm optimization (PSO), or genetic algorithm (GA) — can help it tackle this non-convex optimization problem.
3.5 Performance Metrics To evaluate the optimization outcomes, the following metrics are adopted: The major performance metrics used to evaluate the launch optimization is presented in Table 2. Total
Journal of Automation, Mobile Robotics and Intelligent Systems
Table 2. Metrics to evaluate the optimization outcomes Sr. No.
Metric
Description
1
Total Energy (J)
Integral of actuator Joules power
Smoothness
Mean-squared jerk
2 3 4
Task Time (T) Optimization Time
Unit
Duration of trajectory
Seconds
Algorithm runtime
Seconds
m2/s5
Figure 4. Energy consumption profile — baseline vs. optimized trajectory
VOLUME 20,
N° 3
2026
that optimization performs a lot of energy consumption while continuing the performance and the launch more easily.
4. Intelligent Optimization Algorithms
4.1 Selection and Justification of Algorithms Controlling collaborative robots in a way that saves energy requires algorithms that can solve nonlinear equations and optimization problems with more than one goal. Particle Swarm Optimization (PSO), Genetic Algorithm (GA), and Reinforcement Learning (RL) are the greatest ways to improve the paths of robots, as they are strong, adaptable, and have been demonstrated to function well [1]. PSO works like birds flocking together, which helps it discover a healthy balance between exploring and exploiting search areas that don’t end. Genetic algorithms (GA) are great for optimization problems that involve selection, crossover, and mutation. They are inspired by biological evolution. Reinforcement learning (RL) — deep reinforcement learning (DRL) specifically — learns the optimum energy-efficient control rules by trying things out and seeing what works in the robot’s changing environment [2]. 4.2 Adaptation to Energy Optimization The purpose of robot trajectory optimization is to consume less energy overall, E = ∫P(t)dt. while obeying regulations concerning safety, job accuracy, and joint restrictions. Each algorithm changes this in its own way:
Figure 5. Performance metrics— baseline vs. optimized energy (j) represents an integral part of the actuator power over time— the total energy consumption. The action time (tasks) measures the total time period required to complete the proposal. By reflecting the consistency and relief of speed, the proportion of smoothness (m2/s5) is determined by using an average-tight jerk. Optimization time (s) represents the computer runtime of optimization algorithm, evaluating its efficiency and practical viability for real-time robotic applications. Figures 4 and 5 show how energy adaptation affects the speed of the machinery. Figure 4 shows an immediate power profile for baseline and optimization routes. The favorable approach indicates a major reduction in the peak of power, which means that the speed is more systematic and uses less energy. Figure 5 shows the overall performance of the metropolis. This shows that the total energy consumption has been reduced by about 25%; the time to complete the project is the same; and the measurement of the Piscesspeed smoothness has decreased. These results show
• PSO modifies the positions of particles that stand for trajectory control points to lower the energy and jerk costs of the actuator.
• GA changes the sets of trajectory parameters (joint velocities and accelerations) to optimize the balance of smoothness and energy.
• RL agents get a reward signal that is opposite to how much energy and jerk they consume. This makes them take smoother, more energy-efficient courses [3]. 4.3 Algorithm Workflow Figure 6 shows the main steps of the optimization process. The first step is to think about plausible responses, which may be particles, chromosomes, or policy settings. An objective function looks at each solution’s energy use, its pace of task completion, and length of time it takes to complete. Algorithms use either heuristic or learning-based methods to repeatedly iterate on solutions until they meet convergence requirements, which could be a steady-state error or the lowest energy. 4.4 Multi-Objective Optimization Robots work in the real world, which presents a lot of diverse needs that don’t always line up. For instance, robots need to be quick and precise while using as lit-
137
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Table 3. Comparison of Optimization Algorithms for Robot Energy Efficiency Sr. No.
Algorithm
Strengths
Limitations
Suitable Application
1
PSO
2
GA
Fast convergence, easy implementation
May stagnate in local minima
Continuous trajectory optimization
3
RL
Adaptive, model-free learning
Requires extensive training data Dynamic, uncertain environments
Strong global search, flexible encoding
Computationally heavy
Complex multi-variable problems
setup, test cases, and performance measurements that were utilized to see how well the algorithms function.
Figure 6. Generic workflow of intelligent optimization algorithms tle energy as possible. NSGA-II (Non-dominated Sorting Genetic Algorithm) and MOPSO (Multi-Objective PSO) are two examples of multi-objective algorithms that create a Pareto front. This front has a curve that shows how performance and energy measurements are related. The person in charge can choose the best place based on how important a task is. They might, for example, choose slower but more energy-efficient actions to put things together and faster movements to move items [4]. Table 3 compares three smart optimization methods for robotic energy efficiency: swarm optimization, genetic algorithms, and reinforcement learning. Particle swarm optimization is useful for continuous trajectory optimization because it converges quickly and is easy to use. However, it can become stuck in local minima. Genetic algorithms are effective at searching the whole space and solving complex problems with many variables, but they need a lot of computing power. Reinforcement learning embodies adaptable, norm-free learning and thrives in dynamic or uncertain contexts, but it also requires substantial training data. Each algorithm offers a distinct trade-off between computational expense, adaptability, and optimization efficacy for energy-efficient robotic functions.
5. Simulation and Experimental Setup
138
Both, virtual modeling and actual prototyping are required to employ energy-efficient optimization approaches for collaborative robots (cobots). This section talks about the computer tools, experimental
5.1 Software and Hardware Platforms We utilized ROS–Gazebo and MATLAB/Simulink to run the simulation and see how well the robot could model and control itself. The Robotics System Toolbox made it easier to model kinematics and dynamics in MATLAB. There were ways to make things better in the Optimization Toolbox, such as Particle Swarm Optimization and Genetic Algorithms. ROS allowed us to operate things in real time, while Gazebo enabled us to watch and interact with simulations that were based on real physics. We tested the hardware by using it with a UR5e collaborative robot that had a 6-DOF manipulator and a two-finger parallel gripper. The robot came with current sensors for each joint actuator and a torque feedback loop that displayed how much power was being utilized at any given moment.
5.2 Test Tasks and Scenarios Two examples of cooperative projects were the subject of this review: Pick-and-Place Operation: The collaborative robot had to move cylindrical items weighing 0.3 kg from one location to another that was 0.6 m away. This process takes a lot of energy because it requires a lot of speeding up and slowing down. Assembly Operation: The robot had to do a pegin-hole activity that required precise alignment and torque control. This task is used to demonstrate how people and robots can work together in manufacturing. We employed both baseline PID control and better trajectory control, picking the optimum method for each case.
5.3 Optimization Algorithm Parameters The optimization was set up to make sure that calculations were quick and that convergence was always possible. The table below demonstrates some of the most frequent techniques to set up algorithms: The reinforcement learning approach in this work is classical Q-learning, as differentiated from deep reinforcement learning (DRL). The state space includes discretized robot motion and energy variables, while the action space comprises adjustments for energy reduction. The reward function penalizes high energy usage and promotes task completion. An ε-greedy exploration strategy balances
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Table 4. Algorithm setting Sr. No. 1 2 3
Algorithm PSO GA
RL (Q-Learning)
Population Size
Iterations
30
100
50 —
150
500 episodes
Learning Rate / Crossover Prob.
Convergence Criterion
0.7 (inertia w)
ΔE < 1 × 10⁻³ J
0.8 (crossover) / 0.05 (mutation) α = 0.1, γ = 0.9
Fitness change < 0.01 % Average reward > 0.95 Rₘₐₓ
5.4 Metrics for Evaluation The following things were used to evaluate performance: • Energy Saved (%): The difference between baseline energy use and the optimal energy use.
• Task Accuracy (mm): The distance between the end-effector’s final location and the correct one. Figure 7. Performance comparison of optimization algorithms
• Computational Cost (s): The amount of time it took for an algorithm to reach a solution.
• Motion Smoothness (Mean Jerk): How smooth the path was and how much stress it put on the equipment.
5.5 Visualization Figure 7 compares the performance of PSO, GA, and RL on the pick-and-place test. The bar chart indicates that PSO achieved the correct balance between saving energy (28%) and processing time. GA used much less energy (22%), but took longer to run. RL, on the other hand, was more flexible and could keep decreasing energy use by 25% as long as it got adequate training.
6. Results and Discussion
Figure 8. Energy consumption profile — baseline vs. optimized
exploration and exploitation, and learning stability is maintained through bounded rewards, learning rate control, and convergence monitoring during training. Table 4 shows the settings used with PSO, GA and RL to help the robots consume less power. The PSO had 30 people, 100 iterations and an inertia weight of 0.7. It stopped when the energy change fell below 1×10−3J. The crossover and mutation probabilities were 0.8 and 0.05, respectively, with 50 individuals and 150 replicates in GA. When the change in fitness was less than 0.01%, it was terminated. It ran 500 episodes of reinforcement learning (Q-learning) with a discount factor of γ = 0.9 and a learning rate of α = 0.1. It works when the average reward is greater than 95% of the maximum value (Rmax). After a sensitivity assessment, these settings were chosen to establish the optimal balance between speed and energy savings.
Clever optimization methods like Particle Swarm Optimization (PSO), Genetic Algorithm (GA), and Reinforcement Learning (RL) make it much easier for robots to work together and save energy. Simulations and hardware-in-the-loop testing were performed using a UR5e cobot conducting a normal pick-and-place task on the MATLAB/Simulink and ROS-Gazebo platforms.
6.1 Energy Consumption Comparison Figure 8 demonstrates how energy use changes before and after optimization. Because the joints were moving swiftly, the baseline trajectory featured a lot of high power peaks, up to 130 W. After using PSO for optimization, the instantaneous power peaks dropped by more than 25%, and the total energy use dropped from 480 J to 340 J — a 29% decline. GA lowered energy expenditure by 22%, and after enough training sessions, RL exhibited adaptive gains that led to a 25% cut. Table 5 demonstrates that PSO was the optimal approach to save energy and keep expenses low. RL made paths smoother and less jerky, but required longer to master. 6.2 Task Performance and Efficiency Not only did optimization lead to lower energy consumption, but it also made the path smoother and the
139
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Table 5. Energy Consumption and Task Efficiency Comparison Sr. No.
Algorithm
Total Energy (J)
Reduction (%)
Task Time (s)
Mean Jerk (m²/s⁵)
1
Baseline
480
—
7.2
1.25×10⁴
4
RL
360
25
7.0
7.8×10³
2 3
PSO GA
340 375
29
7.1
22
7.4
8.5×10³ 9.2×10³
• For RL, reduced learning rates made convergence more constant but adaptation slower. • PSO could deal with changes in parameters and always converged in 30 iterations or less.
This research reveals that changing the settings is highly crucial for getting the proper balance between cost and performance.
7. Conclusion Figure 9. Task completion time and smoothness comparison
Figure 10. Sensitivity analysis — effect of algorithm parameters work more accurate. The new paths made the joints move more smoothly, which put less strain on the actuators so that the machines would last longer. Figure 9 illustrates how optimized trajectories kept the time it took to finish a job the same or better; saving energy didn’t entail losing speed or accuracy. These results show that we can achieve energy-efficient trajectories without slowing down the work rate, which is particularly significant for industrial and collaborative robotics. 6.3 Sensitivity Analysis We did a sensitivity study to see how the algorithmic settings affected the results. Figure 10 shows how the proportion of energy savings changes with the size of the population (for PSO and GA) and the learning rate (for RL). 140
• The calculations were more accurate with more particles, although they took longer than 50 particles.
This study presents an advanced optimization framework to achieve energy-efficient control of collaborative robots through the integration of contemporary methods such as particle swarm optimization (PSO), genetic methods (GA), and reinforcement learning (RL). The comprehensive analysis of system modeling, issue formulation, simulation, and experimental validation in this study shows that energy optimization can be effectively achieved without compromising performance, speed, or safety. The main results show that using PSO and RL to optimize robot paths can reduce energy usage by 25-30% with no impact on task completion time. PSO works well when activities are well-organized, while RL was ideal for adaptive control and changing conditions. Kinematic and dynamic modeling of the UR5e robot, along with extensive energy modeling at the actuator level, provided us with a good platform for optimization. Sensitivity studies have shown that population size, iteration limit, and the structure of the reward function all significantly influence the performance of the optimization. The major contributions of this research are as follows: 1. D evelopment of a comprehensive energy utilization paradigm involving mechanical and electrical subsystems of collaborative robots.
2. D evelopment of a framework for multi-objective optimization that balances energy consumption, speed, smoothness, and job length. 3. U sing items 1 and 2 in a production assembly line, showing how they work in the real world by saving money on electricity and making robots last longer. 4. C omparison of PSO, GA, and RL algorithms, showing their advantages and disadvantages for various tasks.
This study contributes to the broader goals of green manufacturing and Industry 5.0 from both
Journal of Automation, Mobile Robotics and Intelligent Systems
industrial and sustainability perspectives. It focuses on how people and robots can work together using less energy. The optimization was evaluated primarily using a single robot model (UR5e) with specific task categories, such as item manipulation and assembly. Its usefulness at the scope of multi-robot systems, complex layouts, and unstructured human interaction scenarios remain to be confirmed. Developing a realtime implementation remains impossible because the algorithm is so difficult to understand. Future research should use AI-driven predictive maintenance, hybrid optimization techniques (integrating PSO-RL or GA-RL), and digital twin simulation. Such modifications could lead to the development of robots that can change their behavior and adapt to their environment, saving energy over time. This will help create sophisticated industrial robots that will last for a long time. AUTHORS Nandkishor Marotrao Sawai – Sandip Institute of Technology & Research Centre (SITRC), Mahiravani, Trimbak Road, Nashik, Maharashtra 422213, India, email: sawai.nm@gmail.com. Satpalsing Devising Rajput – Pimpri Chinchwad University, Sate, Maval, Pune 412106, Maharashtra, India, email: rajputsatpal@gmail.com. Minal Vilas Gade – Sandip Institute of Technology & Research Centre (SITRC), Mahiravani, Trimbak Road, Nashik, Maharashtra 422213, India, email: minalgade2002@gmail.com. Dipak D. Bage – Pimpri Chinchwad University, Sate, Maval, Pune 412106, Maharashtra, India, email: dipakdbage@gmail.com. Aniruddha S. Rumale*– Sandip Institute of Technology & Research Centre (SITRC), Mahiravani, Trimbak Road, Nashik, Maharashtra 422213, India, email: arumale@ gmail.com. ∗ Corresponding author
References
[1] A. Ijaz, H. Haghbayan, A. Malik, E. Nigussie, and J. Plosila, “Towards Optimizing Communication Cost in Energy Efficient IoT Devices for Swarm Robotics,” Procedia Computer Science, vol. 265, 2025, pp. 49–56. [2] M. He, Y. Chen, M. Liu, X. Fan, and Y. Zhu, “Reliable and Energy-Efficient Communications in Mobile Robotic Networks by Collaborative Beamforming,” ACM Transactions on Sensor Networks, vol. 20, no. 5, 2024, pp. 1–24. [3] S. M. Shafaei and H. Mousazadeh, “An Operational Scrutinization of Autonomous Tractor-Trailer Robot Considering Motion Resistance Force of Rubber Tracked Undercarriage,” Cognitive Robotics, vol. 3, 2023, pp. 173–184. [4] P. Li and L. Yang, “Conflict-Free and Energy-Efficient Path Planning for Multi-Robots Based on Priority Free Ant Colony Optimization,” Math-
VOLUME 20,
N° 3
2026
ematical Biosciences and Engineering, vol. 20, 2023, pp. 3528–3565. [5] K. Nonoyama, Z. Liu, T. Fujiwara, M. M. Alam, and T. Nishi, “Energy-Efficient Robot Configuration and Motion Planning Using Genetic Algorithm and Particle Swarm Optimization,” Energies, vol. 15, 2022, 2074. [6] F. Yu, C. Lu, L. Yin, and B. Zhang, “A Self-Learning Memetic Algorithm for Human–Robot Collaboration Scheduling in Energy-Efficient Distributed Mixed Fuzzy Welding Shop,” IEEE Transactions on Automation Science and Engineering, 22, 2024, pp. 6595–6607. [7] M. Zhang, M. Zhou, L. Zhang, and Z. Zhang, “Dual-Self-Learning Co-Evolutionary Algorithm for Energy-Efficient Flexible Job Shop Scheduling Problem with Processing-Transportation Composite Robots,” Scientific Reports, vol. 15, no. 1, 2025, 28716. [8] X. Wen, Y. Sun, H. L. Ma, and S. H. Chung, “Green Smart Manufacturing: Energy-Efficient Robotic Job Shop Scheduling Models,” International Journal of Production Research, vol. 61, no. 17, 2023, pp. 5791–5805. [9] C. Lu, R. Gao, L. Yin, and B. Zhang, “Human–Robot Collaborative Scheduling in Energy-Efficient Welding Shop,” IEEE Transactions on Industrial Informatics, vol. 20, no. 1, 2023, pp. 963–971. [10] Q. Li, T. W. Fan, L. S. Kei, and Z. Li, “Scalable and Energy-Efficient Task Allocation in Industry 4.0: Leveraging Distributed Auction and IBPSO,” PLOS ONE, vol. 20, no. 1, 2025, e0314347. [11] P. Morasso, “Mental Simulation of Actions for Learning Optimal Poses,” Cognitive Robotics, vol. 3, 2023, pp. 185–200. [12] M. Rahmati, “Dynamic Role-Adaptive Collaborative Robots for Sustainable Smart Manufacturing: An AI-Driven Approach,” Journal of Intelligent Manufacturing and Special Equipment, vol. 6, no. 2, 2025, pp. 105–115. [13] P. Luo, X. Zhang, and R. Meng, “Co-Attention Learning Cross Time and Frequency Domains for Fault Diagnosis,” Cognitive Robotics, vol. 3, 2023, pp. 34–44. [14] M. Wu, C. F. Yeong, E. L. Su, W. Holderbaum, and C. Yang, “A Review on Energy Efficiency in Autonomous Mobile Robots,” Robotic Intelligence and Automation, vol. 43, no. 6, 2023, pp. 648–668. [15] X. Jin, Y. Rong, K. Liu, C. Xiao, and X. Zhang, “A Colorization Method for Historical Videos,” Cognitive Robotics, vol. 3, 2023, pp. 201–207. [16] N. M. Sawai, A. R. Mankar, R. Mahant, K. Ingole, Y. R. Mahulkar, and S. M. Asutkar, “Optimization of HOG Algorithm for Computer Vision in Industrial Robots,” in Proc. IEEE International Conference on Computing, Power and Communication Technologies (IC2PCT), vol. 5, 2024, pp. 572–577.
141
Journal of Automation, Mobile Robotics and Intelligent Systems
[17] L. Yang and S. Deng, “LP-BT: A Location Privacy Protection Algorithm Based on Ball Trees,” Cognitive Robotics, vol. 3, 2023, pp. 127–134. [18] S. Sahoo, R. K. Dalei, S. K. Rath, and U. K. Sahu, “Selection of PSO Parameters Based on Taguchi Design–ANOVA–ANN Methodology for Missile Gliding Trajectory Optimization,” Cognitive Robotics, vol. 3, 2023, pp. 158–172.
142
VOLUME 20,
N° 3
2026
[19] M. Soori, B. Arezoo, and R. Dastres, “Optimization of Energy Consumption in Industrial Robots: A Review,” Cognitive Robotics, vol. 3, 2023, pp. 142–157. [20] S. Miranda, C. R. Vá� zquez, and M. Navarro-Gutiérrez, “Energy Consumption Analysis and Optimization in Collaborative Robots,” Frontiers in Robotics and AI, vol. 12, 2025, 1671336.
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
A GENERAL FORM OF CAUTIOUS APPROXIMATE REASONING FOR SYMBOLIC AND COMPLEX DATA Submitted: 29th May 2024; accepted: 23rd July 2024
Saoussen Bel Hadj Kacem DOI: 10.14313/jamris-2026-046 Abstract: Decision-makers are increasingly confronted with the problem of data imprecision. Thereby,the representation and manipulation of such knowledge play a crucial role in the performance of intelligent systems. Among the wellknown logics for imprecise data, we can find symbolic multi-valued logic. This logic extends classical logic to consider a scale of degrees including intermediate symbolic degrees, between True and False. It is based on multi-set theory, where every predicate is modeled by a multi-set. Approximate reasoning in that context consists of inferring with an observation whose degree is different from that of the rule premise. We are interested in an approximate reasoning based on the implication operator. In this paper, we prove that this approximate reasoning checks the axiomatics of approximate reasoning. Furthermore, we improve this approximate reasoning to deal with multi-sets having different scales bases. And thus, propositions of the rule premise and rule conclusion can be associated with different scale bases. We then give solutions for deduction schema in more complex cases. More precisely, we pointed out the presence of complex rules and the combination of conclusions. Keywords: Approximate reasoning, Generalized Modus Ponens, Multi-valued logic, Rule-based systems, complex rules
1. Introduction The main challenge of rule-based systems is to consider the imprecision of the handled knowledge. It becomes a necessity since humans generally provide this type of data. It is no longer adequate to use Modus Ponens, which for a rule “If 𝑋 is 𝒜 then 𝑌 is ℬ”, it can only infer with an observation equal to the rule premise “𝑋 is 𝒜”. Therefore, approximate reasoning was proposed by Zadeh [1] in fuzzy context [2] in order to consider imprecise knowledge in the inference process. It is based on the generalization of Modus Ponens of classical logic, called Generalized Modus Ponens (GMP). The general schema of GMP is as follows: If 𝑋 is 𝒜 then 𝑌 is ℬ 𝑋 is 𝒜′ 𝑌 is ℬ ′
(1)
where 𝑋 and 𝑌 are linguistic variables and 𝒜, 𝒜′ , ℬ and ℬ ′ are predicates.
In classical logic, a rule is triggered only if there is a perfect match between its antecedent “𝑋 is 𝒜” and the observation. However, with GMP, an observation 𝑋 is “𝒜′ ” can differ from the antecedent “𝑋 is 𝒜”. And thus, the result “𝑌 is ℬ ′ ” will be not equal to the rule consequence “𝑌 is ℬ”. The aim is to determine the conclusion “𝑌 is ℬ ′ ” so that it will be as close as possible to what an expert would decide. For that purpose, some authors [3] proposed criteria that should be verified by the conclusion to consider that it is valid and effective, and this was in the fuzzy context. In [4], the authors proposed the following generalization of approximate reasoning axiomatics which appeared in [3]: 𝒜′ = 𝒜 ⇒ ℬ′ = ℬ 𝒜′ is a reinforcement of 𝒜 ⇒ the more 𝒜′ is a reinforcement of 𝒜, the more ℬ ′ is a reinforcement of ℬ C II-2 𝒜′ is a reinforcement of 𝒜 ⇒ ℬ ′ = ℬ C III 𝒜′ is a weakening of 𝒜 ⇒ the more 𝒜′ is a weakening of 𝒜, the more ℬ ′ is a weakening of ℬ (2) This set of criteria describes the behavior that any approach of approximate reasoning should have, with the aim of obtaining relevant and coherent conclusions. Criterion I ensures the verification of classical Modus Ponens. Indeed, when the observation is equal to the rule premise 𝒜, the conclusion must also be equal to the rule conclusion ℬ. This is a natural and logical behavior that is in concordance with classical logic as well as human thinking. Criterion II-1 and II-2 are interested in the case where the observation is a reinforcement of the premise. So, two positions can be adopted. The first one is to consider that the inference conclusion ℬ ′ must also be a reinforcement of the rule conclusion ℬ. Whereas the second one imposes that the inference conclusion ℬ ′ is equal to the rule conclusion ℬ. In effect, the choice between criterion II-1 and criterion II-2 depends on the causality between the premise and the conclusion of the rule. For example, with the rule “If a tomato is red then it is ripe”, we know that the more a tomato is red, the more it is ripe. With the observation “the tomato is very red” we can conclude that “The tomato is very ripe”. However, the expert can choose to have a cautious attitude by keeping as inference conclusion the rule conclusion as it is without considering the reinforcement of the CI C II-1
Open Access. © 2026 Saoussen Bel Hadj Kacem, published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP.
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License
143
Journal of Automation, Mobile Robotics and Intelligent Systems
observation. For example, with the rule “If the fog is slight then the speed is high” and the observation “The fog is very slight”, it is more adequate to deduce that “The speed is high” rather than “The speed is very high”. Indeed, following the reinforcement of the premise can, in this case, create a risk of accident. Finally, criterion III focuses on the case where the observation is a weakening of the premise. In that case, the inference conclusion ℬ ′ must also be a weakening of the rule conclusion ℬ. According to these criteria, two types of approximate reasoning can be released [3]: Type 1: Type 2:
criteria I, II-1, III criteria I, II-2, III
(3)
These types are differentiated by the behavior of approximate reasoning when the observation is a reinforcement of the antecedent. The first adopts criterion II-1, and the second adopts criterion II-2. Criteria I and III must be checked in any case. The approximate reasoning type is fixed by the expert, who is knowledgeable of the nature of manipulated data and their behavior. Indeed, the type depends on the causality between the antecedent and the conclusion, which can only be evaluated by the expert. Fuzzy reasoning has proven its performance in several applications, especially in fuzzy controls [5–9]. But some disadvantages of fuzzy set theory were pointed out by many researchers [10–16]. The principal limit is the numerical definition of fuzzy sets’ universe, which complicates the formalization of abstract terms like beautiful, intelligent, etc., since they do not have a numerical universe. Moreover, in fuzzy rule-based systems, it remains a problem to have a linguistic interpretation of the fuzzy set 𝐵′ , the result of the inference. For that, multi-valued logic [13, 17–20] has been defined in order to handle imprecise knowledge in a symbolic way. This logic, as fuzzy logic, allows treating imprecise knowledge by authorizing intermediate degrees between the full membership and the non-membership. As proposed by De Glas [14], it is based on multi-sets theory. It evaluates knowledge imprecision in a qualitative way, which is more natural and is part of human thinking. We focus in our work on approximate reasoning in the context of symbolic multi-valued logic. As a result of our bibliographic research within the framework of symbolic multi-valued logic, we have found that only the work [21] considers the axiomatics and the typology (3) to define approximate reasoning approaches. Indeed, two Generalized Modus Ponens for multivalued knowledge have been proposed: the first one is of type 1, and the second is of type 2. The proposed approximate reasoning of type 1 was studied and extended in other works [22, 23]. It was also applied and used in a diagnosis aid system for autism [24–26] and has given very satisfactory results. The success of this system is due first to the verification of axiomatics. And also, psychiatry handles principally abstract data, where symptoms are non-measurable and are based mainly on the perception of doctors. 144
VOLUME 20,
N∘ 3
2026
Using symbolic multi-valued logic in that environment is then appropriate compared to fuzzy logic. We are interested in this paper in another problem that concerns the form of manipulated predicates in symbolic approximate reasoning of type 2. In main and known approaches of symbolic approximate reasoning [10, 21, 27], all the multi-sets of the knowledge base must have the same degree scale. This may present a limit for the user in the sense that he will not be free in the definition of the multisets of the system. For that, our aim in this paper is to give a new model of approximate reasoning that can consider multi-sets that do not necessarily have the same scale base. We also treat another problem in the inference process, which is the form of rules. Indeed, experts can express rules that are not simple. For example, the antecedent of a rule can be complex, so that is composed of a conjunction or a disjunction of propositions. Another inference model is then proposed to deal with these types of rules. This paper is organized as follows. In section 2., we present the symbolic multi-valued logic, which represents the context of our work. We also illustrate some reasoning models in that context and present the model of approximate reasoning of type 2. Section 3. gives proofs that this inference schema checks the axiomatics of approximate reasoning. We then propose in Section 4. an improvement of our multi-valued approximate reasoning for multi-bases knowledge. In section 5., we treat the case of reasoning with complex rules. We extend this reasoning so that it supports the simultaneous presence of complex rules and multibase knowledge in Section 6.. A discussion is given in Section 7., we finally complete this paper with a conclusion in Section 8..
2.
Reasoning in symbolic context
We present in this section the symbolic multivalued logic according to the definition of Akdag et al. [17]. Then, we present some approximate reasoning models in that context. 2.1.
Symbolic multi-valued logic
Symbolic multi-valued logic was proposed as a solution to consider imprecise knowledge in intelligent systems. Unlike classical logic, it allows assigning to an affirmation a value that is between the True and the False. Thus, the set of degrees forms an ordered list ℒ𝑀 = {𝜏0 , ..., 𝜏𝑖 , ..., 𝜏𝑀−1 } with 𝑀 the number of truthdegrees, with the total order relation: 𝜏𝑖 ≤ 𝜏𝑗 ⇔ 𝑖 ≤ 𝑗. By correspondence with classical logic, the False is the smallest element 𝜏0 , and the True is the greatest one 𝜏𝑀−1 [13, 17, 27]. If we have, for example, a scale ℒ5 = {𝜏0 , 𝜏1 , 𝜏2 , 𝜏3 , 𝜏4 }, 𝜏0 is the False, 𝜏4 is the True and 𝜏1 , 𝜏2 , 𝜏3 are intermediate degrees between them. The predicate in multi-valued logic is formalized by a multi-set 𝐴 accompanied by a degree 𝜏𝛼 from ℒ𝑀 . So the general form of a proposition is: “𝑋 is 𝐴” is 𝜏𝛼 -true
Journal of Automation, Mobile Robotics and Intelligent Systems
0
1
2
3
4
not-at-all
little
moderately
enough
completely
Figure 1. Representation of the scale ℒ5 with 𝑋 a linguistic variable, 𝐴 a multi-set, and 𝜏𝛼 a symbolic degree. For example, the statement “John is tall” is 𝜏3 -true, with 𝜏3 ∈ ℒ5 , means that John satisfies the multi-set tall with the degree 𝜏3 . To approach the human language, we can associate with each symbolic degree 𝜏𝛼 a linguistic degree 𝑣𝛼 . For example, we can assign to ℒ5 the base ℒ5 = {notat-all, little, moderately, enough, completely} (see Figure 1). So a proposition can have the following form: 𝑋 is 𝑣𝛼 𝐴 The statement “John is tall” is 𝜏3 -true can become “John is enough tall”. As in fuzzy logic, in multi-valued logic, we sometimes need to aggregate several truth degrees (for example in a conjunction of propositions). T-norms, Tconorms and implication operators of fuzzy logic can be used in the context of multi-valued logic. However, instead of manipulating numbers belonging to [0, 1], we aggregate symbolic degrees belonging to ℒ𝑀 . Thus, in multi-valued logic, T-norms, T-conorms and implication operators are applications of: ℒ𝑀 × ℒ𝑀 → ℒ𝑀 . They have the same properties and definitions as in fuzzy logic. In multi-valued logic, the aggregation functions of Łukasiewicz are often preferred. In this context and with 𝑀 truth-degrees, they are defined by:
2.2.
𝑇𝐿 (𝜏𝛼 , 𝜏𝛽 ) = 𝜏𝑚𝑎𝑥(𝛼+𝛽−𝑀+1,0)
(4)
𝑆𝐿 (𝜏𝛼 , 𝜏𝛽 ) = 𝜏𝑚𝑖𝑛(𝛼+𝛽,𝑀−1)
(5)
ℐ𝐿 (𝜏𝛼 , 𝜏𝛽 ) = 𝜏𝑚𝑖𝑛(𝑀−1−𝛼+𝛽,𝑀−1)
(6)
Linguistic modifiers
The notion of linguistic modifiers [28] was proposed by Zadeh in the fuzzy logic context in order to extend the use of fuzzy sets. In fact, they correspond to functions that, when applied to a fuzzy set, generate new ones. This notion is adopted in symbolic multi-valued logic by Akdag et al. [29–31] and named Generalized Symbolic Modifiers (GSM). Since a term is modeled by a multi-set (a degree on a scale), the modification of data in multi-valued logic is done by a transformation of the degree or/and the scale of the multi-set. This can classify the GSM according to two criteria: the variation of the base size 𝑀 and the 𝑖 variation of the proportion . The variation of the 𝑀−1 base size engenders an erosion or a dilation of the original scale, which is translated by a decrease or an increase in its size 𝑀 respectively. The variation of the proportion leads to a weakening or a reinforcement of the notion expressed by the multi-set and is translated by the decrease or the increase of the proportion, respectively. Table 1 illustrates the defined modifiers
VOLUME 20,
N∘ 3
2026
in [30]. We added to the table a new operator 𝐶𝐶 (for Central Conservative) that we introduced in [23]. It is a neutral operator, which means that it does not affect the multi-set and changes neither the base nor the degree. The radius 𝜌 of the modifier 𝑚𝜌 expresses its power, i.e., how strong the effect of the modifier is on the multi-set. For example, suppose that we have a couple of degree/base (𝜏3 , ℒ5 ). We aim to apply to this couple the modifier 𝐷𝑊2 . As we can see in table 1, 𝐷𝑊 refers to the Dilation Weakening modifier. According to the definition of this modifier, the resulted degree after the modification is 𝜏𝑖 ′ = 𝜏𝑖 , which means that the degree will be the same and it will not be changed. However, 𝐷𝑊 acts on the base size as follows: ℒ𝑀′ = ℒ𝑀+𝜌 , which means that it increases the base size by 𝜌. Thus, the modification of the couple (𝜏3 , ℒ5 ) by the operator 𝐷𝑊2 gives the couple (𝜏3 , ℒ7 ). We can see from this example that the modifier 𝐷𝑊2 erodes the base ℒ5 because its size is increased by 2 and becomes ℒ7 . Also, this modifier is weakening because the pro3 3 portion is decreased from = 0.75 to = 0.5. 5−1
2.3.
7−1
Symbolic approximate reasoning
In multi-valued logic based on multi-set theory, approximate reasoning consists of considering an observation with a degree different from the premise degree. The first generation of approximate reasoning in the symbolic multi-valued framework was proposed by Akdag [10]. Its schema is the following: If 𝑋 is 𝐴 then 𝑌 is 𝐵 is 𝜏𝛼 -true 𝑋 is 𝐴 is 𝜏𝛽 -true 𝑌 is 𝐵 is 𝜏𝛾 -true with 𝜏𝛾 = 𝑇(𝜏𝛼 , 𝜏𝛽 )
(7)
where 𝐴 and 𝐵 multi-sets, 𝜏𝛼 , 𝜏𝛽 and 𝜏𝛾 ∈ ℒ𝑀 and 𝑇 a T-norm. This reasoning considers only free multi-valued rules, i.e., rules whose premise is completely true. For that, Khoukhi [27] proposed another approximate reasoning with strong rules, i.e., rules whose both premise and conclusion are accompanied by an imprecision degree: If 𝑋 is 𝑣𝛼 𝐴 then 𝑌 is 𝑣𝛽 𝐵 𝑋 is 𝑣𝛾 𝐴 𝑌 is 𝑣𝜆 𝐵 with 𝜏𝜆 = 𝑇(𝑆𝑖𝑚(𝜏𝛼 , 𝜏𝛾 ), 𝜏𝛽 )
(8)
with 𝑆𝑖𝑚 a similarity measure defined by 𝑆𝑖𝑚(𝜏𝑎 , 𝜏𝑏 ) = 𝑚𝑖𝑛(ℐ(𝜏𝑎 , 𝜏𝑏 ), ℐ(𝜏𝑏 , 𝜏𝑎 )). Here, 𝑣𝛼 , 𝑣𝛽 , and 𝑣𝛾 are linguistic degrees associated with multi-valued degrees 𝜏𝛼 , 𝜏𝛽 , and 𝜏𝛾 . Approximate reasoning of Khoukhi (8) gives more flexibility than approximate reasoning of Akdag (7) to model domain expertise. But unfortunately, the fact that the premise is accompanied by a truth degree engenders a new problem, which is the accordance with the axiomatics of approximate reasoning. Indeed, it was proved in [4] that the approximate reasoning (8) 145
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 1. Definitions of weakening, reinforcing and central modifiers MODE NATURE Erosion
Weakening 𝜏𝑖 ′ = 𝜏𝑚𝑎𝑥(0,𝑖−𝜌)
𝐸𝑊𝜌
ℒ𝑀′ = ℒ𝑚𝑎𝑥(2,𝑀−𝜌)
Reinforcing 𝜏𝑖 ′ = 𝜏 𝑖 ℒ𝑀′ = ℒ𝑚𝑎𝑥(𝑖+1,𝑀−𝜌)
Central 𝐸𝑅𝜌
𝜏𝑖 ′ = 𝜏𝑚𝑎𝑥(⌊ 𝑖 ⌋,1) 𝜌
𝜏𝑖 ′ = 𝜏𝑚𝑖𝑛(𝑖+𝜌,𝑀−𝜌−1)
𝐸𝐶𝜌 *
𝐸𝑅𝜌′
ℒ𝑀′ = ℒ𝑚𝑎𝑥(⌊ 𝑀 ⌋+1,2)
𝜏𝑖 ′ = 𝜏𝑖+𝜌 ℒ𝑀′ = ℒ𝑀+𝜌
𝐷𝑅𝜌
𝜏𝑖 ′ = 𝜏𝑖𝜌 ℒ𝑀′ = ℒ𝑀𝜌−𝜌+1
𝐷𝐶𝜌
𝜏𝑖 ′ = 𝜏𝑚𝑖𝑛(𝑖+𝜌,𝑀−1) ℒ 𝑀′ = ℒ 𝑀
𝐶𝑅𝜌
𝜏𝑖 ′ = 𝜏 𝑖 ℒ 𝑀′ = ℒ 𝑀
𝐶𝐶
𝜌
ℒ𝑀′ = ℒ𝑚𝑎𝑥(2,𝑀−𝜌) 𝜏𝑖 ′ = 𝜏 𝑖 ℒ𝑀′ = ℒ𝑀+𝜌 Dilation 𝜏𝑖 ′ = 𝜏𝑚𝑎𝑥(0,𝑖−𝜌) ℒ𝑀′ = ℒ𝑀+𝜌 𝜏𝑖 ′ = 𝜏𝑚𝑎𝑥(0,𝑖−𝜌) Conservation ℒ 𝑀′ = ℒ 𝑀 * ⌊.⌋ is the floor function.
𝐷𝑊𝜌 𝐷𝑊𝜌′ 𝐶𝑊𝜌
does not check the axiomatics of approximate reasoning (2). In fact, this reasoning verifies criteria I and III but does not verify criterion II. More precisely, when the observation is a reinforcement of the antecedent, the conclusion is a weakening of the consequence. To overcome this problem, it was proposed in [21] two models of approximate reasoning. The first one verifies criteria I, II-1, and III of the axiomatics (2), and so, it is of type 1 of the typology (3). In this inference model, when the observation is a reinforcement of the antecedent, the inference conclusion is also a reinforcement of the rule consequence. The GMP of the first approach is the following:
If “𝑋 is 𝑣𝛼 𝐴” then “𝑌 is 𝑣𝛽 𝐵” “𝑋 is 𝑚(𝑣𝛼 𝐴)” “𝑌 is 𝑚(𝑣𝛽 𝐵)”
Example 1. In this example, we consider the Łukasiewicz implication ℐ𝐿 and the Łukasiewicz T-norm 𝑇𝐿 . Given the list of truth-degrees ℒ7 = {notat-all, very-little, little, moderately, enough, very, completely}, we obtain with the following data:
(9)
where 𝑚 is a linguistic modifier. The principle is to determine the modifier 𝑚 that produces the modification from the rule antecedent to the observation. After that, to have the inference conclusion, the determined modifier 𝑚 is applied to the rule consequence. Nevertheless, in practice, rules are complex, and their antecedents are often a conjunction or a disjunction of propositions. This approach of approximate reasoning was thus improved to be used with complex rules in [22, 23]. It was then integrated into a rulebased system shell called RAMOLI [26], with giving a solution for rule chaining in this context. In order to validate the proposed approach, a diagnostic system for autism, called DAS-AUTISM, was performed in [24, 25], which gave very satisfactory results. This proved the effectiveness of the approximate reasoning of type 1 which is based on linguistic modifiers. The second approach that was proposed in [21] verifies criterion I, II-2, and III. It is thus of type 2 of the typology (3). In this inference model, when the observation is a reinforcement of the antecedent, the inference conclusion is equal to the rule consequence. Our approach is an improvement of the inference schema (8), we have replaced the similarity measure 𝑆𝑖𝑚 by the implication operator ℐ: 146
If 𝑋 is 𝑣𝛼 𝐴 then 𝑌 is 𝑣𝛽 𝐵 𝑋 is 𝑣𝛾 𝐴 𝑌 is 𝑣𝜆 𝐵 with 𝜏𝜆 = 𝑇(ℐ(𝜏𝛼 , 𝜏𝛾 ), 𝜏𝛽 ) (10) The use of the implication operator instead of the similarity measure is explained by the fact that it is asymmetric. And so, the modification undergone from the antecedent to the observation is not considered in the case of reinforcement.
If management is moderately good then job is enough satisfactory (𝜏𝛼 = 𝜏3 , 𝜏𝛽 = 𝜏4 ) Management is little good (𝜏𝛾 = 𝜏2 ) Job is moderately satisfactory (𝜏𝜆 = 𝜏3 )
Indeed, the truth degree of the conclusion is 𝜏𝜆 = 𝑇𝐿 (ℐ𝐿 (𝜏3 , 𝜏2 ), 𝜏4 ) = 𝑇𝐿 (𝜏5 , 𝜏4 ) = 𝜏3 . The resulted degree 𝜏3 corresponds to the linguistic term “moderately”. For the following data: If management is moderately good then job is enough satisfactory (𝜏𝛼 = 𝜏3 , 𝜏𝛽 = 𝜏4 ) Management is enough good (𝜏𝛾 = 𝜏4 ) Job is enough satisfactory (𝜏𝜆 = 𝜏4 )
The truth degree of the conclusion is: 𝜏𝜆 = 𝑇𝐿 (ℐ𝐿 (𝜏3 , 𝜏4 ), 𝜏4 ) = 𝑇𝐿 (𝜏6 , 𝜏4 ) = 𝜏4 , which corresponds to the term “enough”. It is clear in the example that when the observation is a reinforcement of the antecedent (𝜏𝛾 ≥ 𝜏𝛼 ), the conclusion is equal to the consequence (𝜏𝜆 = 𝜏𝛽 ). The reinforcement intensity is retracted due to the use of the implication operator.
Journal of Automation, Mobile Robotics and Intelligent Systems
3. Typology of our symbolic reasoning We give in this section proofs that the approximate reasoning (10) is of type 2. For that, we demonstrate that it verifies criteria I, II-2, and III of the axiomatics (2). Property 1. The GMP (10) verifies criterion I of the approximate reasoning axiomatics (2). Proof 1. Criterion I considers that the observation is equal to the rule premise, thus 𝜏𝛼 = 𝜏𝛾 . In this case, the degree of inference conclusion is: 𝜏𝜆 = 𝑇(ℐ(𝜏𝛼 , 𝜏𝛼 ), 𝜏𝛽 ). We know that for any implication operator, ℐ(𝜏𝛼 , 𝜏𝛼 ) = 𝜏𝑀−1 , so 𝜏𝜆 = 𝑇(𝜏𝑀−1 , 𝜏𝛽 ). We also know that the neutral element of T-norm is 𝜏𝑀−1 , thus 𝜏𝜆 = 𝜏𝛽 . We conclude that the inference conclusion is equal to the rule conclusion. Property 2. The GMP (10) verifies criterion II-2 of the approximate reasoning axiomatics (2). Proof 2. Criterion II-2 treats the case where the observation is a weakening of the rule premise, so 𝜏𝛼 < 𝜏𝛾 . We know that for any implication operator, if 𝜏𝛼 ≤ 𝜏𝛾 then ℐ(𝜏𝛼 , 𝜏𝛾 ) = 𝜏𝑀−1 . Thus, the degree of inference conclusion is: 𝜏𝜆 = 𝑇(ℐ(𝜏𝛼 , 𝜏𝛾 ), 𝜏𝛽 ) = 𝑇(𝜏𝑀−1 , 𝜏𝛽 ). Furthermore, the neutral element of T-norm is 𝜏𝑀−1 , thus 𝜏𝜆 = 𝜏𝛽 . We conclude that the inference conclusion is equal to the rule conclusion. Property 3. The GMP (10) verifies criterion III of the approximate reasoning axiomatics (2). Proof 3. Criteria III concerns the case where the observation is a weakening of the rule premise, i.e., 𝜏𝛾 < 𝜏𝛼 . Moreover, it ensures that the more 𝒜′ is a weakening of 𝒜, the more ℬ ′ is a weakening of ℬ. Let us consider two different observations that weaken the premise and which have the degrees 𝜏𝛾1 and 𝜏𝛾2 , with 𝜏𝛾1 < 𝜏𝛾2 < 𝜏𝛼 . Let us demonstrate that degrees 𝜏𝜆1 and 𝜏𝜆2 of the inference conclusions ℬ1′ and ℬ2′ are so that 𝜏𝜆1 ≤ 𝜏𝜆2 ≤ 𝜏𝛽 . We know that every implication operator and Tnorm verify monotonicity property. A monotonic function is a function that is either entirely non-increasing or non-decreasing in its domain. So in [0, 1], a T-norm will be a monotonic and non-decreasing function and will verify: ∀𝑥, 𝑦, 𝑧, if 𝑥 ≤ 𝑦 then 𝑇(𝑧, 𝑥) ≤ 𝑇(𝑧, 𝑦). In the same way, an implication operator is also monotonic and non-decreasing, which means that ∀𝑥, 𝑦, 𝑧, if 𝑥 ≤ 𝑦 then ℐ(𝑧, 𝑥) ≤ ℐ(𝑧, 𝑦). So: 𝜏𝛾1 < 𝜏𝛾2
⇒ ⇒ ⇒
ℐ(𝜏𝛼 , 𝜏𝛾1 ) ≤ ℐ(𝜏𝛼 , 𝜏𝛾2 ) 𝑇(ℐ(𝜏𝛼 , 𝜏𝛾1 ), 𝜏𝛽 ) ≤ 𝑇(ℐ(𝜏𝛼 , 𝜏𝛾2 ), 𝜏𝛽 ) 𝜏𝜆 1 ≤ 𝜏 𝜆 2
We know also that for any T-norm 𝑇, ∀𝑥, 𝑦 𝑇(𝑥, 𝑦) ≤ 𝑚𝑖𝑛(𝑥, 𝑦), So: 𝜏𝜆2 = 𝑇(ℐ(𝜏𝛼 , 𝜏𝛾2 ), 𝜏𝛽 ) ≤ 𝑚𝑖𝑛(ℐ(𝜏𝛼 , 𝜏𝛾2 ), 𝜏𝛽 ) ≤ 𝜏𝛽 ⟹ 𝜏 𝜆 2 ≤ 𝜏 𝛽 . We deduce that 𝜏𝜆1 ≤ 𝜏𝜆2 ≤ 𝜏𝛽 . This allows us to deduce that the more an observation is a weakening of the rule premise, the more the inference conclusion is a weakening of the rule conclusion.
VOLUME 20,
4.
N∘ 3
2026
Approximate reasoning with multi-bases knowledge
Approximate reasoning (10) imposes that the multi-sets 𝐴 and 𝐵 have the same base ℒ𝑀 . So, when constructing a rule base for a rule-based system, all the defined rules must have antecedents and consequences of the unique multi-valued base. Moreover, all the facts must also have this same base. This constraint could represent a limit for the expert, since it will prevent him from freely expressing his expertise as he assesses. For example, he can judge that these data are as follows: If the sky is moderately clear then the weather is enough nice The sky is little clear The weather is ? nice (11)
with clear and nice are multi-sets with respectively the bases ℒ7 and ℒ5 : ℒ7 ={not-at-all, very-little, little, moderately, enough, very, completely} ℒ5 ={not-at-all, little, moderately, enough, completely} In this circumstance, the GMP (10) cannot be applied directly. Indeed, the inferred conclusion is determined by the T-norm and the implication operators: 𝜏𝜆 = 𝑇(ℐ(𝜏𝛼 , 𝜏𝛾 ), 𝜏𝛽 ). These operators are calculated, by definition, using the base size 𝑀 (see (4) and (6)). But we have in the schema (11) two base sizes: 7 and 5. The problem consists of which base size will be used in the calculation. We aim in this section to define a new approximate reasoning that is able to infer with multi-sets with different bases [32]. The problem can be formalized as the following GMP: If 𝑋 is 𝑣𝛼 𝐴 then 𝑌 is 𝑣𝛽 𝐵 𝑋 is 𝑣𝛾 𝐴 𝑌 is 𝑣𝜆 𝐵
(12)
with 𝑣𝛼 , 𝑣𝛽 , 𝑣𝛾 and 𝑣𝜆 are symbolic degrees, and 𝐴 and 𝐵 are multi-sets. Here, we suppose that 𝐴 and 𝐵 have not the same base. So, we denote by ℒ𝑀𝐴 and ℒ𝑀𝐵 the bases of respectively 𝐴 and 𝐵. And we will consequently have: 𝜏𝛼 ∈ ℒ𝑀𝐴 , 𝜏𝛾 ∈ ℒ𝑀𝐴 , 𝜏𝛽 ∈ ℒ𝑀𝐵 and 𝜏𝜆 ∈ ℒ𝑀𝐵 . To infer in this situation, we propose a method composed of three stages. The first stage is to convert all the manipulated multi-sets of the inference schema to a common base ℒ𝑀 . After that, we infer our approximate reasoning (10) to obtain a conclusion on the common base. Finally, this result is converted to the original base of the rule conclusion ℒ𝑀𝐵 . Let us note that we used these three stages in our approach of approximate reasoning of type 1 that we proposed in [23]. However, the work proposed in [23] concerns an extension of an approximate reasoning whose GMP is different from the one considered in this paper. Thus, the principle of these stages, their contents, and their formulas are different from those of our previous work. Figure 2 describes the different steps of our solution. 147
Journal of Automation, Mobile Robotics and Intelligent Systems
Antecedent
Consequence
N∘ 3
VOLUME 20,
Observation
τ0
τ1
τ3
τ2
2026
τ4
τ0 τ1 τ2 τ3 τ4 τ5 τ6 τ7 τ8 τ9 τ10 τ11 τ12 Interfacing
Antecedent
Consequence
τ0
Observation
τ1
τ3
τ2
τ5
τ4
τ6
Figure 3. Example of determination of the common base ℒ𝑀 with LCM linguistic modifier must change the base to become ℒ𝑀 . The new base size 𝑀 is in all cases a multiple of the base size of the multi-set (subtracted by 1), since it is a LCM of all base sizes. The nature of the linguistic modifier is consequently dilation. As mentioned in the table 1, the modifier that conserves the proportion and increases the base size is 𝐷𝐶 (Dilating Central). We conclude that the linguistic modifier that we will use in the translation is 𝐷𝐶. Let us recall its definition:
Inference
Conclusion
Re-interfacing
𝐷𝐶𝜌 = � Conclusion
Figure 2. Inference with multi-bases multi-sets 4.1.
Interfacing of manipulated multi-sets
The aim of this stage is to translate the manipulated multi-sets of the GMP schema (those of the antecedent, the consequence, and the observation) to a unique base ℒ𝑀 . However, this conversion to the new base must guarantee non-information loss. For that, we propose to calculate the base size 𝑀 using the least common multiple (LCM) [33] of the different base sizes 𝑀𝐴 and 𝑀𝐵 . We subtract 1 from every base size because the numeration of degrees in a multivalued base starts with 0. The size of the new common base is obtained as follows: 𝑀 = 𝐿𝐶𝑀(𝑀𝐴 − 1, 𝑀𝐵 − 1) + 1
(13)
With the use of LCM, every degree from ℒ𝑀𝐴 or ℒ𝑀𝐵 will have its corresponding degree in the base ℒ𝑀 . For example, having the bases ℒ7 and ℒ5 , the common base size is 𝑀 = 𝐿𝐶𝑀(7 − 1, 5 − 1) + 1 = 13. Figure 3 illustrates visually this example. We can see that the degree 𝜏1 of the base ℒ5 is equivalent to the degree 𝜏3 in the base ℒ13 , and 𝜏4 of the base ℒ7 is equivalent to 𝜏8 in the base ℒ13 , and so on. The multi-sets 𝑣𝛼 𝐴, 𝑣𝛽 𝐵, and 𝑣𝛾 𝐴 must then be translated to the common base ℒ𝑀 . To do that, we propose to apply a linguistic modifier to the considered multi-sets. To avoid information loss, this linguistic modifier must conserve their proportions. This will guarantee that the notion expressed by a multi-set will not be modified. So the mode of this linguistic modifier must be central (see section 2.2.). Moreover, the 148
𝜏𝑖 ′ = 𝜏𝑖.𝜌 ℒ𝑀′ = ℒ𝑀.𝜌−𝜌+1
(14)
The modifier 𝐷𝐶 changes the degree 𝜏𝑖 to 𝜏𝑖.𝜌 , and changes the base size to 𝑀′ = 𝑀.𝜌 − 𝜌 + 1 = 𝜌.(𝑀 − 1)+1. So it acts as a zoom in by multiplying the degree and the base size by 𝜌 (with subtracting it by 1 for the base size). It remains to determine the radius 𝜌 of the modifier 𝐷𝐶 for the interfacing, which represents its intensity. Suppose that we would translate the proposition 𝑣𝛼 𝐴 into the base ℒ𝑀𝐴 . The operator 𝐷𝐶 transforms a base to its multiple. Moreover, the intended base ℒ𝑀 is a multiple of ℒ𝑀𝐴 since it is its LCM. So the 𝑀−1
ratio of 𝐷𝐶 must be . This ratio will also be used 𝑀𝐴 −1 for the conversion of 𝑣𝛾 𝐴 since they have the same base. In the same way, the ratio of the conversion for 𝑀−1 the proposition 𝑣𝛽 𝐵 will be . If we return to the 𝑀𝐵 −1 example of figure 3, to translate the degree 𝜏1 of the base ℒ5 to the base ℒ13 , the used modifier is 𝐷𝐶 13−1 = 5−1
𝐷𝐶3 . This will give the new multi-set: 𝐷𝐶3 (𝜏1 , ℒ5 ) = (𝜏1∗3 , ℒ5∗3−3+1 ) = (𝜏3 , ℒ13 ). 4.2.
Symbolic inference
Having all the multi-sets on the same base ℒ𝑀 , we can apply our approximate reasoning (10). So, the GMP (12) on the base ℒ𝑀 becomes the following: If 𝑋 is 𝐷𝐶 𝑀−1 (𝑣𝛼 𝐴) then 𝑌 is 𝐷𝐶 𝑀−1 (𝑣𝛽 𝐵) 𝑀𝐴 −1
𝑀𝐵 −1
𝑋 is 𝐷𝐶 𝑀−1 (𝑣𝛾 𝐴) 𝑀𝐴 −1
𝑌 is 𝑣𝜃 𝐵 (15) The degree 𝜏𝜃 of the inference conclusion is determined using The GMP (10). Thus, it is obtained by the following formula: 𝜏𝜃 = 𝑇(ℐ(𝜏𝛼. 𝑀−1 , 𝜏𝛾. 𝑀−1 ), 𝜏𝛽. 𝑀−1 ) 𝑀𝐴 −1
𝑀𝐴 −1
𝑀𝐵 −1
(16)
Journal of Automation, Mobile Robotics and Intelligent Systems
4.3.
Re-interfacing of the result
The obtained inference conclusion in the previous stage is on the common base ℒ𝑀 . It is then necessary to give it back to the base ℒ𝑀𝐵 . For that, we propose to also do that with a linguistic modifier. This modifier should conserve the proportion of the multi-set to avoid information loss, so it must be central. Moreover, it will change the base size 𝑀 of the intermediate result to 𝑀𝐵 . More precisely, it will decrease the base size, since the base ℒ𝑀 is a LCM of the targeted base ℒ𝑀𝐵 . For that, the nature of the used linguistic modifier should be erosion. Consequently, we conclude that the used modifier for the re-interfacing is 𝐸𝐶 (Eroded Central) (see table 1). Its definition is the following:
𝐸𝐶𝜌 = �
𝜏𝑖 ′ = 𝜏𝑚𝑎𝑥(⌊ 𝑖 ⌋,1) 𝜌
(17)
ℒ𝑀′ = ℒ𝑚𝑎𝑥(⌊ 𝑀 ⌋+1,2) 𝜌
Here, with this modifier, the degree and the base size are divided by 𝜌. Thus, it acts like a zoom out of the multi-sets. The obtained degree and base size after the modification are the dividers of the original ones. The problem that then arises is to determine the radius of the modifier 𝐸𝐶. The radius represents its intensity. It must lead the intermediate conclusion, which is in the base ℒ𝑀 , to its original base, which is ℒ𝑀𝐵 . As 𝑀−1 is the LCM of 𝑀𝐵 −1, 𝑀𝐵 −1 is a divider of
𝜃 𝑀𝐵 − 1 𝜆 = 𝑚𝑎𝑥 �� 𝑀−1 � , 1� = 𝑚𝑎𝑥 ��𝜃. � , 1� 𝑀−1 𝑀𝐵 −1
(18) For example, if the intermediate result is the degree 𝜏10 in the base ℒ13 (see fig. 3), its re-interfacing 7−1 � , 1� = 5. So the to the base ℒ7 gives 𝑚𝑎𝑥 ��10 × 13−1 deduced degree is 𝜏5 in the base ℒ7 . 4.4.
If skin is more-or-less pale then blood flow is enough low (𝜏𝛼 = 𝜏5 , 𝜏𝛽 = 𝜏4 ) Skin is moderately pale (𝜏𝛾 = 𝜏4 ) blood flow is moderately low (𝜏𝜆 = 𝜏3 ) We calculate first the base size 𝑀: 𝑀 = 𝐿𝐶𝑀(𝑀𝐴 − 1, 𝑀𝐵 − 1) + 1 = 𝐿𝐶𝑀(9 − 1, 7 − 1) + 1 = 25 The intermediate conclusion 𝜏𝜃 is determined as follows: 𝜏𝜃 = 𝑇𝐿 �ℐ𝐿 �𝜏𝛼. 𝑀−1 , 𝜏𝛾. 𝑀−1 � , 𝜏𝛽. 𝑀−1 � 𝑀𝐴 −1
𝑀𝐵 −1
9−1
9−1
7−1
= 𝑇𝐿 (ℐ𝐿 (𝜏15 , 𝜏12 ) , 𝜏16 ) = 𝑇𝐿 (𝜏21 , 𝜏16 ) = 𝜏13 𝑀 −1 Then we deduce that: 𝜆 = max ��𝜃. 𝐵 � , 1� = 𝑀−1
7−1
max ��13. � , 1� = 3 25−1 This corresponds to the linguistic degree moderately in the scale ℒ7 . If we try the same method with another observation, we obtain: If skin is more-or-less pale then blood flow is enough low (𝜏𝛼 = 𝜏5 , 𝜏𝛽 = 𝜏4 ) Skin is very pale (𝜏𝛾 = 𝜏7 ) blood flow is enough low (𝜏𝜆 = 𝜏4 ) The degree 𝜏𝜃 is determined as follows:
We summarize in the proposition 1 the procedure to determine the inference conclusion from the inference schema (12) where the multi-sets 𝐴 and 𝐵 have different bases. Proposition 1. Given the multi-sets 𝐴 and 𝐵 of the bases ℒ𝑀𝐴 and ℒ𝑀𝐵 respectively. With a rule and an observation of the form: - “If 𝑋 is 𝑣𝛼 𝐴 then 𝑌 is 𝑣𝛽 𝐵” - “𝑋 is 𝑣𝛾 𝐴” we can conclude that “𝑌 is 𝑣𝜆 𝐵” with: 𝜆 = 𝑀 −1 max ��𝜃. 𝐵 � , 1�, 𝑀−1
where 𝜏𝜃 = 𝑇 �ℐ �𝜏𝛼. 𝑀−1 , 𝜏𝛾. 𝑀−1 � , 𝜏𝛽. 𝑀−1 � 𝑀𝐴 −1
𝑀𝐴 −1
= 𝑇𝐿 �ℐ𝐿 �𝜏5. 25−1 , 𝜏4. 25−1 � , 𝜏4. 25−1 �
Summary form of the method
𝑀𝐴 −1
2026
ℒ9 ={not-at-all, very-little, little, somewhat, moderately, more or less, enough, very, completely} ℒ7 ={not-at-all, very-little, little, moderately, enough, very, completely} The multi-sets pale and blood flow are associated to the bases ℒ9 and ℒ7 respectively. With the following data, we obtain this result:
𝑀−1
𝑀−1. So the ratio 𝜌 of 𝐸𝐶 must be . Consequently, 𝑀𝐵 −1 the degree 𝜏𝜆 of the final conclusion is obtained by the following formula:
N∘ 3
VOLUME 20,
𝑀𝐵 −1
and 𝑀 = 𝐿𝐶𝑀(𝑀𝐴 − 1, 𝑀𝐵 − 1) + 1. Example 2. Consider the bases ℒ9 and ℒ7 with the following scales:
𝜏𝜃
=
�ℐ𝐿 �𝜏5. 25−1 , 𝜏7. 25−1 � , 𝜏4. 25−1 � 9−1
9−1
=
7−1
𝑇𝐿 (ℐ𝐿 (𝜏15 , 𝜏21 ) , 𝜏16 ) = 𝑇𝐿 (𝜏24 , 𝜏16 ) = 𝜏16 We conclude that the rank degree of the final result is: 𝑀 −1 7−1 𝜆 = max ��𝜃. 𝐵 � , 1� = max ��16. � , 1� = 4 𝑀−1 25−1 This corresponds to the linguistic degree enough in the scale ℒ7 .
5.
Consideration of complex rules
The approximate reasoning that we proposed in the previous section is a simplified form of what we can find in rule-based systems. It handles a simple rule whose antecedent and consequence are composed of a unique proposition each. In practice, rule-based systems contain rules whose premise and conclusion can be a conjunction or a disjunction of propositions, or combinations of them. We discuss in this section these 149
Journal of Automation, Mobile Robotics and Intelligent Systems
cases, and we specify how to infer with them. We will suppose here that all the manipulated multi-sets have the same base. The multi-bases problem is examined in the next section. 5.1.
Conjunctive rules
A conjunctive rule has a premise that is a conjunction of propositions. To infer with such a rule, the observation must also be a conjunction of propositions whose predicates are the same as those of the premise. Thus, the reasoning schema becomes the following: If 𝑋1 is 𝑣𝛼1 𝐴1 and …and 𝑋𝑛 is 𝑣𝛼𝑛 𝐴𝑛 then 𝑌 is 𝑣𝛽 𝐵 𝑋1 is 𝑣𝛾1 𝐴1 and …and 𝑋𝑛 is 𝑣𝛾𝑛 𝐴𝑛 𝑌 is 𝑣𝜆 𝐵
(19)
The problem is how to evaluate the transformation between the premise and the observation, and how to translate it in order to calculate the degree 𝜏𝜆 of the inference conclusion. In other words, how can the term and be mathematically modeled to calculate the degree 𝜏𝜆 ? In fuzzy logic, the term and is generally formalized by a fuzzy connective [34–37]. We say that a binary operation 𝑓 is a fuzzy connective if it verifies the following axioms: - 𝑓 gives the same results of the classical logic connectives ∨ and ∧ when the degrees are precise and certain; - 𝑓 is commutative, monotone and associative. These properties are verified by T-norms and Tconorms. For simplicity reasons, fuzzy community generally uses the operators 𝑚𝑖𝑛 and 𝑚𝑎𝑥 as fuzzy connectives to replace, respectively, ∨ and ∧ in fuzzy context. We choose in this work to consider 𝑚𝑖𝑛 and 𝑚𝑎𝑥 as symbolic multi-valued connectives. We note them respectively by ∨ and ∧ as for classical logic. So, for the case of conjunctive rules, the principle of our solution is to aggregate by the connective ∧ the different transformations between every couple of propositions from the premise and the observation that have the same predicate. Proposition 2. Given a rule and an observation of the form: - “If 𝑋1 is 𝑣𝛼1 𝐴1 and …and 𝑋𝑛 is 𝑣𝛼𝑛 𝐴𝑛 then 𝑌 is 𝑣𝛽 𝐵” - “𝑋1 is 𝑣𝛾1 𝐴1 and …and 𝑋𝑛 is 𝑣𝛾𝑛 𝐴𝑛 ” where 𝐴1 … 𝐴𝑛 and 𝐵 are multi-sets on the base ℒ𝑀 . We can deduce that “Y is 𝑣𝜆 𝐵” with: 𝑛
𝜏𝜆 = 𝑇(𝜏𝛿 , 𝜏𝛽 ) where 𝜏𝛿 = � ℐ(𝜏𝛼𝑖 , 𝜏𝛾𝑖 )
VOLUME 20,
N∘ 3
2026
If headache is moderately intense and dizziness is little significant then malaria is enough severe (𝜏𝛼1 = 𝜏3 , 𝜏𝛼2 = 𝜏2 , 𝜏𝛽 = 𝜏4 ) Headache is very-little intense and dizziness is very-little significant (𝜏𝛾1 = 𝜏1 , 𝛾2 = 𝜏1 ) Malaria is little severe (𝜏𝜆 = 𝜏2 )
We use the Łukasiewicz T-norm 𝑇𝐿 (4) and implication operator ℐ𝐿 (6) to determine the conclusion by the proposition 2: 𝜏𝛿 = ℐ𝐿 (𝜏𝛼1 , 𝜏𝛾1 ) ∧ ℐ𝐿 (𝜏𝛼2 , 𝜏𝛾2 ) = ℐ𝐿 (𝜏3 , 𝜏1 ) ∧ ℐ𝐿 (𝜏2 , 𝜏1 ) = 𝜏4 Which allows determining the following: 𝜏𝜆 = 𝑇𝐿 (𝜏𝛿 , 𝜏𝛽 ) = 𝑇𝐿 (𝜏4 , 𝜏4 ) = 𝜏𝑚𝑎𝑥(4+4−7+1,0) = 𝜏2 The symbolic degree 𝜏2 corresponds to the term “little” from the base ℒ7 , so the conclusion is: “Malaria is little severe”. And if we consider the following data we can deduce: If headache is moderately intense and dizziness is little significant then malaria is enough severe (𝜏𝛼1 = 𝜏3 , 𝜏𝛼2 = 𝜏2 , 𝜏𝛽 = 𝜏4 ) Headache is very intense and dizziness is moderately significant (𝜏𝛾1 = 𝜏5 , 𝛾2 = 𝜏3 ) Malaria is enough severe (𝜏𝜆 = 𝜏4 )
The degree of the conclusion is determined as follows: 𝜏𝛿 = ℐ𝐿 (𝜏𝛼1 , 𝜏𝛾1 ) ∧ ℐ𝐿 (𝜏𝛼2 , 𝜏𝛾2 ) = ℐ𝐿 (𝜏3 , 𝜏5 ) ∧ ℐ𝐿 (𝜏2 , 𝜏3 ) = 𝜏6 The resulted degree is: 𝜏𝜆 = 𝑇𝐿 (𝜏𝛿 , 𝜏𝛽 ) = 𝑇𝐿 (𝜏6 , 𝜏4 ) = 𝜏4 The symbolic degree 𝜏4 corresponds to the term “enough”, so the conclusion is: “Malaria is enough severe”. As we can see in the example, our approximate reasoning with the conjunctive rule has the same behavior as the simple rule. Indeed, we remark that if the propositions of the observation are a weakening of the propositions of the rule premise, then the inference conclusion is a weakening of the rule conclusion. Otherwise, the inference conclusion obtains the same value as the rule conclusion. This corresponds to our expectations in order to verify the criterion II-2 of the axiomatics (2). 5.2.
Disjunctive rules
A disjunctive rule is a rule whose premise is a disjunction of propositions. The disjunction is formulated by the term or between propositions. The Generalized Modus Ponens with such a rule is the following: If 𝑋1 is 𝑣𝛼1 𝐴1 or …or 𝑋𝑛 is 𝑣𝛼𝑛 𝐴𝑛 then 𝑌 is 𝑣𝛽 𝐵 𝑋1 is 𝑣𝛾1 𝐴1 or …or 𝑋𝑛 is 𝑣𝛾𝑛 𝐴𝑛 𝑌 is 𝑣𝜆 𝐵
(20)
𝑖=1
Example 3. Having the scale base ℒ7 = {not-at-all, very-little, little, moderately, enough, very, completely}, the following rule and fact allow deducing that: 150
As for the conjunction case, the aim is to aggregate the transformation between the propositions of the premise and of the conclusion. We consider for that the connective ∨.
Journal of Automation, Mobile Robotics and Intelligent Systems
Proposition 3. Given a rule and an observation of the form: - “If 𝑋1 is 𝑣𝛼1 𝐴1 or …or 𝑋𝑛 is 𝑣𝛼𝑛 𝐴𝑛 then 𝑌 is 𝑣𝛽 𝐵” - “𝑋1 is 𝑣𝛾1 𝐴1 or …or 𝑋𝑛 is 𝑣𝛾𝑛 𝐴𝑛 ” where 𝐴1 … 𝐴𝑛 and 𝐵 are multi-sets on the base ℒ𝑀 . We can deduce that “Y is 𝑣𝜆 𝐵” with:
VOLUME 20,
If 𝑋1 is 𝑣𝛼1 𝐴1 and …and 𝑋𝑛 is 𝑣𝛼𝑛 𝐴𝑛 then 𝑌 is 𝑣𝛽 𝐵 𝑋1 is 𝑣𝛾1 𝐴1 and …and 𝑋𝑛 is 𝑣𝛾𝑛 𝐴𝑛 𝑌 is 𝑣𝜆 𝐵
𝜏𝜆 = 𝑇(𝜏𝛿 , 𝜏𝛽 ) where 𝜏𝛿 = � ℐ(𝜏𝛼𝑖 , 𝜏𝛾𝑖 ) Example 4. Having the list of truth-degrees ℒ7 = {not-at-all, very-little, little, moderately, enough, very, completely}, with the following knowledge we conclude: If nausea is moderately strong or jaundice is little intense then malaria is moderately severe (𝜏𝛼1 = 𝜏3 , 𝜏𝛼2 = 𝜏4 , 𝜏𝛽 = 𝜏3 ) Nausea is very strong and jaundice is enough intense (𝜏𝛾1 = 𝜏5 , 𝜏𝛾2 = 𝜏4 ) Malaria is moderately severe (𝜏𝜆 = 𝜏3 )
One obtains the conclusion with proposition 3 as follows: 𝜏𝛿 = ℐ𝐿 (𝜏𝛼1 , 𝜏𝛾1 ) ∨ ℐ𝐿 (𝜏𝛼2 , 𝜏𝛾2 ) = ℐ𝐿 (𝜏3 , 𝜏5 ) ∨ ℐ𝐿 (𝜏4 , 𝜏4 ) = 𝜏6 We can then determine the inferred degree: 𝜏𝜆 = 𝑇𝐿 (𝜏𝛿 , 𝜏𝛽 ) = 𝑇𝐿 (𝜏6 , 𝜏3 ) = 𝜏𝑚𝑎𝑥(6+3−7+1,0) = 𝜏3 The determined degree 𝜏3 represents the term “moderately” in the scale base ℒ7 . So, the deducted conclusion is: “Malaria is moderately severe”.
6. Reasoning with multi-bases knowledge and complex rules After treating the presence of multi-bases multisets and complex rules separately, we consider in this section the presence of multi-bases multi-sets and complex rules in the same inference model. We consider in this section the case where the antecedent is a conjunction of propositions, which is the more used form. When a rule is disjunctive (the premise is with OR connective), the same procedure and formulas can be adopted. Just must replace the connective ∧ by ∨, and it will be modeled by the 𝑚𝑎𝑥 operator instead of the 𝑚𝑖𝑛 operator.
2026
For the conjunctive rule, the observation must also be composed by a conjunction of propositions whose predicates are the same as that of the antecedent. Approximate reasoning in this case is formalized by the following schema:
𝑛
𝑖=1
N∘ 3
(21)
with 𝐴1 , …, 𝐴𝑛 and 𝐵 are multi-sets. In this situation, these multi-sets have different bases, which we denote by ℒ𝐴1 , …, ℒ𝐴𝑛 and ℒ𝐵 respectively. In order to infer with these knowledge, we propose to follow the same principle as the reasoning with multi-bases multi-sets described in Section 4., while aggregating the imprecision of the composed rule premise and the composed observation as explained in Section 5.. Consequently, tree stages are necessary, which are: interfacing of manipulated multi-sets; inference; and re-interfacing of the result. 6.1.
Interfacing of manipulated multi-sets
The interfacing of the multi-sets consists of translating them to a common base ℒ𝑀 . The size 𝑀 of this base is obtained by the Lower Common Multiple of all the base sizes that appear in the GMP (21), which are 𝑀𝐴1 , …, 𝑀𝐴𝑛 and 𝑀𝐵 . Thus, it is determined by the following formula:
𝑀 = 𝐿𝐶𝑀(𝑀𝐴1 − 1, … , 𝑀𝐴𝑛 − 1, 𝑀𝐵 − 1) + 1 (22) The conversion of the considered multi-sets to the base ℒ𝑀 is made by the modifier 𝐷𝐶 (Dilating Central in table 1). Indeed, this modifier will dilate their bases in order to reach the base size 𝑀. Every multi-set will have its own radius 𝜌 according to its base size. The radius 𝜌 of the modification will be, for every multi-set, 𝑀 − 1 divided by the base size of the multi-set minus one. So for a multi-set 𝐴𝑖 , 𝑖 ∈ [1, 𝑛], the value of the 𝑀−1 𝑀−1 radius is , and for 𝐵 it is . 𝑀𝐴𝑖 −1
6.2.
𝑀𝐵 −1
Symbolic inference
Once all the multi-sets are on the same base, approximate reasoning can be executed. This will determine a new assertion from the following schema:
If 𝑋1 is 𝐷𝐶 𝑀−1 (𝑣𝛼1 𝐴1 ) and … and 𝑋𝑛 is 𝐷𝐶 𝑀−1 (𝑣𝛼𝑛 𝐴𝑛 ) then 𝑌 is 𝐷𝐶 𝑀−1 (𝑣𝛽 𝐵) 𝑀𝐴 −1 1
𝑀𝐴 −1 𝑛
𝑀𝐵 −1
𝑋1 is 𝐷𝐶 𝑀−1 (𝑣𝛾1 𝐴1 ) and … and 𝑋𝑛 is 𝐷𝐶 𝑀−1 (𝑣𝛾𝑛 𝐴𝑛 ) 𝑀𝐴 −1 1
Here, the multi-set 𝐵′ is the equivalent of the multiset 𝐵 on the base ℒ𝑀 . The problem encountered here is the consideration of complex rules in the inference process. The challenge is how to aggregate the imprecision of the whole antecedent and the observation to consider it in the determination of the inference
𝑀𝐴 −1 𝑛
(23) 𝑌 is 𝑣𝜃 𝐵′
conclusion. We already proposed in Section 5. a solution to reason with complex rule in the case of monobase (see proposition 2). This reasoning model can be used here since all the manipulated multi-sets have the same base. Consequently, the resulted degree 𝜏𝜃 of the inference after interfacing is obtained as follows:
151
Journal of Automation, Mobile Robotics and Intelligent Systems
𝜏𝜃 = 𝑇(𝜏𝛿 , 𝜏𝛽. 𝑀−1 ) with: 𝑀𝐵 −1
𝑛
(24)
𝜏𝛿 = � ℐ(𝜏𝛼 . 𝑀−1 , 𝜏𝛾 . 𝑀−1 ) 𝑖=1
6.3.
𝑖 𝑀
𝐴𝑖 −1
𝑖 𝑀 −1 𝐴𝑖
Re-interfacing of the result
The obtained result 𝑣𝜃 𝐵′ in the previous step is on the base ℒ𝑀 . It is then necessary to translate it under the conclusion base ℒ𝑀𝐵 , as it is imposed by the expert. In order to do that, we propose to use the modifier 𝐸𝐶 (Eroded Central) as for the simple inference schema described in Section 4.. This modifier will erode the base of the multi-set 𝑣𝜃 𝐵′ by dividing its size 𝑀 until reaching the desired base size 𝑀𝐵 . Consequently, and using the definition of 𝐸𝐶 (see table 1), the degree of the inference conclusion is obtained as follows: (25)
2026
If skin is moderately pale and limbs are little numb, then blood flow is enough low (𝜏𝛼1 = 𝜏3 , 𝜏𝛼2 = 𝜏2 , 𝜏𝛽 = 𝜏3 ) Skin is little pale and limbs are very-little numb (𝜏𝛾1 = 𝜏2 , 𝛾2 = 𝜏1 ) blood flow is moderately low (𝜏𝜆 = 𝜏2 ) The conclusion degree 𝜏𝜆 of the inference with these data are obtained by calculating first the base size: 𝑀 = 𝐿𝐶𝑀(𝑀𝐴1 −1, 𝑀𝐴2 −1, 𝑀𝐵 −1)+1 = 𝐿𝐶𝑀(9− 1, 7 − 1, 5 − 1) + 1 = 25. The intermediate conclusion 𝜏𝜃 is determined as follows: 𝑛
𝜏𝜃 = 𝑇𝐿 � ⋀ ℐ𝐿 (𝜏𝛼 . 𝑀−1 , 𝜏𝛾 . 𝑀−1 ), 𝜏𝛽. 𝑀−1 � 𝑖=1
𝑖 𝑀 −1 𝐴𝑖
𝑖 𝑀 −1 𝐴𝑖
𝑀𝐵 −1
= 𝑇𝐿 �ℐ𝐿 (𝜏4. 25−1 , 𝜏2. 25−1 ) ∧ ℐ𝐿 (𝜏2. 25−1 , 𝜏1. 25−1 ), 𝜏3. 25−1 � 9−1
𝑀𝐵 − 1 𝜆 = 𝑚𝑎𝑥 ��𝜃. � , 1� 𝑀−1
N∘ 3
VOLUME 20,
9−1
7−1
7−1
5−1
= 𝑇𝐿 (ℐ𝐿 (𝜏12 , 𝜏6 ) ∧ ℐ𝐿 (𝜏8 , 𝜏4 ), 𝜏18 ) = 𝑇𝐿 (𝜏18 ∧ 𝜏20 , 𝜏18 )
6.4.
Summary of the method
We present below a proposition that summarizes the procedure to determine a conclusion in the presence of a conjunctive rule and multi-bases multi-sets. Proposition 4. Given a conjunctive rule and an observation of the form: - “If 𝑋1 is 𝑣𝛼1 𝐴1 and …and 𝑋𝑛 is 𝑣𝛼𝑛 𝐴𝑛 then 𝑌 is 𝑣𝛽 𝐵” - “𝑋1 is 𝑣𝛾1 𝐴1 and …and 𝑋𝑛 is 𝑣𝛾𝑛 𝐴𝑛 ” where 𝐴1 , …, 𝐴𝑛 and 𝐵 are multi-sets of the bases ℒ𝑀𝐴1 , …, ℒ𝑀𝐴𝑛 and ℒ𝑀𝐵 . We can conclude that “𝑌 is 𝑣𝜆 𝐵” with: 𝜆 = max ��𝜃.
𝑀𝐵 − 1 � , 1� 𝑀−1
𝑛
= 𝜏12 So, the degree 𝜏𝜆 of the final conclusion is: 5−1
𝑀 −1
� , 1� = 2. 𝜆 = max ��𝜃. 𝐵 � , 1� = max ��12. 𝑀−1 25−1 This corresponds to the linguistic degree moderately in the scale ℒ5 . If the input is the following, we obtain:
If skin is moderately pale and limbs are little numb, then blood flow is enough low (𝜏𝛼1 = 𝜏4 , 𝜏𝛼2 = 𝜏2 , 𝜏𝛽 = 𝜏3 ) Skin is completely pale and limbs are completely numb (𝜏𝛾1 = 𝜏8 , 𝜏𝛾2 = 𝜏6 ) blood flow is enough low (𝜏𝜆 = 𝜏3 )
where 𝜏𝜃 = 𝑇 � ⋀ ℐ(𝜏𝛼 . 𝑀−1 , 𝜏𝛾 . 𝑀−1 ), 𝜏𝛽. 𝑀−1 � 𝑖=1
𝑖 𝑀 −1 𝐴𝑖
𝑖 𝑀
𝐴𝑖 −1
𝑀𝐵 −1
and 𝑀 = 𝐿𝐶𝑀(𝑀𝐴1 − 1, … , 𝑀𝐴𝑛 − 1, 𝑀𝐵 − 1) + 1
The intermediate conclusion 𝜏𝜃 is determined as follows:
In the same way, when the rule is disjunctive instead of conjunctive, the conclusion of the inference is obtained by following the proposition 4 with replacing ∨ by ∧.
𝜏𝜃 = 𝑇(ℐ(𝜏𝛼 . 𝑀−1 , 𝜏𝛾 . 𝑀−1 ) ∧ ℐ(𝜏𝛼 . 𝑀−1 , 𝜏𝛾 . 𝑀−1 ), 1 𝑀 −1 𝐴1
1 𝑀
𝐴1 −1
2 𝑀
𝐴2 −1
2 𝑀 −1 𝐴2
𝜏𝛽. 𝑀−1 ) 𝑀𝐵 −1
Example 5. Having the lists of truth-degrees: ℒ9 ={not-at-all, very-little, little, somewhat, moderately, more or less, enough, very, completely}, ℒ7 ={not-at-all, very-little, little, moderately, enough, very, completely}, ℒ5 ={not-at-all, little, moderately, enough, completely}, and the multi-sets pale, numb and low in the basis ℒ9 , ℒ7 and ℒ5 respectively. The following rule and observation give as result: 152
= 𝑇 (ℐ(𝜏12 , 𝜏24 ) ∧ ℐ(𝜏8 , 𝜏24 ), 𝜏18 ) = 𝑇 (𝜏24 ∧ 𝜏24 , 𝜏18 ) = 𝜏18 So, the degree 𝜏𝜆 of the final conclusion is calculated as follows: 𝑀 −1
5−1
𝜆 = max ��𝜃. 𝐵 � , 1� = max ��18. � , 1� = 3 𝑀−1 25−1 This corresponds to the linguistic degree enough in the scale ℒ5 .
Journal of Automation, Mobile Robotics and Intelligent Systems
7. Discussion The proposition 4 allows inferring with both multibases multi-sets and conjunctive rules. Thus, the determination of the conclusion by the proposition 4 generalizes that of the proposition 1 when the number 𝑛 of propositions in the rule antecedent is equal to 1. Also, it generalizes that of the proposition 2 when the base sizes 𝑀𝐴𝑖 (𝑖 ∈ [1 … 𝑛]) and 𝑀𝐵 are equal. Moreover, the determination of the conclusion by the proposition 1 generalizes that of the Generalized Modus Ponens (10) which considers a simple rule when 𝐴 and 𝐵 have the same base. We deduce that the proposition 4 gives a general form for our cautious approximate reasoning. In a rule-based system, considering data imprecision can be performed by integrating this approximate reasoning into its inference engine. The purpose of the inference engine in rule-based systems is to simulate a human deduction based on a set of facts and a set of rules. Its principle is to enchain cycles, each comprising approximate reasoning phase in order to consider rule chaining. A multi-set can be present in more than one rule consequence. For that, it is necessary to make combination of conclusions. This can be made by the 𝑚𝑎𝑥 operator, as in fuzzy rule-based systems. As we mentioned before, the particularity of this approximate reasoning compared to other approaches of approximate reasoning is that it gives a conclusion equal to the rule consequence when the observation is a reinforcement of the rule antecedent. It is then useful for specific rules whose antecedent and consequence do not have a big causality and do not then require gradual reasoning. Its use must be decided by the domain expert. Indeed, he knows the characteristics of the knowledge in the antecedent and the consequence of every rule, as well as their causality relation. This leads to realize that in the same rule-based system, we may have to define rules which do not necessarily have the same behavior. For that, we propose to associate with each defined rule in the rule-based system a parameter to specify the behavior of the rule in order to know which approximate reasoning will be used for this rule: gradual reasoning (type 1) or cautious reasoning (type 2).
8. Conclusion In this paper, we focused on the approximate reasoning based on implication operator that is proposed in the multi-valued context [21]. We gave proofs that this model checks the axiomatics of approximate reasoning defined in the literature. This approximate reasoning nevertheless considers multi-sets having the same base. Moreover, it infers only with a simple fact and a simple rule, whose premise and conclusion consist each of a single proposition. For that, we considered other types of knowledge in the inference process. We improved this approximate reasoning to handle multi-bases multi-sets. Moreover, we treated the case of complex rules, where the premise is a conjunction or a disjunction of propositions. We finally
VOLUME 20,
N∘ 3
2026
proposed a general form of cautious approximate reasoning that generalizes all the treated cases in the paper. A perspective of this work is to build a multi-valued rule-based system that integrates our two types of approximate reasoning.
AUTHOR Saoussen Bel Hadj Kacem∗ – University of Carthage, Tunisia, e-mail: saoussen.belhadjkacem@ensiuma.tn, https://orcid.org/0000-0001-9772-4969. ∗
Corresponding author
References [1] L. A. Zadeh, “The concept of a linguistic variable and its application to approximate reasoning i - ii - iii”, Information Sciences, 1975, 8:199–249, 8:301–357, 9:43–80. [2] L. A. Zadeh, “Fuzzy sets”, Information and Control, vol. 8, no. 3, 1965, 338–353. [3] S. Fukami, M. Mizumoto, and K. Tanaka, “Some considerations of fuzzy conditional inference”, Fuzzy Sets and Systems, vol. 4, no. 3, 1980, 243–273. [4] S. B. H. Kacem, A. Borgi, and K. Ghé dira, “Generalized modus ponens based on linguistic modifiers in a symbolic multi-valued framework”. In: Proceeding of the 38th IEEE International Symposium on Multiple-Valued Logic, Dallas, USA, 2008, 150–155. [5] F. Lafont and J.-F. Balmat, “Optimized fuzzy control of a greenhouse”, Fuzzy Sets and Systems, vol. 128, no. 1, 2002, 47 – 59. [6] L. Ljung, R. Palm, D. Driankov, B. Graham, H. Hellendoorn, A. Ollero, and M. Reinfrank, An Introduction to Fuzzy Control, Springer Berlin Heidelberg, 2013. [7] M. Maeda, Y. Maeda, and S. Murakami, “Fuzzy drive control of an autonomous mobile robot”, Fuzzy Sets and Systems, vol. 39, no. 2, 1991, 195 – 204. [8] R.-E. Precup and H. Hellendoorn, “A survey on industrial applications of fuzzy control”, Computers in Industry, vol. 62, no. 3, 2011, 213 – 226. [9] Y. Takahashi, T. Ishii, C. Todoroki, Y. Maeda, and T. Nakamura, “Fuzzy control for a kite-based tethered flying robot”, Journal of Advanced Computational Intelligence and Intelligent Informatics, vol. 19, no. 3, 2015, 349–358. [10] H. Akdag. Une approche logique du raisonnement incertain. PhD thesis, University of Paris VI, 1992. [11] H.-T. Chung and D. G. Schwartz, “A resolutionbased system for symbolic approximate reasoning”, Int. J. Approx. Reasoning, vol. 13, no. 3, 1995, 201–246. 153
Journal of Automation, Mobile Robotics and Intelligent Systems
[12] M. El-Sayed and D. Pacholczyk, “Towards a symbolic interpretation of approximate reasoning”, Electr. Notes Theor. Comput. Sci., vol. 82, no. 4, 2003, 1–12. [13] M. L. Ginsberg, “Multivalued logics: a uniform approach to reasoning in artificial intelligence”, Computational Intelligence, vol. 4, no. 3, 1988, 265–316. [14] M. D. Glas. “Knowledge representation in a fuzzy setting”. Technical Report 89–48, LAFORIA, University of Paris VI, 1989. [15] D. Pacholczyk. Contribution au traitement logicosymbolique de la connaissance. PhD thesis, University of Paris VI, 1992. [16] D. G. Schwartz, “A system for reasoning with imprecise linguistic information”, Int. J. Approx. Reasoning, vol. 5, no. 5, 1991, 463–488. [17] H. Akdag, M. D. Glas, and D. Pacholczyk, “A qualitative theory of uncertainty”, Fundam. Inform., vol. 17, no. 4, 1992, 333–362. [18] N. Cat Ho and W. Wechler, “Hedge algebras: An algebraic approach to structure of sets of linguistic truth values”, Fuzzy Sets and Systems, vol. 35, no. 3, 1990, 281–293. [19] V. H. Le, F. Liu, and D. K. Tran, “Fuzzy linguistic logic programming and its applications”, Theory and Practice of Logic Programming, vol. 9, no. 3, 2009, 309–341, 10.1017/S1471068409003779. [20] M. Ying, “A logic for approximate reasoning”, The Journal of Symbolic Logic, vol. 59, no. 3, 1994, 830–837. [21] A. Borgi, S. B. H. Kacem, and K. Ghé dira. “Approximate reasoning in a symbolic multi-valued framework”. In: Computer and Information Science [outstanding papers from IEEE/ACIS ICIS/IWEA 2008], volume 131 of Studies in Computational Intelligence, 203–217. Springer, 2008. [22] S. B. H. Kacem, A. Borgi, and M. Tagina, “On some properties of generalized symbolic modifiers and their role in symbolic approximate reasoning.”. In: International Conference on Intelligent Computing (ICIC), vol. 5755, 2009, 190–208. [23] S. B. H. Kacem, A. Borgi, and M. Tagina, “Extended symbolic approximate reasoning based on linguistic modifiers”, Knowledge and Information Systems, vol. 42, no. 3, 2015, 633–661. [24] S. B. H. Kacem, A. Borgi, and S. Othman, “A diagnosis aid system of autism in a multi-valued framework”. In: W. Scientific, ed., Uncertainty Modelling in Knowledge Engineering and Decision Making: Proceedings of the 12th International FLINS Conference (FLINS 2016), 2016, 405–411. 154
VOLUME 20,
N∘ 3
2026
[25] S. B. H. Kacem, A. Borgi, and S. Othman. “Dasautism: a rule-based system to diagnose autism with multi-valued logic”. In: Smart Systems for E-health, Advanced Information and Knowledge Processing, 183–200. Springer, 2021. [26] S. B. H. Kacem, A. Borgi, and M. Tagina, “Ramoli: A generic knowledge-based systems shell for symbolic data”. In: 2013 World Congress on Computer and Information Technology (WCCIT), 2013, 1–6. [27] F. Khoukhi. Approche logico-symbolique dans le traitement des connaissances incertaines et imprécises dans les systèmes à base de connaissances. PhD thesis, Université de Reims, France, 1996. [28] L. A. Zadeh, “A fuzzy-set-theoretic interpretation of linguistic hedges”, Journal of Cybernetics, vol. 2, no. 3, 1972, 4–34. [29] H. Akdag, N. Mellouli, and A. Borgi, “A symbolic approach of linguistic modifiers”. In: Information Processing and Management of Uncertainty in Knowledge-Based Systems, Madrid, 2000, 1713–1719. [30] H. Akdag, I. Truck, A. Borgi, and N. Mellouli, “Linguistic modifiers in a symbolic framework”, International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, vol. 9, no. Supplement, 2001, 49–61. [31] I. Truck, A. Borgi, and H. Akdag, “Generalized modifiers as an interval scale: Towards adaptive colorimetric alterations”. In: Advances in Artificial Intelligence - IBERAMIA 2002, 8th IberoAmerican Conference on AI, 2002, 111–120. [32] S. B. H. Kacem, “A new approximate reasoning for multi-bases symbolic data”. In: 2017 IEEE/ACS 14th International Conference on Computer Systems and Applications (AICCSA), 2017, 1450–1453, 10.1109/AICCSA.2017.16. [33] L. N. Childs, A concrete introduction to higher algebra, Springer, 2009. [34] J. Dombi, “A general class of fuzzy operators, the demorgan class of fuzzy operators and fuzziness measures induced by fuzzy operators”, Fuzzy Sets and Systems, vol. 8, no. 2, 1982, 149–163. [35] D. Dubois and H. Prade, “A review of fuzzy set aggregation connectives”, Fuzzy Sets and Systems, vol. 36, 1985, 85–121. [36] J. Fodor and M. Roubens, Fuzzy Preference Modelling and Multicriteria Decision Support, Kluwer Academic Publishers, 1994. [37] M. Grabisch, J.-L. Marichal, R. Mesiar, and E. Pap, “Aggregation functions: Construction methods, conjunctive, disjunctive and mixed classes”, Information Sciences, vol. 181, no. 1, 2011, 23–43.
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
REVOLUTIONIZING BIG DATA ASSESSMENT FOR HUMAN ACTIVITY RECOGNITION WITH METAPATH CONTEXT AND BI-DIRECTIONAL CASCADE NETWORKS Submitted: 5th November 2024; accepted: 20th May 2025
Praveen S. Banasode, Sunita Padmannavar DOI: 10.14313/jamris-2026-047 Abstract: Human activity recognition focuses on automatically detecting and classifying what individuals are doing, utilizing data gathered from sensors or video feeds. Challenges include interpreting complex activities, handling imbalanced data classes, ensuring privacy in videobased methods, and achieving efficiency with limited computational resources. To solve these issues, this research develops a novel model named the Metapath Contexted Convoluted Bi-directional Cascade Network (MCCBCN), which leverages meta paths to enhance accuracy and robustness in capturing contextual information for human activity recognition. By leveraging both convolutional and bi-directional cascade architectures, MCCBCN makes a significant leap in the field by effectively capturing temporal dependencies. With the addition of the Enhanced Football Team Training Optimization Algorithm (EFTTOA), training efficiency gets a boost, allowing the model to better understand the complex temporal and contextual nuances in human activity data. Additionally, utilizing Large-scale Synthetic Minority Over-Sampling Technique (LSMOST) for dataset expansion ensures a more comprehensive and balanced representation of minority classes, mitigating biases from imbalanced class distributions and improving model generalizability. The proposed method achieved impressive evaluation metrics with an accuracy of 99.50%, F1score of 99.9%, precision of 99.42%, and recall of 99.3%, outperforming existing techniques. Also, the proposed MCCBCN time complexity is 10.2 seconds. The proposed MCCBCN model effectively addresses key challenges in human activity recognition by enhancing contextual understanding, improving training efficiency, and ensuring robust performance across imbalanced datasets. Keywords: Human Activity Recognition, Enhanced Football Team Training Optimization Algorithm, MobileNetV2-Lite, Dig Data, Metapath Contexted Convoluted Bi-directional Cascade Network
1. Introduction Human Activity Recognition (HAR) offers valuable insights for numerous sectors, including health monitoring, assisted living, sports and fitness, and surveillance [1]. Video-based HAR has been successfully implemented across many domains. Additionally, recent advancements in sensor technology have made sensors indispensable for
analyzing human behavior in healthcare, fitness tracking, behavior investigation, assisted living, and therapy [2]. These systems stand central to pervasive computing, where a multitude of sensors watch the environment and conduct activities [3]. To enable HAR, on-body sensors, namely accelerometers, gyroscopes, and magnetometers, gather motion data that depict the physical movements of persons. These data are important for monitoring medical conditions, fitness, assisted living, and behavioral evaluation [4]. Among the different sensors, wearable sensors show the most interest owing to their omnipresent notion of seamless integration with life. Recognizing human activity is of utmost concern, as it logs behaviors of individuals, generating data that permits computing systems to monitor, analyze, and support daily activities [5]. Systems predominantly fall into two groups: video-based and sensor-based. However, privacy concerns, particularly related to camera installations in personal environments, have prompted the prevalence of sensor-based alternatives [6]. These systems leverage on-body or ambient sensors to precisely detect motion details or record activity patterns while upholding privacy standards. The Big Data Assessment for HAR leverages Deep Learning (DL) systems to examine huge quantities of data for recognizing and categorizing human activities. The advent of big data in HAR has led to the adoption of DL models such as Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) [7], facilitating automatic classification from raw sensor data. These approaches facilitate the automatic extraction and classification of attributes from original sensor data. Despite the authority, DL paradigms require big annotated datasets, often not generalizing over various environments due to sensor placements, activity types, and user behaviors. Data augmentation, transfer learning, and adaptive feature learning are considered a few remedies by researchers. Sensor-based HAR systems are a hot topic in health care, wherein they provide a platform for patient monitoring on a continuous and nonintrusive basis. Similarly, these systems measure fitness performance and prevent injuries in fitness applications [8], while, in assisted living, systems respond to caregiver alerts triggered by anomalies in the routines of elderly or disabled persons [9]. To analyze these problems, many hybrid methods in related computer vision problems have been
Open Access. © 2026 Praveen S. Banasode and Sunita Padmannavar, published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP. 4.0 License
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives
155
Journal of Automation, Mobile Robotics and Intelligent Systems
considered. The combination of robust Principal Component Analysis (PCA) with metaheuristic optimization algorithms aimed at feature extraction and noise reduction, for scenarios involving occlusion or incomplete data, kind of problem faced in HAR, with sensor readings either missing or corrupted [10]. Following this pattern, an architecture combining pretrained detection/meters association/ segmentation modules and lightweight classifiers like K-Nearest Neighbors (KNN) demonstrated an improvement in recognition performance while preserving computational efficiency [11]. Swarm intelligence-based metaheuristics such as Grey Wolf Optimization (GWO) are very useful for feature selection and dimensionality reduction, so systems attend only to the most informative sensor features and end up reducing overfitting [12]. Improvements in skin detection and regional feature localization certainly teach a lot about how to cleanly segment activity-relevant bits from raw input data [13]. Together, these ideas lead to a HAR framework that is more adaptive, interpretable, and resource-efficient to deal with real-world variations [14]. Therefore, to overcome these issues, a novel technique named MCCBCN is introduced, which includes enhanced contextual understanding through rich relational data, improved accuracy by capturing dependencies in both directions, and the ability to handle complex, sequential data more effectively. To enhance the performance of MCCBCN, ETTOA is introduced. These techniques collectively lead to more precise and reliable recognition of human activities, making them highly valuable for analyzing extensive and intricate datasets in HAR.
Contributions • This research introduces the novel MCCBCN for HAR, effectively capturing contextual information through metapaths to enhance accuracy and robustness. Additionally, this approach of MCCBCN, leveraging both convolutional and bi-directional cascade architectures, contributes significantly to the field. • The utilization of LSMOST for dataset expansion ensures a more comprehensive and balanced representation of minority classes. By preventing biases caused by class distribution imbalances, it enhances the strength of the model, while empowering it to generalize easily. • The integration of the MobileNetV2-Lite backbone with Graph Convolutional Networks (MGCN) enables the extraction of human-centric features, while multi-view feature computation captures diverse activity patterns. Selection criteria based on their parameters ensure the inclusion of the most relevant features, enhancing model performance and interpretability. • The incorporation of compression techniques enhances the FTTOA, named EFTTOA. This enhancement reduces the complexity and cost 156
VOLUME 20,
N∘ 3
2026
associated with integrating advanced pose estimation techniques. This algorithm has the ability to provide precise activity recognition while minimizing computational resources, making it efficient for applications. The structure of the work is delineated as follows: Section 2 offers a literature survey, presenting a summary of earlier research and findings in the field. Section 3 elucidates the proposed method, delineating the approaches utilized in the research. Section 4 illustrates the outcomes and discussion, analyzing the outcomes of the methodology and deliberating on their implications. Section 5 provides the conclusion.
2.
Literature Survey
This section delves into recent literature on HAR, expanding on prior research efforts. Numerous studies have emerged, and this discussion focuses on showcasing current developments in the field. Nguyen et al. [10] developed the Triple-feature Double-motion Network (TD-Net), an enhanced iteration of the Double-feature Double-motion Network (DD-Net) featuring an additional branch. This new branch incorporated standardized joint coordinates to enhance spatial information. Zhang et al. [11] incorporated the Discrete Wavelet Transform (DWT) method into the discrete transform system to enhance the expressive features of human actions. By investigating pre-trained dual-stream Convolutional Neural Networks (CNN)- Recurrent Neural Networks (CNN-RNN), learned features combined with handcrafted ones to address the severe demands within the domain. Challa et al. [12] developed a vigorous classification approach for HAR employing data from wearable sensors, employing a hybrid technique that combines CNN and Bidirectional Long Short-Term Memory (BiLSTM). This multibranch CNN-BiLSTM network is intended to extract features directly from raw sensor data with negligible preprocessing required. Luptakova et al. [13] developed the HAR transformer model, originally designed for repurposing for analyzing time-series motion signals. Leveraging its self-attention mechanism, the transformer captured individual dependencies among signal values within the time series. Khatun et al. [14] developed a hybrid DL that coupled CNN and LSTM (CNN-LSTM) and enhanced it with a self-attention algorithm to boost the model’s predictive abilities. In addition to the database, the approach was assessed using benchmark databases like MHEALTH and UCI-HAR to validate its comparative performance. Shanableh [15] developed a new method for extracting features for HAR in videos, utilizing video encoding principles like motion compensation and attribute variables based on coding. These features were then used with DL for model classification. Ahmad [16] suggested HAR utilizing deep CNN features. Subsequently, the critical features within deep representations were selected, and these selected
Journal of Automation, Mobile Robotics and Intelligent Systems
features were fed to learn temporal dynamics. Hassan et al. [17] presented a dynamic HAR technique based on a Deep BiLSTM model and feature extraction via pre-trained transfer learning. The approach initially employs CNN models, specifically MobileNetV2, to extract deep-level features from video frames. Liu et al. [18] developed a new device-free method called TransTM (Time-streaming Multiscale Transformer). This model harnesses the data-fitting capabilities of transformers to directly use raw RFID RSSI data as input, bypassing the need for pre-processing. Cob-Parro et al. [19] presented a method that was devised for detecting people and recognizing their actions in natural settings, tailored and optimized for real-time operation on edge devices. HAR was implemented using an RNN, specifically an LSTM. The LSTM received as input a lightweight, ad hoc feature vector extracted from the bounding box around every detected individual in video surveillance footage. Table 1 presents a comparative analysis of HAR against established methods. 2.1.
Problem statement
In recent years, researchers have increasingly employed DL-based frameworks for HAR. However, persistent issues within these frameworks highlight various limitations, necessitating the development of improved approaches. Previous methods have identified challenges like deficiencies in activity recognition, low accuracy, performance disparities across various datasets, higher time, and lower convergence. These shortcomings underscore the motivation for the MCCBCN methodology in HAR. The primary goal is to address these issues and enhance overall HAR performance by leveraging optimization techniques, aiming for a more robust and effective solution that produces precise and contextually relevant recognition results.
VOLUME 20,
3.1.
N∘ 3
2026
Dataset collection
Input images are sourced from four diverse datasets like JHMDB [20], UCF11 [21], UCF YouTube action [22], and HMDB51 [23], comprising 928 video sequences categorized into 21 action classes. It includes physically marked skeleton data extracted from RGB videos and is partitioned into random train/test splits. The UCF11 dataset, also known as the YouTube action dataset, includes 1,160 videos in 11 action classes and presents challenges due to diverse conditions in the clips. The UCF YouTube action dataset, with 11 action classes like horse riding and biking, is challenging due to variations in camera movement, viewpoint changes, diverse poses, intricate backgrounds, and was compiled by 25 different groups. The HMDB51 dataset contains 6,849 videos in 51 classes, each with at least 101 videos, sourced from movie clips and YouTube. 3.2.
Big data assessment using local linear synthetic minority over-sampling technique
After data collection, the LSMOST is applied to address class imbalance in the dataset. This involves identifying minority class instances that are underrepresented compared to the majority classes. LSMOST then generates synthetic samples specifically for these minority classes by including existing instances and their nearest neighbors in a locally linear manner. This process aims to raise the number of minority class samples, effectively balancing the class distribution in the dataset. By doing so, LSMOST ensures that all classes are adequately represented throughout the training phase of machine learning models, thereby mitigating the potential bias towards majority classes and improving the model’s overall performance and accuracy. This establishes a clear sequence in the process where the results from the big data using LSMOST are directed towards the feature extraction [24].
3. Proposed Methodology for Human Activity Recognition
3.3.
In this research, four human action datasets are used: JHMDB, UCF11, UCF YouTube Action, and HMDB51.These datasets underwent expansion using the LSMOST to generate synthetic samples for minority classes, creating extensive and balanced data to mitigate biases during model training. Features were extracted using a MobileNet v2-Lite backbone with GCN, followed by computing multi-view features from gradients in both horizontal and vertical directions, with features oriented vertically. Fig. 1 illustrates the block diagram of the introduced work. These extracted features were then combined and selected based on relative entropy, mutual information, and the SCC via a maximum probabilitybased threshold function. In the next phase, the selected features were fed into the MCCBCN to capture time-based, spatial, and behavioral relationships in video data streams. The model’s accuracy and robustness were further enhanced with the integration of the EFTTOA.
The collected dig data undergoes feature extraction using a MobileNetV2-lite backbone with Graph Convolutional Networks (MGCN). The MobileNetv2-Lite is a lightweight and efficient feature extractor for images or videos. The MobileNetv2-Lite produces small but rich features that capture important properties. These features are then fused with the MGCNs, which are capable of modeling relationships in graphstructured data. The proposed method rests on leveraging the efficiency of MobileNetv2-Lite as a feature extractor and the relational learning capacity of MGCNs to create a rich and expressive overall feature representation or embedding. Figure 2 displays the MobileNetV2 with the MGCN model. During this stage, MGCNs further process the features generated from MobileNetv2-Lite, supplementing the feature representation with contextual data important for downstream processing. Finally,
MobileNetV2-lite backbone with graph convolutional networks for feature extraction
157
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 1. Comparison of the existing work with HAR Ref
Methods
Pros
Cons
[10]
TD-Net
The method was effective in both dimensions, delivering strong performance on extensive datasets.
It failed to address the computational complexity associated with implementation, limiting its practical application in scenarios requiring immediate processing.
[11]
CNN-RNN
Absence of temporal domain diversity in terms of time scales.
[12]
CNNBiLSTM
[13]
HAR transformer
[14]
CNNLSTM
[15]
ViCoMoCo-DL
[16]
CNN-BiGRU
The DT model enhances the descriptive human action features by integrating the DWT technique. This approach improves the performance of action recognition by leveraging both engineered features and learned ones, thereby harnessing the advantages of both methodologies. The approach was intended to seize both short-range and extended dependencies within sequential data, ensuring a comprehensive understanding of temporal relationships across various time scales. Greater temporal spacing between features in the time series provided enhanced contextual comprehension across extended sequences. The recognition accuracy was improved by leveraging user similarity and weighting training data, allowing for a more refined learning process that accounted for individual user characteristics and preferences. This method effectively utilizes video encoding and motion estimation techniques, combined with DL, for enhanced HAR. This technique leverages DL through CNN and Bi-GRUs for improved HAR performance.
[17]
Deep BiLSTM
[18]
TransTM
[19]
RNNLSTM
The utilization of the model improved through transfer-learning-based attribute extraction for dynamic HAR. By utilizing a time-streaming multiscale transformer for device-free HAR.
Its capability to perform DL based HAR directly on edge devices.
the combined approach produces a final feature representation that incorporates the original features obtained from MobileNetv2-Lite into the relational information stored in MGCNs. The feature
158
Performance measures used Accuracy-79.3% F1-score-64.15% Training time-0.5ms Inference time-0.05ms Params-1.88M GFLOPs-0.028 Accuracy-95.68
Inability to incorporate contextual information from multiple branches into the model architecture hindered its capability to effectively seize nuanced outlines in the data.
Accuracy-96.37 F1-score-96.31 Training time-53.68 Params-631,014
Lack of testing the modified transformer on a prolonged database, preferably integrating diverse sensor data, limited a comprehensive evaluation. It failed to focus on the advances observed in wearable and phone-based tracking systems.
Accuracy-99.2 Precision-99.2 Recall-99
Absence of evaluation across diverse datasets limits the generalizability of the findings.
Accuracy-97%
Lack of comprehensive feature selection techniques, theoretically leading to suboptimal performance and increased computational complexity. Inability to evaluate across diverse datasets, potentially limiting the generalizability of the model’s performance. Lack of extensive validation across diverse scenarios, potentially limiting the generalizability of its performance outside controlled environments. Inability to constrain resources in edge environments, potentially limiting the model’s performance or scalability on such platforms.
Accuracy-93.38 Average time-7.15s
Accuracy-98.76 Params-634,052 Recall-99.93
Accuracy-99.2%
Accuracy-99.1% F1-score-99.1% FLOPs-27.98M Params-1.82M Accuracy-99.31% Recall-99.09% F1-score-99.06% Time-0.298ms
extraction method with a MobileNetv2-Lite backbone and MGCNs enables efficient processing of intricate graph-structured data, providing important insights for comprehension and decision-making in different
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
JHMDB dataset
Big data assessment
UCF 11
Local Linear Synthetic Minority Over-sampling Technique
UCF youtube action HMDB51 dataset
N∘ 3
2026
Feature extraction Bottleneck 1 Bottleneck 3
Bottleneck 5 Bottleneck 7 Avgpool 7
7
Input
Conv 1
Bottleneck2
Bottleneck 4
Bottleneck 6
Conv 2
GCM
Conv 3
MobileNetversion 2-Lite backbone with Graph Convolutional Networks Enhanced Football Team Training Optimization Algorithm Feature selection
To optimize the parameters of MCCBCN Relative entropy
Mutual information
Activity recognition Metapath Contexed Convoluted Bi-directional Cascade Network
Construction of metapath
Activity prediction
Strong Correlation Coefficient
Convolutional metapath
metapath metapath metapath metapath
Selected features
Bidirectional cascade network
Figure 1. Workflow of the proposed work
Bottleneck 1 Bottleneck 3 Bottleneck 5 Bottleneck 7 Avgpool 224
112 224
3 Input
112
56
28 14
14
7
7
7
28
7 7 7 14 14 112 64 96 160 320 128 1280 16 24 32 Bottleneck 4 Bottleneck 6 Conv2 GCM 32 Bottleneck2 Conv1 112
56
1 1280 Conv3
1
Figure 2. MobileNetV2 with GCN model application fields. The extracted features are then transferred to feature selection to pull out the more important features [25, 26]. 3.4.
Feature selection using parameters
The procedure for identifying the most pertinent features involves combining all extracted features. This selection process relies on three key parameters: relative entropy, mutual information, and the SCC, which are assessed using a maximum probabilitybased threshold function.
• Relative entropy Relative entropy measures the difference between two probability distributions. In this feature selection, it evaluates the information gain provided by each feature compared to a baseline distribution. Features with higher relative entropy values are considered more relevant, as they contribute significantly to distinguishing between different categories in the dataset. The relative entropy is expressed in Eqn. (1): 𝐺𝐿𝐾 ( 𝑄‖ 𝑃) = Σ𝑥 𝑃 (𝑥) log �
𝑃 (𝑥) � 𝑄 (𝑥)
(1)
159
Journal of Automation, Mobile Robotics and Intelligent Systems
where, 𝐺𝐿𝐾 ( 𝑄‖ 𝑃) is the relative entropy between distributions of 𝑄 and 𝑃, 𝑃(𝑥) is the probability of a feature 𝑥 occurring in the database and 𝑄(𝑥) is the probability of a feature 𝑥 occurring in the baseline distribution. Higher relative entropy values indicate features that contribute more to discriminating between different categories in the dataset. • Mutual information It quantifies the dependency among each feature 𝑎 and the target variable 𝑦 (e.g., action labels in human motion data). It measures how much information one variable delivers regarding another variable. Therefore, the expression for mutual information is given in Eqn. (2): 𝐼(𝐴; 𝐵) = � � 𝑃(𝐴, 𝐵) log � 𝐴
𝐵
𝑃(𝐴, 𝐵) � 𝑃(𝐴) 𝑃(𝐵)
(2)
where, 𝐼(𝐴; 𝐵) is the mutual information between features 𝐴 and the target variable 𝐵, Higher mutual information values indicate stronger associations among features and the target variable. • Strong correlation coefficient The SCC evaluates the linear correlation strength between pairs of features. It identifies features that exhibit strong correlations with each other, which indicate redundancy in the feature set. To avoid redundancy and overfitting, a maximum probability-based threshold function is applied to the SCC values. The Pearson correlation coefficient formula for SCC is expressed in Eqn. (3): 𝜌 (𝐴, 𝐵) =
𝐶𝑜𝑣 (𝐴, 𝐵) 𝜎𝑎 𝜎𝑏
(3)
where, 𝜌 (𝐴, 𝐵) is the Pearson correlation coefficient between features 𝐴 and 𝐵, is the covariance between features 𝐴 and 𝐵, 𝜎𝑎 and 𝜎𝑏 are the standard deviations of features 𝐴 and 𝐵, respectively. By setting a threshold 𝜏, features with SCC values above this threshold are selected, as expressed in Eqn. (4): 𝑆𝑒𝑙𝑒𝑐𝑡𝑒𝑑 𝑓𝑒𝑎𝑡𝑢𝑟𝑒𝑠 = { 𝑋𝑖 | 𝜌 (𝐴, 𝐵) > 𝜏}
(4)
By considering these parameters and their corresponding equations, the feature selection process ensures that the selected attributes are extremely informative, non-redundant, and relevant for the given DL task. The chosen attributes are then moved to MCCBCN for activity recognition. 3.5.
Activity recognition using metapath contexted convoluted bi-directional cascade network
During the HAR process, the extracted features, characterized by fixed sequence lengths, are forwarded to the proposed method, MCCBCN. This phase is critical for capturing the complex dependencies that characterize human activities in video data streams. The MCCBCN is designed to effectively model and analyze these dependencies by leveraging a combination 160
VOLUME 20,
N∘ 3
2026
of advanced techniques. In metapath contextual learning, metapaths are sequences of actions or poses over time, capturing the semantic relationships between different activities. The MCCBCN model is proposed to boost human activity recognition by combining graph-based context modeling and deep sequence learning. It utilizes metapath-guided context extraction to capture semantic high-level interactions between heterogeneous activity features, allowing the model to identify minute interdependencies between actions. The convolutional layers are nonlocal spatial feature extractors from input streams (e.g., sensor or video frames), and the bidirectional cascade modules yield efficient temporal aggregation by progressively combining forward and backward temporal flows. The cascade mechanism facilitates hierarchical fusion, in which outputs from previous layers are iteratively refined by later-stage features to ensure deep temporal coherence. The model’s architecture prioritizes cross-context attention propagation and multi-level abstraction to be able to well capture both long-range activity dependencies and short-term transitions in an end-to-end trainable, scalable manner. Mathematically, a metapath 𝜌 in a diverse evidence denoted as a sequence of node types 𝑣1 , 𝑣2 , ....., 𝑣𝑘 and edge types 𝑒1 , 𝑒2 , ....., 𝑒𝑘−1 . The HAR is expressed in Eqn. (5): 𝜌 = �𝑣𝑠𝑖𝑡𝑡𝑖𝑛𝑔 , 𝑒𝑡𝑟𝑎𝑛𝑠𝑖𝑡𝑖𝑜𝑛 , 𝑣𝑠 tan 𝑑𝑖𝑛𝑔 , 𝑒𝑡𝑟𝑎𝑛𝑠𝑖𝑡𝑖𝑜𝑛 , 𝑣𝑤𝑎𝑙𝑘𝑖𝑛𝑔 �
(5)
By utilizing these metapaths, the network understands and controls the temporal context in which each activity occurs. The contextual information is crucial for distinguishing activities that appear similar in isolated frames but differ in their temporal sequence. It also involves convolutional processing, where GCNs are applied to nodes and edges defined by the metapaths. In HAR, nodes represent individual frames or body joints, and edges represent transitions between these nodes over time. The convolution operation in a GCN for a node 𝑖 is expressed in Eqn. (6): (𝑙+1)
ℎ𝑖
(𝑙)
= 𝜎 �Σ𝑗∈𝑁(𝑖)
1 �𝑑𝑖 𝑑𝑗
(𝑙)
𝑊 (𝑙) ℎ𝑗 �
(6)
where, ℎ𝑗 is the attribute vector of the node 𝑖 at layer 𝑙, 𝑁(𝑖) is the set of neighbors of the node 𝑖, 𝑑𝑖 and 𝑑𝑗 are the degrees of the nodes of 𝑖 and 𝑗, 𝑊 (𝑙) is the weight matrix at layer 𝑙, 𝜎 is an activation function. This operation aggregates features from neighboring nodes, capturing spatial dependencies and learning robust features representing spatial configurations and movement patterns. Bi-directional learning processes data in both forward and backward directions along the metapaths, analyzing sequences of actions from the past to the future and vice versa.
Journal of Automation, Mobile Robotics and Intelligent Systems
In the bi-directional learning method, the hidden state ℎ𝑡 at time 𝑡 is a mixture of the forward hid← den state ℎ⃗ 𝑡 and backward hidden state ℎ 𝑡 which is expressed in Eqn. (7): ←
ℎ𝑡 = �ℎ⃗ 𝑡 ; ℎ 𝑡 �
(7)
The forward and backward hidden states are expressed as in Eqns. (8) and (9): ℎ⃗ 𝑡 = 𝑓 �𝑊𝑥 𝑥𝑡 + 𝑊ℎ ℎ⃗ 𝑡−1 + 𝑏�
(8)
←
⃖��𝑡+1 + 𝑏� ℎ 𝑡 = 𝑓 �𝑊𝑥 𝑥𝑡 + 𝑊ℎ ℎ
(9)
where, 𝑥𝑡 is the input at time 𝑡, 𝑊𝑥 and 𝑊ℎ are weight matrices, 𝑏 is the bias vector, and 𝑓 is the activation function. Bi-directional learning captures dependencies that span across different time frames, providing a more comprehensive understanding of the activity. MCCBCN employs a multi-layered approach, where the outcome from one layer is used as the input for the subsequent layer. This cascading structure allows for progressive refinement of the learned features. The input of the 𝑙𝑡ℎ layer 𝐻(𝑙) is expressed as in Eqn. (10): 𝐻(𝑙+1) = 𝜎 �𝑊 (𝑙) 𝐻(𝑙) + 𝑏(𝑙) �
(10)
where, 𝐻(𝑙+1) is the output of the layer 𝑙, 𝑊 (𝑙) is the weight matrix, 𝑏(𝑙) is the bias vector. In HAR, the first layer detects basic actions like “aising hand” or “moving leg,” the next layer recognizes more complex sequences like “waving” or “walking,” and the final layer identifies overall activities like “greeting” or “exercising.” This multi-stage processing builds increasingly complex and discriminative representations. MCCBCN is a powerful method for HAR that leverages metapath-based contextual learning, convolutional processing, bidirectional learning, and cascading networks. By capturing temporal, spatial, and behavioral dependencies, MCCBCN accurately recognizes and differentiates various activities, even in complex environments. The network’s capability to learn from metapaths, aggregate spatial features, process bidirectional sequences, and refine features through cascading layers results in robust HAR [27, 28]. 3.6.
Enhanced football team training optimization algorithm to boost the performance of MCCBCN
The introduction of EFTTOA aims to enhance the performance of MCCBCN in HAR. The EFTTOA, implemented as the training policy for MCCBCN, further increases the adaptability and convergence efficiency of the model. EFTTOA mimics the dynamics of a football team, with players (solution agents) assigned roles and strategies through cooperative learning and competitive performance. Agents experience repeated formation change, role exchange, and ability adaptation, allowing for effective global exploration in the parameter space. Adaptive balance between exploration (searching new solutions) and
VOLUME 20,
N∘ 3
2026
exploitation (optimizing promising regions) in EFTTOA prevents premature convergence and encourages optimal hyperparameter tuning. It works especially well with the high-dimensional optimization problems posed by the MCCBCN’s deep structure, providing faster convergence rates and better generalization without degrading model stability or increasing computational costs. EFTTOA is a modification of FTTOA [29] that leverages principles from football team training methodologies to optimize the training process of MCCBCN and improve its effectiveness in identifying human emotions from video data streams. To address the challenges of complexity and cost associated with integrating advanced pose estimation techniques in FTTOA, implementing model compression techniques is a suitable enhancement. The fitness function is expressed as in Eqn. (11): 𝐹𝑖𝑡𝑛𝑒𝑠𝑠𝑓𝑢𝑛𝑐𝑡𝑖𝑜𝑛 = min{𝑊 (𝑙) }
(11)
By compressing DL models like Convolutional Pose Machines (CPMs), the computational complexity and resource requirements are significantly reduced without compromising performance, which makes the models more efficient and cost-effective to implement. The mathematical model compression is expressed as in Eqn. (12): Compressed model = arg min 𝑀′ 𝐿(𝑀′ , 𝐷) + 𝜆 ⋅ Compression_loss(𝑀′ ) (12) where 𝐿 represents the original loss function, 𝐷 is the training dataset, 𝑀′ is the compressed model, and 𝐶𝑜𝑚𝑝𝑟𝑒𝑠𝑠𝑖𝑜𝑛_𝑙𝑜𝑠𝑠 is a regularization term that penalizes complexity. The hyperparameter 𝜆 controls the trade-off between accuracy and compression. This enhancement reduces the computational burden and resource requirements, making the integration of pose estimation models more feasible for football teams and organizations. Table 2 displays the pseudo code of EFTTO. Overall, the activity recognition utilizes an MCCBCN, which is designed to effectively capture and analyze complex relational data to understand activity contexts more accurately. This network structure processes data in both directions across several layers. This algorithm optimizes the training process by fine-tuning the network’s parameters, resulting in more accurate and efficient activity recognition. This integrated method is designed for applications that need detailed and accurate activity recognition.
4.
Results and Discussions
In this section, the efficiency of the introduced framework for HAR is estimated by relating it to existing methods across four challenging datasets, such as JHMDB [20], UCF 11 [21], UCF YouTube Action [22], and HMDB51 [23]. This research is implemented using Python. The parameter setting of the introduced approach is explained as follows: In this training setup, the maximum detection size for the input data is set 161
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 2. Pseudo code of EFTTO Initialize FTTA CLC // Function parameters objectives set Set 𝑀𝑎𝑥𝑔𝑒𝑛 = 500 // Iterations numbers Set 𝑛𝑝𝑜𝑝 = 50 // Population size Set 𝑃𝑠𝑡𝑢𝑑𝑦 = 0.2 // Probability of study Set 𝑃𝑐𝑜𝑚𝑚 = 0.2 // Probability of communication Set 𝑃𝑒𝑟𝑟𝑜𝑟 = 0.001 // Probability of error // Initialization For I = 1 to 𝑀𝑎𝑥𝑔𝑒𝑛 Do // Initialization loop starts here For each 𝑖 from 1 to 𝑛𝑝𝑜𝑝 Do The fitness value of each football player is calculated, which is expressed in Eqn. (11): End for Find the finest player & the worst player // Collective Training Phase begins Collective training:
⎧ ⎪ 𝑘 𝐻𝑖,𝑗 𝑛𝑒𝑤 =
𝑘 𝑘 𝑘 𝐻𝑖,𝑗 𝑜𝑙𝑑 + 𝑟𝑎𝑛𝑑 ∗ �𝐻𝑏𝑒𝑠𝑡,𝑗 − 𝐻𝑖,𝑗 𝑜𝑙𝑑� 𝑘 𝑘 𝑘 𝑘 𝑘 𝐻𝑖,𝑗 𝑜𝑙𝑑 + 𝑟𝑎𝑛𝑑1 ∗ �𝐻𝑏𝑒𝑠𝑡,𝑗 − 𝐻𝑖,𝑗 𝑜𝑙𝑑� − 𝑟𝑎𝑛𝑑2 ∗ �𝐻𝑤𝑜𝑟𝑠𝑡,𝑗 − 𝐻𝑖,𝑗 𝑜𝑙𝑑�
⎨ ⎪ ⎩
(13)
𝑘 𝑘 𝑘 𝐻𝑖,𝑗 𝑜𝑙𝑑 + 𝑟𝑎𝑛𝑑 ∗ �𝐻𝑏𝑒𝑠𝑡,𝑗 − 𝐻𝑤𝑜𝑟𝑠𝑡,𝑗 � 𝑘 𝑜𝑙𝑑 ∗ (1 + 𝐾 (𝑡)) 𝐻𝑖,𝑗
𝑘,𝑡𝑒𝑎𝑚1
𝐻𝑖,𝑗
𝑛𝑒𝑤 =
𝑘,𝑡𝑒𝑎𝑚1
𝐻𝑖,𝑗
𝑘,𝑡𝑒𝑎𝑚
𝐻𝑏𝑒𝑠𝑡,𝑗 1 𝑖𝑓 𝑟𝑎𝑛𝑑 ≤ 𝑝𝑠𝑡𝑢𝑑𝑦
⎧ ⎪
𝑘,𝑡𝑒𝑎𝑚
1 𝐻𝑟𝑎𝑛𝑑𝑜𝑚,𝑗 𝑖𝑓 𝑟𝑎𝑛𝑑 ≤ 𝑝𝑠𝑡𝑢𝑑𝑦 ⎨ 𝑘,𝑡𝑒𝑎𝑚 1 ⎪ 𝐻 ⎩ 𝑟𝑎𝑛𝑑𝑜𝑚,𝑗∗(1+𝑟𝑎𝑛𝑑𝑛) 𝑖𝑓 𝑟𝑎𝑛𝑑≤𝑃𝑐𝑜𝑚𝑚
𝑘,𝑡𝑒𝑎𝑚
𝑛𝑒𝑤 = 𝐻𝑟𝑎𝑛𝑑𝑜𝑚1 1,𝑟𝑎𝑛𝑑𝑜𝑚2 𝑖𝑓 𝑟𝑎𝑛𝑑 ≤ 𝑝𝑒𝑟𝑟𝑜𝑟
(14)
(15)
For every 𝑖 ∶ 1 ≤ 𝑖 ≤ 𝑛Pop do End for Update the players Find the best player Additional individualized training:
𝑘 𝐻𝑏𝑒𝑠𝑡 𝑛𝑒𝑤 = �
1
1
𝑘
𝑘
𝑘 𝐻𝑏𝑒𝑠𝑡 𝑜𝑙𝑑 ∗ �1 + � � ∗ 𝑔𝑎𝑢𝑠𝑠 +
𝑘 𝑘 ∗ 𝑐𝑎𝑢𝑐ℎ𝑦� 𝑖𝑓 𝐻 �𝐻𝑏𝑒𝑠𝑡 𝑛𝑒𝑤� < 𝐻 �𝐻𝑏𝑒𝑠𝑡 𝑜𝑙𝑑�
(16)
𝑘 𝑘 𝑘 𝐻𝑏𝑒𝑠𝑡 𝑜𝑙𝑑 𝑖𝑓 𝐻 �𝐻𝑏𝑒𝑠𝑡 𝑛𝑒𝑤� > 𝐻 �𝐻𝑏𝑒𝑠𝑡 𝑜𝑙𝑑�
End for 𝑇 =𝑇+1 Exit the while loop Exit the for loop
to 256 × 256 pixels. The model is trained over 100 epochs using the EFTTOA optimizer. Each epoch consists of 100 iterations with a batch size of 64, meaning 64 models are handled before the approach’s parameters are reorganized. The learning rate is set to 0.0001. Additionally, a dropout rate of 0.03 is applied to avoid overfitting by randomly dropping 3% of the neurons during each iteration. Figures 3 (a), (b), (c), and (d) display the confusion matrix of the corresponding dataset. This confusion matrix is utilized to assess the performance of a recognition model by exhibiting the counts of actual versus predicted values. 162
Fig. 4 shows the class-wise accuracies of the corresponding datasets. 4.1.
Qualitative analysis
Figures 5 (a), (b), (c), and (d) have four videos selected from the test set to evaluate and analyze the experimental quality of HAR generated by the proposed method, comparing it with the corresponding datasets. In the first, a girl is seen combing her hair. Upon examining all tags for HAR, it was noted that this specific activity is absent from the training data.
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
(a) JHMDB dataset
(b) UCF 11 dataset
(c) UCF YouTube action dataset
(d) HMDB51 dataset
N∘ 3
2026
Figure 3. Confusion matrix of the corresponding datasets
(a)
(b)
(c)
Figure 4. (a, b and c). Class-wise accuracies of the corresponding datasets
163
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
Human activity recognition of dataset 1
N∘ 3
2026
Human activity recognition of dataset 2
Ground truth: Brush_hair Recognition: Brush_hair
Ground truth: Trampoline_jumping Recognition: Trampoline_jumping
Ground truth: Tennis_swing Recognition: Tennis_swing
Ground truth: Climb_ stairs Recognition: Climb_ stairs
Ground truth: Horse_riding Recognition: Horse_riding Ground truth: Pull_up Recognition: Pull_up
(b)
(a) Human activity recognition of dataset 4
Human activity recognition of dataset 3
Ground truth: Basket ball Recognition: Basket ball
Ground truth: Laugh Recognition: Laugh
Ground truth: Soccer_ juggling Recognition: Soccer_ juggling
Ground truth: Eat Recognition: Eat
Ground truth: Volleyball_ spiking Recognition: Volleyball_ spiking
Ground truth: Kick_ball Recognition: Kick_ball
(d)
(c)
Figure 5. (a), (b), (c) and (d). Activity recognition outcomes using the proposed method Therefore, the proposed method correctly recognizes the activity. The investigation underscores the critical influence of input video data on the proposed method. Through accurate activity recognition, this approach enhances surveillance systems by enabling nuanced analysis of human behavior, improving decisionmaking and bolstering safety and security measures. As a result of its precise recognition capabilities, the proposed model significantly improves overall surveillance system performance. 4.2.
Performance comparison with JHMDB and HMDB51 datasets
Comparative evaluations are performed by contrasting the proposed method with corresponding datasets, employing these performance metrics. This evaluation assesses the performance of the MCCBCN system against traditional approaches such as TDnet [10], MF + MVF [13], ViCo-MoCo [15], and Deep BiLSTM [17]. Table 3 shows the accuracy of different methods on a specific dataset. TD-net gets 79.3% accuracy, proving it works well for predictions. Deep BiLSTM reaches 76.3% accuracy, while ViCo-MoCo hits 78.5%. MF + MVF does better with 87.76% accuracy, showing it predicts more. The new method, MCCBCN, beats them all with 99.34% accuracy, showing it’s the best at predicting results for this dataset.
164
Table 4 shows how well different methods perform, based on their accuracy percentages for a specific dataset. ViCo-MoCo has an accuracy of 71.4%. MF + MVF and Bi-GRU have accuracies of 71.28% and 71.89%, respectively. Also, CNN-RNN reaches 70.8% accuracy. The “MCCBCN” (Proposed) method tops the list with a remarkable 99.45% accuracy. 4.3.
Performance analysis of human activity recognition
The efficiency is assessed using several metrics across relevant datasets. Comparative evaluations are performed by contrasting the proposed method with Table 3. Accuracy analysis comparison of the JHMDB dataset Techniques TD-net [10] MF + MVF [13] ViCo-MoCo [15] Deep BiLSTM [17] MCCBCN (Proposed)
Accuracy (%) 79.3 87.76 78.5 76.3 99.34
Table 4. Accuracy analysis comparison of the HMDB51 dataset Techniques CNN-RNN [11] MF + MVF [13] ViCo-MoCo [15] CNN Bi-GRU [16] MCCBCN (Proposed)
Accuracy (%) 70.8 71.28 71.4 71.89 99.45
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
(a)
(b)
(c)
(d)
N∘ 3
2026
Figure 6. (a) accuracy, (b) precision, (c) recall, and (d) F1-score on recent HAR techniques established techniques, employing these performance metrics. This evaluation compares the performance of the proposed system with traditional approaches such as CNN BiLSTM [12], HAR transformer [13], CNN LSTM [14], TransTM [18], and RNN LSTM [19]. Figures 6 (a), (b), (c), and (d) illustrate various HAR models assessed using key performance metrics: accuracy, precision, recall, and F1-score. The CNNBiLSTM model achieved an accuracy of 96.37%, with corresponding precision, recall, and F1-score values of 96.14%, 96.14%, and 96.31%, respectively. The HAR transformer model established high performance across all metrics, attaining an accuracy of 99.2% along with consistent precision, recall, and F1-score values of 99.2%. Similarly, the CNN-LSTM model showed superior accuracy at 99.8%, with precision, recall, and F1-score values of 98.89%, 98.78%, and 99.7%. The DCNN model achieved an accuracy of 99.80%, with precision, recall, and F1-score values of 99.31%. TransTM achieved an accuracy of 99.2% with a precision, recall, and F1-score of 99%, 99%, and 99.1%, respectively. MCCBCN outperformed with an accuracy of 99.50% and precision, recall, and F1-score values of 99.42%, 99.3%, and 99.9%, respectively. Fig. 7 illustrates the time complexity of the introduced approach compared to existing techniques like TD-Net, MF+MVF, and ViCo-MoCo. The time complexities for these methods are as follows: TD-Net takes 35 seconds, MF+MVF takes 42 seconds, and ViCo-MoCo takes 27.9 seconds. In contrast, the proposed method, MCCBCN, significantly outperforms these techniques with a time complexity of just 10.2 seconds. The introduced MCCBCN model efficiently surpasses current HAR methods by solving major challenges with contextual perception, temporal dependency modeling, and class imbalance treatment. In contrast to
Figure 7. Time complexity analysis
models like CNN-BiLSTM or HAR Transformer, which either do not leverage multi-branch contextual information or are not evaluated across diverse datasets, MCCBCN presents a metapath-based contextual learning approach that improves its capacity to identify intricate and overlapping activities. The incorporation of a bidirectional cascade structure further enhances temporal feature extraction by modeling forward and backward dependencies across activity sequences, a capability not always achieved by most current models. The application of the EFTTOA also simplifies training and enhances learning efficiency with faster convergence and better generalization over diverse datasets. Additionally, MCCBCN addresses the vital problem of class imbalance by using LSMOST to effectively represent minority-class activities and learn them during training. This ensures the model performs more balanced than other approaches, such as ViCo-MoCoDL or TransTM, which have not received sufficient validation on diverse or imbalanced datasets. Having an accuracy of 99.50%, an F1-score of 99.9%, and a reasonable time complexity of 10.2 seconds, 165
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 8. K-fold validation analysis 4.4.
Figure 9. Fitness curve
MCCBCN raises the bar in HAR by achieving robustness, precision, and scalability combined. Its end-toend design promises higher performance not just in conventional metrics but also in real-world usability, where subtle activity recognition and generalization are paramount. Fig. 8 depicts performance metrics for the proposed model, MCCBCN, evaluated using the average 10-fold cross-validation. Fig. 9 compares the performance of four optimization algorithms, like EFTTOA (proposed), FTTOA, Hierarchical Particle Swarm Optimization algorithm (HPSOA) [30], and Stochastic Gradient Descent Optimization Algorithm (SGDOA) [31], over 100 iterations, with fitness values plotted against iterations. The proposed EFTTOA algorithm shows significantly faster convergence, reaching a lower fitness value in fewer iterations, compared to the existing algorithms. This recommends that EFTTOA is more efficient and effective in optimizing the given problem, as it consistently achieves better results more quickly and stabilizes earlier than FTTOA, HPSOA, and SGDOA. 166
Ablation study
This research on HAR investigates the impact of removing specific components from the MCCBCN architecture on accuracy, precision, recall, and F1score across three datasets, such as JHMDB, UCF 11, UCF YouTube action, and HMDB51. Table 5 shows the performance of the MCCBCN architecture for HAR across four datasets, evaluated alongside various model configurations. The full MCCBCN architecture achieves exceptionally high accuracy across all datasets, ranging from 99.34% to 99.45%, indicating its robustness and effectiveness. When specific components of the MCCBCN architecture are excluded, like metapath context, Bidirectional cascade, EFTTOA Integration, LSMOST, and MobileNetV2-Lite & MGCN Integration, there is a noticeable decrease in accuracy. For instance, taking out the metapath context makes the accuracy fall between 85.9% and 86.7%. This shows how important it is to capture context information to boost performance. In the same way, removing the Bi-directional cascade increases accuracy from 86.2% to 89.3%. This points out how crucial this setup is to make the model more accurate.
5.
Conclusion
This research introduced the MCCBCN method, which leverages metapaths to enhance the accuracy and robustness of contextual information capture for HAR. Significant advancements in the field were achieved through the utilization of both convolutional and bi-directional cascade architectures. The model’s accuracy and robustness were further boosted by integrating the EFTTOA. Additionally, the LSMOST was employed for dataset expansion, ensuring a comprehensive and balanced representation of minority classes, mitigating biases from imbalanced class distributions, and improving model
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Table 5. Ablation study of the introduced method with their datasets Model Configuration Full MCCBCN architecture Without Metapath Context Without Bi-directional Cascade Without EFTTOA Integration Without LSMOST Without MobileNetV2-Lite & GCN Integration
JHMDB Accuracy (%) 99.34 85.9 89.3 89.1 88.5 88.9
generalizability. The integration of the MobileNetV2Lite backbone with MGCN allowed for the extraction of human-centric features, while multi-view feature computation captured diverse activity patterns. Selection criteria based on relative entropy, mutual information, and the SCC ensured the inclusion of the most relevant features, thereby enhancing model performance and interpretability. The performances were recorded as 99.50% (accuracy), 99.9% (F1-score), 99.42% (precision), and 99.3% (recall), respectively. The accuracy comparison between the JHMDB and HMDB51 datasets showed scores of 99.34% for both datasets. The time complexity of the proposed methods is 10.2 seconds. These validate the dominance of the MCCBCN model in achieving highly accurate, robust, and reliable HAR, outperforming existing methods across key evaluation metrics such as FDR and FOR, especially under varying prevalence conditions. Future scope In future work, we aim to enable the proposed HAR system to support real-time processing, allowing instantaneous analysis and decision-making on widely varying, continuous data streams. The goals will be to modify the MCCBCN architecture for low-latency applications, to combine edge computing so that localized processing takes place, and to investigate novel compression and transmission techniques to efficiently manage big data from sensors. Additionally, the model’s flexibility to varying implementations in real-world situations and hardware platforms is ensured to ensure wider use and deployment. Data availability statements Data sharing not applicable to this article as no datasets were generated or analyzed during the current study. Funding This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
UCF 11 Accuracy (%) 99. 39 86.7 86.2 88.1 85.5 88.1
UCF YouTube Action Accuracy (%) 99.42 85.9 87.7 88.5 86.1 86.7
HMDB51 Accuracy (%) 99.45 85.3 88.3 87.2 87.8 87.7
AUTHORS Praveen S. Banasode∗ – Department of Master of Computer Applications, Jain College of Engineering, Belagavi, Karnataka 590014, India, e-mail: praveenb.jce@gmail.com. Sunita Padmannavar – Department of Master of Computer Applications, Gogte Institute of Technology, Belagavi 590006, Karnataka, e-mail: sunitapdm@gmail.com. ∗
Corresponding author
ACKNOWLEDGEMENTS None
References [1] Hussain, S. U. Khan, N. Khan, M. Shabaz, and S. W. Baik, “AI-driven behavior biometrics framework for robust human activity recognition in surveillance systems,” Engineering Applications of Artificial Intelligence, vol. 127, p. 107218, Jan. 2024. doi: 10.1016/j.engappai.2023.107218 [2] Kabiruzzaman, M. Shidujaman, S. Hassan Shifat, P. Debnath, and S. Hossain, “Time series analysis of Care Records data for nurse activity recognition in the wild,” Human Activity and Behavior Analysis, pp. 405–415, Mar. 2024. doi: 10.1201 /9781003371540-28 [3] A. Kushwaha, A. Khare, and O. Prakash, “Human activity recognition algorithm in video sequences based on the fusion of multiple features for realistic and multiview environment,” Multimedia Tools and Applications, vol. 83, no. 8, pp. 22727–22748, Aug. 2023. doi: 10.1007/s11042-023-16364-z [4] S. Ankalaki and M. N. Thippeswamy, “Optimized convolutional neural network using hierarchical particle swarm optimization for sensor based human activity recognition,” SN Computer Science, vol. 5, no. 5, Apr. 2024. doi: 10.1007/s42 979-024-02794-5 [5] Md. M. Islam, S. Nooruddin, F. Karray, and G. Muhammad, “Multi-level feature Fusion for multimodal human activity recognition in internet of healthcare things,” Information Fusion, vol. 94, pp. 17–31, Jun. 2023. doi: 10.1016/j.inffus.2023 .01.015
167
Journal of Automation, Mobile Robotics and Intelligent Systems
[6] T. R. Mim et al., “Gru-Inc: An inception-attention based approach using GRU for human activity recognition,” Expert Systems with Applications, vol. 216, p. 119419, Apr. 2023. doi: 10.1016/j.e swa.2022.119419 [7] Z. Amiri, A. Heidari, N. J. Navimipour, M. Unal, and A. Mousavi, “Adventures in data analysis: A systematic review of deep learning techniques for pattern recognition in cyber-physical-social systems,” Multimedia Tools and Applications, vol. 83, no. 8, pp. 22909–22973, Aug. 2023. doi: 10. 1007/s11042-023-16382-x [8] A. Ray, M. H. Kolekar, R. Balasubramanian, and A. Hafiane, “Transfer learning enhanced vision-based human activity recognition: A decade-long analysis,” International Journal of Information Management Data Insights, vol. 3, no. 1, p. 100142, Apr. 2023. doi: 10.1016/j.jjimei.2022.100142 [9] M. K. Jannat, Md. S. Islam, S.-H. Yang, and H. Liu, “Efficient Wi-Fi-based human activity recognition using adaptive antenna elimination,” IEEE Access, vol. 11, pp. 105440–105454, 2023. doi: 10.1109/access.2023.3320069 [10] T. Nguyen, D. Pham, H. Vu, and T. Le, “A robust and efficient method for skeleton-based Human action recognition and its application for crossdataset evaluation,” IET Computer Vision, vol. 16, no. 8, pp. 709–726, Jul. 2022. doi: 10.1049/cvi2. 12119 [11] C. Zhang, Y. Xu, Z. Xu, J. Huang, and J. Lu, “Hybrid handcrafted and learned feature framework for Human Action Recognition,” Applied Intelligence, vol. 52, no. 11, pp. 12771–12787, Feb. 2022. doi: 10.1007/s10489-021-03068-w [12] S. K. Challa, A. Kumar, and V. B. Semwal, “A multibranch CNN-BILSTM model for human activity recognition using wearable sensor data,” The Visual Computer, vol. 38, no. 12, pp. 4095–4109, Aug. 2021. doi: 10.1007/s00371-021-02283-3 [13] I. Dirgová Luptá ková , M. Kubovč ı́k, and J. Pospı́chal, “Wearable sensor-based human activity recognition with Transformer model,” Sensors, vol. 22, no. 5, p. 1911, Mar. 2022. doi: 10.3390/s22051911 [14] Mst. A. Khatun et al., “Deep CNN-LSTM with selfattention model for human activity recognition using wearable sensor,” IEEE Journal of Translational Engineering in Health and Medicine, vol. 10, pp. 1–16, 2022. doi: 10.1109/jtehm.2022. 3177710
168
VOLUME 20,
N∘ 3
2026
gated recurrent unit with features selection,” IEEE Access, vol. 11, pp. 33148–33159, 2023. doi: 10.1109/access.2023.3263155 [17] N. Hassan, A. S. Miah, and J. Shin, “A deep bidirectional LSTM model enhanced by transferlearning-based feature extraction for dynamic human activity recognition,” Applied Sciences, vol. 14, no. 2, p. 603, Jan. 2024. doi: 10.3390/a pp14020603 [18] Y. Liu et al., “TransTM: A device-free method based on time-streaming multiscale transformer for human activity recognition,” Defence Technology, vol. 32, pp. 619–628, Feb. 2024. doi: 10.1016 /j.dt.2023.02.021 [19] A. C. Cob-Parro, C. Losada-Gutié rrez, M. Marró nRomera, A. Gardel-Vicente, and I. Bravo-Muñ oz, “A new framework for deep learning video based human action recognition on the edge,” Expert Systems with Applications, vol. 238, p. 122220, Mar. 2024. doi: 10.1016/j.eswa.2023.122220 [20] https://figshare.com/articles/dataset/JHMDB/ 19179260 [21] https://www.kaggle.com/datasets/pypiahma d/realistic-action-recognition-ucf50-dataset [22] https://www.kaggle.com/datasets/pypiahma d/ucf-youtube-action-data-set [23] https://www.kaggle.com/datasets/easonlll/h mdb51 [24] K. Wang, T. Zhou, M. Luo, X. Li, and Z. Cai, “Generative adversarial minority enlargement—a local linear over-sampling synthetic method,” Expert Systems with Applications, vol. 237, p. 121696, Mar. 2024. doi: 10.1016/j.eswa.2023.121696 [25] L. Si et al., “A novel coal-gangue recognition method for top coal caving face based on IALOVMD and improved mobilenetv2 network,” IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1–16, 2023. doi: 10.1109/tim.2 023.3316250 [26] Z. Chen et al., “Learnable graph convolutional network and feature fusion for multi-view learning,” Information Fusion, vol. 95, pp. 109–119, Jul. 2023. doi: 10.1016/j.inffus.2023.02.013 [27] X. Fu and I. King, “MECCH: Metapath context convolution-based heterogeneous graph neural networks,” Neural Networks, vol. 170, pp. 266–275, Feb. 2024. doi: 10.1016/j.neunet.20 23.11.030
[15] T. Shanableh, “Vico-Moco-DL: Video coding and motion compensation solutions for human activity recognition using deep learning,” IEEE Access, vol. 11, pp. 73971–73981, 2023. doi: 10.1109/a ccess.2023.3296252
[28] J. Ye, Z. Yu, Y. Wang, D. Lu, and H. Zhou, “PlantBiCNet: A new paradigm in plant science with bidirectional cascade neural network for detection and counting,” Engineering Applications of Artificial Intelligence, vol. 130, p. 107704, Apr. 2024. doi: 10.1016/j.engappai.2023.107704
[16] T. Ahmad et al., “Human activity recognition based on deep-temporal learning using convolution neural networks features and bidirectional
[29] Z. Tian and M. Gai, “Football team training algorithm: A novel sport-inspired meta-heuristic optimization algorithm for global optimization,”
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Expert Systems with Applications, vol. 245, p. 123088, Jul. 2024. doi: 10.1016/j.eswa.2023.1 23088 [30] Z. Li, Y. Liu, B. Liu, J. Le Kernec, and S. Yang, “A holistic human activity recognition optimisation using AI techniques,” IET Radar, Sonar & Navigation, vol. 18, no. 2, pp. 256–265, Sep. 2023. doi: 10.1049/rsn2.12474 [31] I. Priyadarshini, R. Sharma, D. Bhatt, and M. Al-Numay, “Human activity recognition in cyberphysical systems using optimized machine learning techniques,” Cluster Computing, vol. 26, no. 4, pp. 2199–2215, Aug. 2022. doi: 10.1007/s10586-022-03662-8
169
VOLUME 20, N° 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
Automated Detection and Severity Grading of Knee Osteoarthritis from X-ray Images Using Machine Learning with detailed literature Review Submitted: 14th August 2024; accepted: 23rd April 2025
R.Gokulapriya, M.Balamurugan, Rakoth Kandan Sambandam, Divya Vetriveeran DOI: 10.14313/jamris-2026-048 Abstract: This paper presents a groundbreaking approach to the detection and severity grading of knee osteoarthritis (OA) utilizing X-ray imaging and a Convolutional Neural Network (CNN) architecture, specifically EfficientNet B5. Osteoarthritis, a prevalent degenerative joint disease, requires accurate and timely diagnosis for effective management. The proposed methodology focuses on enhancing diagnostic accuracy through advanced deep learning techniques. The study leverages a diverse dataset of knee X-ray images sourced from various clinical settings, encompassing a spectrum of OA severity levels. The EfficientNet B5 architecture, known for its superior performance and efficiency, is employed for feature extraction and classification. The CNN model is trained and validated on a defined dataset, achieving an impressive accuracy of 95.84 percent in differentiating knee OA from normal conditions across multiple grades like Healthy, Minimal, Doubtful, Moderate, and Severe. The model further enhances clinical utility and incorporates a severity grading mechanism, providing clinicians with a detailed assessment of the disease progression. The severity grading is accomplished through fine-tuning the CNN model on a sub-dataset with annotated severity levels. The integration of severity grading demonstrates the model’s ability to detect knee OA accurately and quantify its severity with high precision. The findings of this research highlight the potential of EfficientNet B5-based CNNs as a robust and efficient tool for knee OA diagnosis and severity grading. The reported accuracy of 95.84 percent positions the proposed model as a promising asset in clinical settings, offering a reliable and automated solution for early detection and precise evaluation of knee OA. This research, with its potential to significantly improve the diagnostic capabilities of musculoskeletal disorders, ultimately enhances patient care and outcomes by providing a more accurate and timely diagnosis.
particularly in the elderly population. The knee joint, a complex hinge joint crucial for mobility, undergoes changes with knee OA that involve the deterioration of cartilage, the formation of bony outgrowths (osteophytes), and inflammation of the synovial membrane. These structural alterations lead to pain, stiffness, and reduced joint function, significantly impacting the quality of life for affected individuals. The application of artificial intelligence (AI) in knee OA diagnosis has emerged as a transformative avenue with significant potential to revolutionize the detection, management, and outcomes of this prevalent musculoskeletal condition. This work stems from the historical challenges of diagnosing knee osteoarthritis promptly and accurately. Traditional methods often fail to comprehensively understand the disease’s progression and deliver precise grading for effective management. Leveraging the power of AI, particularly deep learning techniques, this study aims to revolutionize the diagnosis and management of knee OA by offering a more nuanced and reliable approach. This work, as given in figure 1, presents a new approach to the detection and severity grading of knee OA utilizing X-ray imaging and a Convolutional Neural Network (CNN) architecture, specifically EfficientNet B5. Knee OA requires accurate and timely diagnosis for effective management. The EfficientNet B5 architecture, known for its superior performance and efficiency, is employed for feature extraction and classification. The CNN model is trained and validated on a defined dataset, achieving an impressive accuracy in differentiating knee OA from normal conditions across multiple grades like Healthy, Moderate, and Severe.
Keywords: Knee osteoarthritis, EfficientNet B5, severity grading, patient care
1. Introduction Knee osteoarthritis (OA) is a prevalent degenerative joint disorder characterized by the gradual breakdown of the protective cartilage within the knee joint. This condition affects millions of people worldwide and is a leading cause of pain and disability, 170
Figure 1. Flow diagram describing the model development and testing process
Open Access. © 2026 R.Gokulapriya et al., published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP. This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 License.
Journal of Automation, Mobile Robotics and Intelligent Systems
The study not only introduces a cutting-edge methodology for detection and grading but also compares the performance of the EfficientNet B5 model against conventional models like Random Forest and K-Nearest Neighbors (KNN), as well as advanced neural network architectures such as EfficientNetV2 B3 and Conv2D. Our findings reveal that the EfficientNet B5 model consistently outperforms conventional models Random Forest and KNN. The results showcase a remarkable accuracy for EfficientNetV2 B3 and Conv2D, highlighting the superior accuracy of the EfficientNet B5 model. The motivation behind this technical report lies in the critical need for advanced tools to improve the detection and grading of knee arthritis through x-ray imaging. Traditional methods often face challenges in providing nuanced assessments of arthritis severity, and there is a growing demand for innovative approaches to enhance diagnostic accuracy. Many patients face the challenge of navigating busy schedules and the inconvenience of healthcare appointments. Unnecessary visits to rheumatologists contribute to patient inconvenience and strain healthcare systems grappling with high demand. By harnessing the power of deep-learning models, which can analyze medical data and images with exceptional accuracy, we can offer a transformative solution to this problem. There is potential to provide individuals with an accessible and user-friendly platform enabling them to receive an early diagnosis from their homes. This not only empowers patients to take a proactive role in their healthcare but also has the potential to significantly reduce the traffic to rheumatologists, allowing these specialists to focus on more critical cases and complex scenarios. This report also seeks to contribute to the scientific understanding of applying deep learning, specifically the EfficientNet B5 model, to the detection and severity grading of knee OA in X-ray images. By doing so, we aspire to advance the field of skeletal imaging, providing healthcare practitioners with a valuable resource for more accurate and timely diagnosis, ultimately improving patient outcomes and the overall management of knee OA.
2. Literature Review
The literature on artificial intelligence (AI), and machine learning (ML) applications in knee osteoarthritis (OA) diagnosis is diverse and impactful. Studies focus on various aspects, from diagnostic algorithms to advanced imaging techniques and prediction models. Jang et al. [1] designed a prospective case series study to assess the effectiveness of an automated AI algorithm-based texture analysis in identifying the restoration of subchondral remodeling following therapeutic intervention. Polynucleotide (PN) filler injections were chosen as the therapeutic modality, and treatment outcomes were evaluated through symptom improvement and induced subchondral microstructural changes. The study enrolled 51 participants with knee OA, administering intra-articular PN filler injections weekly for five sessions. Knee X-rays and
VOLUME 20,
N° 3
2026
texture analyses with bone structure values (BSVs) were conducted at the initial screening and three months post-treatment. Visual Analogue Scale and Korean-Western Ontario MacMaster measurements demonstrated decreased scores post-PN treatment, lasting for three months after the final injection. AI algorithm-based texture analysis revealed BSV changes in the tibial bone’s middle and deep layers post-PN injection, suggesting microstructural changes in the subchondral bone were detectable using this method. In conclusion, AI algorithm-based texture analysis shows promise as a tool for detecting and assessing therapeutic outcomes in knee OA. Mohammadi et al. [2] performed a systematic review and meta-analysis to pool the data on diagnostic performance metrics of AI in OA detection and compare them with clinicians’ performances. The researchers conducted a meta-analysis to aggregate data concerning diagnostic performance metrics. In order to identify potential sources of heterogeneity, the researchers undertook a subgroup analysis of the involved joint and a meta-regression based on multiple parameters. The assessment of bias risk was done with the Prediction Model Study Risk of Bias Assessment Tool reporting guidelines. The pooled sensitivities for AI algorithms and clinicians on internal validation test sets were 88 percent and 80 percent. The pooled specificities for AI algorithms and clinicians were 81 percent and 79 percent, respectively. The pooled sensitivity and specificity for AI algorithms at external validation were 94 percent and 91 percent, respectively. In Alshamrani et al. [3], the researchers proposed the utilization of transfer learning models, specifically sequential CNNs, such as Visual Geometry Group 16 (VGG-16) and Residual Neural Network 50, for early detection of OA in knee X-ray images. The analysis reveals that all recommended models exhibit a predictive accuracy exceeding 90 percent in OA detection. Notably, the pre-trained VGG-16 model emerge as the top-performing model, attaining a training accuracy of 99 percent and a testing accuracy of 92 percent. A separate study by Chen et al. [4] systematically employed two deep CNNs to autonomously gauge the severity of knee OA based on the Kellgren-Lawrence (KL) grading system. Initially, utilizing a customized one-stage YOLOv2 network, knee joints were detected in X-ray images, considering the relatively consistent size and distribution of knee joints. Subsequently, popular CNN models, including ResNet variants, VGG, DenseNet, and InceptionV3, were fine-tuned to classify the detected knee joint images. The classification incorporated a novel adjustable ordinal loss, wherein higher penalties were assigned for misclassifications with larger disparities between the predicted and actual KL grades, acknowledging the ordinal nature of the KL grading task. Evaluation of baseline X-ray images from the Osteoarthritis Initiative (OAI) dataset yielded a mean Jaccard index of 0.858 and a recall of 92.2 percent for
171
Journal of Automation, Mobile Robotics and Intelligent Systems
knee joint detection under a Jaccard index threshold of 0.75. For the knee KL grading task, the finetuned VGG-19 model, using the proposed ordinal loss, attained the highest classification accuracy of 69.7 percent and a mean absolute error of 0.344. The study demonstrated state-of-the-art performance in knee OA detection and KL grading. Furthering the research around computer-aided diagnosis (CAD), Brahim et al. [5] presented a fully developed CAD system for early knee OA detection, employing knee X-ray imaging and ML algorithms. The Fourier domain was utilized for initial image preprocessing through a circular Fourier filter. Then, to reduce variability between OA and healthy participants, the researchers used a novel normalizing technique based on predictive modeling with multivariate linear regression. To reduce dimensionality, an independent component analysis technique was used during the feature selection/extraction stage. The classification challenge utilized Random Forest and Gaussian Naï�ve Bayes classifiers. The researchers evaluated the suggested image-based method using 1,024 knee X-ray images from the OAI database. The findings showed a strong predictive classification rate of 82.98 percent accuracy, 87.15 percent sensitivity, and up to 80.65 percent specificity for OA detection. Further research into the use of CAD systems showed promising results in knee OA detection at various stages. Song and Zhang [6] investigated a CAD method for knee osteoarthritis (KOA-CAD) that utilized multivariate information, specifically vibroarthrographic and fundamental physiological signals. The research introduced an improved deep learning model incorporating a novel Laplace distribution-based strategy (LD-S) for classification. Additionally, an aggregated multiscale dilated convolution network (AMD-CNN) was devised to extract features from the multivariate information of knee OA patients. The proposed KOACAD method seamlessly integrated AMD-CNN with LD-S, achieving three computer-assisted diagnosis objectives: automatic knee OA detection, early knee OA detection, and knee OA grading detection. Clinical validation with collected multivariate information demonstrated the method’s effectiveness, yielding automatic knee OA detection accuracy of 93.6 percent, early knee OA detection accuracy of 92.1 percent, and knee OA grading detection accuracy of 84.2 percent. Focusing specifically on knee OA in overweight middle-aged women who initially presented without OA, Lazzarini et al. [7] developed an analytics pipeline to identify small models that could effectively predict the 30-month incidence of knee OA. The comprehensive dataset encompassed clinical variables, responses to food and pain questionnaires, biochemical markers, and imaging-based information. Demonstrating robust performance, all models exhibited high accuracy (area under the curve (AUC) > 0.7), while utilizing only a minimal set of variables. The study delved into the significance and directional impact of each variable within the models. Ultimately, a comparative analysis was conducted, pitting the performance of 172
VOLUME 20,
N° 3
2026
two models against state-of-the-art approaches available in the existing literature. In S.J.O. Rytky et al. [8], the authors aimed to develop and validate a ML approach for the automatic three-dimensional histopathological grading of osteochondral samples imaged with contrast-enhanced micro-computed tomography. A dataset comprising 79 osteochondral cores from 24 total knee arthroplasty patients and two asymptomatic donors was utilized, with imaging conducted using contrast-enhanced micro-computed tomography. Volumes of interest (VOI) within surface, deep, and calcified zones were extracted depth-wise, and a dimensionally reduced Local Binary Pattern-textural feature analysis was applied. Researchers employed regularized linear and logistic regression models and validated the models using nested leave-one-out cross-validation. The models were assessed using mean squared error (MSE) and average precision metrics. An independent test set consisting of 4 mm samples was used for additional validation. V. Pedroia et al.’s paper [9] aims to investigate the capacity of both conventional and deep learning-based T2 relaxometry patterns in distinguishing between knees afflicted with and without radiographic OA. The analysis involved T2 relaxation time maps from 4,384 subjects in the baseline Osteoarthritis Initiative Dataset. Voxel-Based Relaxometry (VBR) facilitated automatic quantification and voxel-based examination of T2 differences between subjects with and without radiographic OA. A Densely Connected CNN (DenseNet) was trained for OA diagnosis using T2 data. To benchmark algorithmic performance, classical feature extraction techniques and shallow classifiers were employed. Evaluation of deep and shallow models, with and without the inclusion of risk factors, was conducted, utilizing Sensitivity and Specificity values along with the McNemar test to compare classifier performance. Innovative algorithmic and deep learning models have shown promising potential in improving knee OA classification efficiency. A. Haseeb et al. [10] introduced a novel approach to classifying knee OA through the integration of deep learning and a whale optimization algorithm, aimed at addressing the time-consuming and costly manual categorization of knee joint disorders in X-ray imaging. Two pretrained deep learning models, EfficientNet-b0 and DenseNet201, were utilized for training and feature extraction through deep transfer learning with fixed hyperparameter values. Fusion employing a canonical correlation approach resulted in a feature vector with enhanced information. An enhanced whale optimization algorithm was then employed for dimensionality reduction. The selected features underwent classification using ML algorithms, including fine-tuned support vector machines (SVMs) and neural networks. Experimental validation on a publicly available dataset yielded a maximum accuracy of 90.1 percent. The system’s interpretability was enhanced through Explainable Artificial Intelligence techniques, specifically
Journal of Automation, Mobile Robotics and Intelligent Systems
occlusion. Comparative analysis with recent research demonstrated a significant improvement in accuracy achieved by the proposed method. Similarly, A. Rehman et al. [11] developed a model capable of effectively diagnosing early-stage OA in knee X-ray images. Utilizing advanced deep learning-based CNNs alongside various ML techniques, the research introduced a novel transfer learning-based feature engineering method known as CRK (CNN, Random Forest, KNN). This technique demonstrated high-performance OA detection. The CRK utilized a 2D-CNN to intelligently extract spatial features from X-ray images, which are then input to Random Forest and KNN techniques, generating a probabilistic feature set. This set was employed in building machine learning-based models. Experimental results indicated that the proposed model achieved a remarkable accuracy score of 99 percent in predicting OA. The performance of each model was rigorously validated through hyperparameter optimization and k-foldbased cross-validation. The study holds the potential to revolutionize OA prediction from X-ray images, showcasing high-performance scores. In a comparative analysis, Teoh et al. [12] developed a multitask model utilizing CNN feature extractors and ML classifiers to identify nine crucial OA features. These features included the KL grade, knee osteophytes and joint-space narrowing, as well as patient-reported pain intensity from plain radiography. A novel feature extraction method was proposed, replacing the fully connected layer with a global average pooling (GAP) layer. The researchers conducted a comparative analysis to assess the effectiveness of 16 different CNN feature extractors and three ML classifiers. The most effective model incorporated the VGG16-GAP feature extractor and KNN classifier. This particular model not only surpassed the performance of the other models under examination, but also outperformed state-ofthe-art methods by demonstrating superior balanced accuracy, higher Cohen’s kappa, greater F1 scores, and lower Mean Squared Error when predicting seven OA features. In their paper, researchers Du, Shan, and Zhang [13] discovered hidden biomedical insights from knee magnetic resonance imaging (MRIs)s to predict OA. The Cartilage Damage Index information was computed from 36 strategic locations in the tibiofemoral compartment through 3D MRI reconstruction. Principal component analysis was applied to process the feature set, which was then utilized along with the original raw feature set as inputs for four distinct ML methods (artificial neural network (ANN), SVM, Random Forest, and Gaussian Naï�ve Bayes). The most optimal performance was observed with ANN, achieving an area under the receiver operating characteristic (ROC) curve of 0.761 and an F-measure of 0.714. Experimental results highlighted that the informative locations on the medial tibiofemoral compartment held more valuable information than those on the lateral tibiofemoral compartment for predicting the severity of osteoarthritis.
VOLUME 20,
N° 3
2026
G.B. Joseph et al. [14] aimed to formulate a machine learning-based predictive model for incident radiographic knee OA over an eight-year period, utilizing MRI-based cartilage biochemical composition, knee joint structure, demographics, and clinical predictors, including muscle strength and symptoms. Three distinct models were developed and compared: Model 1, incorporating 112 predictors based on OA risk factors; Model 2, featuring the top ten predictors determined by feature importance score from Model 1, alongside clinical relevance; and Model 3, representing Model 2 without the inclusion of imaging predictors. Model performance was evaluated using the area under the ROC curve derived from hold-out data. The ten-predictor model (Model 2), inclusive of cartilage and meniscus WORMS scores and cartilage T2, exhibited a slightly lower AUC compared to the model with 112 predictors but significantly outperformed the model without MRI predictors. Further studies address the challenge of automatically detecting knee OA through the development of a computer system. The system developed in M. Kotti et al.’s study [15] takes body kinetics as input and provides an output that not only estimates the presence of knee OA as seen in previous literature, but also identifies the most discriminative parameters and outlines a set of rules guiding the decision-making process. This approach bridges the interpretability gap between medical and engineering perspectives. Participants walked on a walkway equipped with two force plates containing piezoelectric three-component force sensors. Parameters related to vertical, anterior-posterior, and mediolateral ground reaction force—such as mean value, push-off time, and slope—were extracted. Random Forest regressors then mapped these parameters via rule induction to determine the degree of knee OA. To enhance generalization ability, a subject-independent protocol was implemented. A different deep learning model, Deep-KOA, developed by D.N.A. Ningrum et al. [16] aimed to predict the risk of knee OA over the next year using non-imagebased electronic medical record data from the previous three years. A feature matrix was constructed based on the three-year sequential history of diagnoses, drug prescriptions, age, and sex. The risk prediction model was developed using a combination of CNN and ANN deep learning methods. Performance evaluation metrics, such as the area under the ROC, sensitivity, specificity, and precision, were employed to assess the efficacy of Deep-KOA. Furthermore, the study explored important features through stepwise feature selection. L.S. Lee et al. [17] provide a comprehensive overview of the current evidence and recent advancements in the application of AI for diagnosing knee OA and predicting outcomes of total knee arthroplasty. The researchers conducted searches in the PubMed and EMBASE databases for articles published in peer-reviewed journals from January 1, 2010, to May 31, 2021. Exclusions comprised non-English language articles and those lacking English translation. A designated 173
Journal of Automation, Mobile Robotics and Intelligent Systems
reviewer assessed the articles for relevance to the research questions and the strength of evidence. In their paper, A.J. Prabhakar et al. [18] aimed to introduce a novel algorithm designed to classify the presence and severity of knee OA using parameters derived from a force plate. Forty-four sway movement graphs were measured for analysis. Various ML algorithms, including K-Nearest Neighbours (KNN), Logistic Regression, Gaussian Naï�ve Bayes, Support Vector Machine (SVM), Decision Tree Classifier, and the Random Forest (RF) Classifier, were applied to the dataset. The proposed method demonstrated a commendable 91 percent accuracy in detecting sway variation, providing valuable support for rehabilitation specialists in objectively identifying a patient’s condition at an early stage and facilitating patient education regarding disease progression. Returning to predictive OA research, A. Tiulpin et al. introduced [19] a multi-modal machine learning-based model for predicting knee OA progression. This model incorporated raw radiographic data, clinical examination results, and the patient’s previous medical history. The validation of this approach was carried out on an independent test set comprising 3,918 knee images from 2,129 subjects. The method demonstrated notable performance with an area under the ROC curve (AUC) of 0.79 and an average precision (AP) of 0.68. In comparison, a reference approach relying on logistic regression produced an AUC of 0.75 and an AP of 0.62. This proposed method holds the potential to significantly enhance the subject selection process for OA drug development trials and contribute to the formulation of personalized therapeutic plans. In a review of the field of knee OA, P.S.Q. Yeoh et al. [20] offer a broad overview of current two-dimensional and three-dimensional CNN approaches. A total of 74 relevant studies about the classification and segmentation of knee OA were reviewed and extracted from the Web of Science database. The review delves into discussions on various state-of-the-art deep learning approaches proposed in this context, with a particular emphasis on highlighting the potential feasibility of employing three-dimensional CNNs in the field of knee OA. The authors conclude by addressing potential challenges and envisioning advancements associated with the adoption of three-dimensional CNNs in this particular research domain. In their paper, C. Ntakolia et al. [21] use an ML pipeline for predicting knee joint space narrowing (JSN) in patients with knee OA utilizing multi-disciplinary data from the OAI database, the proposed methodology encompasses several key components. Firstly, a clustering process is employed to identify distinct groups of individuals with progressing and non-progressing OA. Subsequently, a robust feature selection process is implemented, comprising filter, wrapper, and embedded techniques, aimed at identifying the most informative risk factors contributing to JSN prediction. Lastly, a decision-making process evaluates and compares various classification algorithms 174
VOLUME 20,
N° 3
2026
to select and develop the final prediction model for JSN. Evaluation criteria included the model’s overall performance, robustness, and highest achieved accuracy. Notably, Logistic Regression achieved accuracies of 78.3 percent and 77.7 percent for the left and right leg, respectively, based on a group of 164 risk factors. Additionally, SVM attained accuracies of 78.3 percent and 77.7 percent for the left and right leg, respectively, utilizing a reduced set of 88 and 90 risk factors. The KL grading scale is a highly useful measure of analysis in knee OA severity research. In one study, utilizing deep CNNs alongside the KL grading system, researchers Ganesh Kumar and Goswami [22] evaluated the severity of knee OA using a form of image filtering. Image sharpening enhances image clarity and addresses noise in knee X-ray images. The evaluation used baseline X-ray images from the OAI. Upon analyzing enhanced images generated through the image sharpening process, a mean accuracy of 91.03 percent was achieved. This marks a significant improvement of 19.03 percent over the previous accuracy rate of 72 percent obtained using original knee X-ray images for OA detection with five gradings. The utilization of image sharpening techniques serves to enhance knee joint recognition and facilitate knee KL grading. V.V. Kishore et al. [23] compared the accuracy of models to identify the most effective model for detecting knee OA. In their paper, knee OA served as a clinical scenario to assess twelve transfer learning deep learning models for their ability to detect the grade of knee OA from radiographs. The evaluated models demonstrated a wide range of accuracies, varying from 30 percent to 98 percent in detecting knee osteoarthritis. Notably, MobileNet exhibited the highest accuracy, achieving 98.36 percent. This model showed high levels of training and validation accuracy. Conversely, EfficientNetB7 displayed the maximum loss among the evaluated models. The study suggests that deep learning approaches developed by experienced radiologists and orthopedic specialists could benefit smaller hospitals, enabling them to enhance their learning capabilities and provide improved emergency room services, particularly in scenarios where medical personnel may be scarce. Recent studies demonstrate the potential of machine learning (ML) in automating knee OA diagnosis and prognosis through knee joint localization, classification of OA severity, and prediction of disease progression. Y.X. Teoh et al. [24] provide an overview of knee OA imaging features across various practices for traditional OA diagnosis, alongside recent advancements in image-based ML approaches for knee OA diagnosis and prognosis. While X-ray remains the standard imaging modality for knee OA diagnosis, it is limited in its ability to capture short-term OA changes, prompting recommendations for MRI exploration to uncover hidden OA-related radiomic features. Additionally, ultrasound imaging features should be further explored for improved point-of-care diagnosis. Traditional knee OA diagnosis, reliant on manual interpretation of medical images using the
Journal of Automation, Mobile Robotics and Intelligent Systems
Kellgren-Lawrence (KL) grading scheme, faces challenges of human resources and time constraints, hindering effective OA prevention. AI-aided diagnostic models significantly enhance the efficiency, reproducibility, and accuracy of knee OA diagnosis. Prognostic capabilities have been exhibited by several prediction models in estimating OA onset, deterioration, pain progression, structural changes, and time to total knee replacement incidence. Despite existing research gaps, ML techniques hold promise for addressing demanding tasks, such as early knee OA detection and forecasting future disease events, as well as fundamental objectives like discovering new imaging features and establishing novel OA status measures. Continuous enhancement of ML models may lead to the discovery of innovative OA treatments in the future. Cigdem and Deniz’s literature review [25] provides a comprehensive overview of the current evidence and recent applications of AIin knee OA detection. The review focuses on articles that examine AI’s role in diagnosing and predicting the prognosis of knee OA and in accelerating image acquisition. Each selected study was analyzed for various factors, including code availability, patient and knee count, imaging modalities, covariates, OA grading methods, utilized models, validation techniques, objectives, and outcomes. A.E. Nelson et al. [26] used innovative ML approaches for knee OA phenotyping in order to identify progression phenotypes potentially more responsive to interventions. Publicly available data from the Foundation for the National Institutes of Health’s Osteoarthritis Biomarkers Consortium were utilized, wherein radiographic and pain progression over 48 months were categorized into four mutually exclusive outcome groups (none, both, pain only, radiographic only), alongside an extensive set of covariates. Distance weighted discrimination, direction-projection-permutation testing, and clustering methods were employed to focus on the contrast (z-scores) between individuals progressing by both criteria (“progressors”) and those not progressing by either (“non-progressors”). A separate investigation aimed to determine if a significant loss of surface stiffness could be detected in early OA superficial zone chondrocyte spatial organizations (SCSO) stages. This study [27] examined the relationship between two sensitive early osteoarthritis (OA) markers: atomic force microscopy -measured human articular cartilage surface stiffness and location-matched SCSOs. Subsequently, the study assessed whether current clinical technology could visualize and accurately diagnose the SCSOs using an approved probe-based confocal laser-endomicroscopic system and a RF model. Another study found that AI-aided diagnostic ratings exhibited a stronger association with the overall KL score and the Knee Injury and Osteoarthritis Outcome Score (KOOS). In this study [28], M. Neubauer et al. used seventy-one DICOMs for AI analysis using KOALA software in a study involving subjects recruited from a physiotherapy trial (MLKOA).
VOLUME 20,
N° 3
2026
At baseline, each subject received a knee X-ray along with an assessment using five main scores: Tegner Scale, KOOS, International Physical Activity Questionnaire, Star Excursion Balance Test, and a Six-Minute Walk Test. Clinical assessments were repeated at weeks 6, 12, and 24. Three physicians analyzed the presented X-rays both with and without AI assistance using KL grading. Interrater reliability (IRR) analyses and Spearman’s Correlation Test were conducted to assess the overall KL score for each individual rater with clinical scores. The extent of improvement due to AI varied among individual raters. In their study, J.J. Lee et al. [29] aimed to delineate distinctive pain trajectories among knee OA patients and also to explore the correlation between MRI biomarkers acquired through a three-dimensional CNN and the identified pain trajectories. Data encompassing repeated measures of KOOS pain scores for both knees over a decade-long period were sourced from the OAI, totaling 4,796 subjects. Three-dimensional Double Echo Steady State images of the knee at baseline were employed for image biomarker discovery. To mitigate inherent noise and address missing data, individual pain curves were temporally smoothed using a regression model fitted with orthogonal polynomials of degrees one and two. Standardized estimated parameters were utilized as input for clustering analysis via a Gaussian Naï�ve Bayes mixture model, with the optimal model selected using the Silhouette approach to effectively capture distinct pain patterns. The study’s CNN architecture is a three-dimensional extension of the DenseNet121, with weights initialized using the He initialization. The network was trained to ascertain posterior probabilities of pain trajectory membership, employing the mean squared error function as the loss function for regression. N. Pongsakonpruttikul et al. [30] employed and evaluated the performance of a CNN aimed at aiding orthopedists and radiologists in detecting and categorizing knee OA across various degrees, following the KL classification system. A dataset comprising 1,650 knee joint radiographs (anteroposterior view) was amassed from the Osteoarthritis Initiative’s (OAI) public resources. Two models were constructed: one for distinguishing normal (KL 0-I) from osteoarthritic knees (KL II-IV); The other for classifying severity as normal (KL 0-I), non-severe (KL II), or severe (KL III-IV). Expert supervision was utilized for labeling regions of interest. The AI models were trained utilizing the You Only Look Once version (YOLOv3) detection algorithm. The initial AI model employing YOLOv3 demonstrated an 85 percent accuracy and 81 percent mean average precision in detecting and classifying normal and osteoarthritic knees on plain knee joint radiographs. The subsequent AI model, focusing on severity classification, achieved an overall accuracy of 86.7 percent and mean average precision of 70.6 percent. It is also useful to look at the use of AI in OA diagnosis in other joints. One paper [31] aimed to develop 175
Journal of Automation, Mobile Robotics and Intelligent Systems
and evaluate an AI model for diagnosing Temporomandibular joint (TMJ) osteoarthritis from conebeam computed tomography (CBCT). W.M. Talaat et al. used a dataset comprising 2,737 CBCT images from 943 patients for training and validation. The model, employing a single CNN and a single regression model for object detection, underwent assessment against a model-testing set of 350 images, established by two experienced evaluators using the Diagnostic Criteria for Temporomandibular Disorders. The diagnosis concluded by the evaluators served as the golden reference for comparison. The diagnostic performance of the AI model was then juxtaposed with that of an experienced oral radiologist, revealing statistically higher agreement between the AI diagnosis and the golden reference compared to the radiologist. Another paper centered on TMJ [34] aimed to develop a diagnostic support tool utilizing pretrained models to classify panoramic images of the TMJ into normal and OA cases. The dataset was randomly divided into training, validation, and testing subsets. Pretrained ResNet152 and EfficientNet-B7 models were employed as transfer learning models. The accuracy, specificity, sensitivity, area under the curve, and gradient-weighted class activation mapping of both trained models were assessed. The performance of the trained models was compared to that of dentists. The classification accuracies of ResNet-152 and EfficientNet-B7 were 0.87 and 0.88, respectively, with the trained models demonstrating the highest accuracy in OA classification. Returning to knee OA, M.W. Brejnebøl et al. [32] sought to validate an AI tool for classifying radiographic knee OA severity using weight-bearing, nonfixed-flexion posterior/anterior knee radiographs from a clinical production PACS system. The index test involved ordinal KL grading by an AI tool, two musculoskeletal radiology consultants, two reporting technologists, and two resident radiologists. All readers repeated grading after at least four weeks. The reference test was the consensus of the two consultants. The primary outcome measure was quadratic weighted kappa, with secondary outcomes including ordinal weighted accuracy, multiclass accuracy, and F1-score. In another study, S. Olsson et al. [33] aimed to assess the effectiveness of an AI system in classifying the severity of knee OA using entire image series, without excluding common visual disturbances such as implants, casts, and non-degenerative pathologies. They gathered 6,103 radiographic knee exams conducted at Danderyd University Hospital between 2002 and 2016 and manually categorized them based on the KL grading scale. Subsequently, a CNN of ResNet architecture was trained using PyTorch. The study’s outcomes were evaluated against a test set comprising 300 exams, independently reviewed by two senior orthopedic surgeons who resolved any inter-observer disagreements through consensus sessions. Seeking to address a topic not previously explored in knee OA research, Sozan Ahmed and Ramadhan 176
VOLUME 20,
N° 3
2026
Mstafa [35] developed a novel method to efficiently diagnose and classify knee OA severity based on X-ray images, addressing the issue of classifying knee OA in binary and multiclass categories. The proposed models were categorized into two frameworks: one utilizing pre-trained convolutional neural networks (CNNs) for feature extraction and the other involving fine-tuning of pre-trained CNNs using transfer learning (TL) techniques. Additionally, traditional ML) classifiers were employed to leverage the enriched feature space for improved knee OA classification. The first framework comprised five class-based models employing a proposed pre-trained CNN for feature extraction, principal component analysis for dimensionality reduction, and SVM for classification. In the second framework, adjustments were made to the steps of the first framework, incorporating TL to calibrate the pretrained CNNs for two-, three-, and four-class-based models. In their paper, H. Bonakdari et al. [36] constructed a comprehensive ML model aimed at linking major OA risk factors with serum levels of adipokines and related inflammatory factors at baseline to predict early progression of knee OA among at-risk patients. A total of 677 subjects were selected from the OAI progression sub-cohort for the study. Probability values indicating the likelihood of being structural progressors were generated using a previously developed prediction model, which incorporated five baseline structural knee features, including two X-rays and three magnetic resonance imaging variables. To identify the most influential variables among the 47 studied in relation to progressive knee OA, the researchers utilized ML feature classification methodology. Among the five supervised ML algorithms tested, SVM exhibited the highest accuracy and was chosen for gender-based classifier development. Feature selection analysis revealed that age, BMI, and the ratios of CRP/MCP-1 and leptin/CRP were the most crucial variables for predicting OA structural progressors in both genders. A paper by S.B. Kwon et al. [37] used patientreported outcome measures to identify gait features related to knee OA and develop regression models using ML algorithms to estimate knee OA severity. Knee OA severity was assessed using the Western Ontario and McMaster Universities Osteoarthritis Index (WOMAC). Linear regression models and a Random Forest regression were employed to estimate WOMAC scores for patients. A range of features were selected from each joint, including 12 from the hip, 1 from the pelvis, 17 from the knee, 9 from the ankle, 1 from the foot, and 3 from spatiotemporal parameters. Gait analysis features were utilized to predict knee OA severity, contributing to the development of an objective estimation model for knee OA severity. A.S. Mohammed et al. [38] conducted binary classification to detect the presence or absence of knee OA and classify the severity of knee OA into three classes. This paper proposed the application of six pre-trained deep neural network (DNN) models, including VGG16,
Journal of Automation, Mobile Robotics and Intelligent Systems
VGG19, ResNet101, MobileNetV2, InceptionResNetV2, and DenseNet121, for knee OA diagnosis using images sourced from the OAI dataset. For comparative analysis, experiments were performed on three datasets— Dataset I, Dataset II, and Dataset III—with five, two, and three classes of knee OA images. The ResNet101 DNN model achieves maximum classification accuracies of 69 percent, 83 percent, and 89 percent for Dataset I, Dataset II, and Dataset III, respectively. The results demonstrate improved performance compared to existing literature. Continuing research into computer-aided diagnosis (CAD), A. Tiulpin et al. [39] present a novel transparent CAD method based on the Deep Siamese CNN to automatically score knee OA severity according to the KL grading scale. The method is trained using data exclusively from the Multicenter Osteoarthritis Study and validated on a randomly selected subset of 3,000 subjects (5,960 knees) from the OAI dataset. The proposed method achieves a quadratic Kappa coefficient of 0.83 and an average multiclass accuracy of 66.71 percent compared to annotations provided by a committee of clinical experts. Additionally, a radiological OA diagnosis area under the ROC curve of 0.93 is reported. S. Abd El-Ghany, M. Elmogy, and A.A. Abd El-Aziz’s [40] model aims to classify the severity of knee OA diseases through multi-classification and binary classification. Their paper proposed a fine-tuned knee OA diagnosis model utilizing the DenseNet169 deep learning) technique to enhance the efficiency of knee OA diagnosis. The model was designed to accurately localize peripheral opacities, diffuse distribution, and vascular thickening, thereby assisting clinicians in understanding the primary causes of knee OA. Pre-processing of the OAI dataset involved artifact removal, resizing, contrast handling, and normalization techniques. The proposed model underwent evaluation and comparison with recent classifiers. In multi-classification, the DenseNet169 model demonstrated performance metrics of 95.93 percent accuracy, 88.77 percent sensitivity, 95.41 percent specificity, 85.8 percent precision, and 87.08 percent F1-score.
VOLUME 20,
N° 3
2026
during training. This comprehensive and systematic approach to data collection and preprocessing provides a solid foundation for the subsequent phases of model development and evaluation. The primary deep learning architecture chosen for model selection is the pre-trained EfficientNet B5 model, tailored explicitly for X-ray detection. The model parameters were adapted to the distinctive features of X-ray images. Leveraging transfer learning, the pre-trained EfficientNet B5 will capitalize on the knowledge gained from general image datasets, enhancing its ability to discern nuanced patterns in knee X-rays. Additionally, alternative ML models, including Random Forest and KNN, chosen for their simplicity and interpretability, are utilized for benchmarking purposes. EfficientNet V2B3, a variant of the primary model, will also be deployed to compare its performance. A primary CNN (Conv2D) with MaxPooling2D layers is used to diversify the comparison, offering insights into the trade-offs between model complexity and performance. This comprehensive model selection strategy aims to evaluate the efficacy of the chosen deep learning architecture against various benchmarks, providing a thorough understanding of the model’s strengths and potential improvements. In the model training phase, the pre-trained EfficientNet B5 underwent fine-tuning on the designated knee OA training dataset using categorical cross-entropy as the loss function for multi-class classification. Early stopping was implemented to mitigate overfitting risks. An overall view of the proposed work is given in Figure 2. Random Forest and KNN were trained on flattened X-ray images, utilizing cross-validation for robust model evaluation. EfficientNet V2B3 and Conv2D models were trained with configurations similar to EfficientNet B5. In the model evaluation phase, a comprehensive set of evaluation metrics, was employed to assess the performance of the trained models on the
3. Proposed Methodology
In the initial phase of data collection and preprocessing, a diverse and well-annotated dataset of knee X-ray images, labeled with OA severity grades, was carefully curated. To mitigate potential biases, particular attention was given to ensuring a balanced distribution across all severity levels. The dataset was then partitioned into training, validation, and test sets while maintaining representative distributions in each subset. During preprocessing, the images were standardized and resized to ensure a uniform format suitable for model input. To improve the model’s generalization capabilities, augmentation techniques, such as rotation, flipping, and scaling, were applied. Additionally, pixel values were normalized to a specific range to support efficient model convergence
Figure 2. Flow chart detailing the data training and validation stages of the methodology 177
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
grow the model’s depth, width, and resolution. It has been described as a new breed of CNNs that, with fewer parameters, are able to achieve very high accuracy. The Random Forest model, known for its interpretability, was chosen to assess its capability to distinguish severity levels. The KNN algorithm, valued for its simplicity, was included as a viable option for knee OA detection. Additionally, the EfficientNet V2B3 model and traditional Conv2D and MaxPooling2D architectures were employed to understand model versatility comprehensively. Highlights of Model Architecture:
• MBConv Blocks: Mobile Inverted Bottleneck Convolution (MBConv) blocks form the core part of the architecture; • Activation Function: Swish (smooth; non-monotonic); • Squeeze-and-Excitation (SE): Applied in every block to enhance the recalibration of features; • Input resolution: 456×456 pixels;
• Total layers: ~570 layers (not all of them are convolution; includes batch norm, pooling,etc.); • Number of Params: ~30 million (29.34M); and
Figure 3. Sample images of each class of knee OA severity
178
validation set, including accuracy, precision, recall, F1-score, and confusion matrices. The final evaluation was based on the independent test set, providing a robust assessment of the model’s effectiveness in accurately classifying and grading knee OA severity. The evaluation initially took place on the validation set, allowing for the fine-tuning of hyperparameters and guarding against overfitting. A diverse dataset comprising knee X-ray images with varying degrees of OA severity was meticulously selected to achieve comprehensive evaluations. The dataset was annotated and graded by medical professionals to ensure accuracy. The dataset is already a benchmark dataset, which can be found in: Chen, Pingjun (2018), “Knee Osteoarthritis Severity Grading Dataset”, Mendeley Data, V1, doi: 10.17632/56rmx5bjcr.1. Sample data from each class of the dataset is given in Figure 3. Standard preprocessing techniques, including resolution standardization and image augmentation, were applied to enhance the robustness of the models. The choice of machine learning models was driven by a desire to explore diverse architectures. EfficientNet B5, renowned for its success in image classification tasks, was selected for its potential in medical image analysis. EfficientNet B5 is a member of the EfficientNet family that uses a compound scaling method to
• Number of MBConv Blocks: 39 MBConv blocks with differing kernel sizes (3×3 or 5×5) and expansion ratios. Fine-tuning Procedure:
• Pretrained on: ImageNet;
• Layers Frozen: During early phases of training, the initial layers were frozen;
• Unfrozen Layers: The last 3 blocks were unfrozen to be fine-tuned for the purpose of specific tasks; • Optimizer: Adam, LR=1e-4, ReduceLROnPlateau; • Batch size: 16;
• Epochs: 50 with early stopping.
A comparative analysis with well-known architecture is given in Table 1 for better justification.
4. Results and Analysis
The application of ML models for the detection of knee OA using X-ray images yielded compelling outcomes. Notably, the EfficientNet B5 model demonstrated exceptional performance, achieving an impressive accuracy of 95.84 percent. This result highlights the efficacy of leveraging state-of-the-art neural network architectures for medical image analysis. The EfficientNet B5 model’s ability to discern intricate features within X-ray images contributed significantly to its superior performance in accurately grading the severity of knee OA. Sample images and a predicted image sample are given in Figure 4 and Figure 5, respectively.
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Table 1. Comparative analysis of models Model
Parameters
Input Size
Conv Layers
Key Features
EfficientNet-B5
~29.3M
456×456
MBConv (39)
Compound scaling, SE blocks, Swish
MobileNetV2
~3.4M
224×224
~53
Lightweight, depthwise separable conv
ResNet50 VGG16
~25.6M ~138M
224×224 224×224
50 13
Residual blocks, skip connections
Deep stack, large fully connected head
Figure 6. Output Graph
Figure 7. Confusion Matrix Efficient Net B5 Figure 4. Sample Images
Figure 5. Predicted Image
The impressive outcomes from applying ML models for knee OA detection using X-ray images underscore the significance of advanced neural network architectures. The model’s ability to discern intricate features within X-ray images was crucial in achieving superior performance, showcasing its potential for accurate severity grading in knee osteoarthritis cases. The generated graph is given in Figure 6. Figures 7, 8, and 9 show the confusion matrices of the algorithms under comparison. Comparatively, the Random Forest model exhibited a moderate accuracy of 45 percent, indicating a reasonable capability to distinguish between different severity levels. While not as high as the EfficientNet B5, this accuracy suggests a pragmatic performance level for severity classification. The KNN algorithm achieved an accuracy of 41 percent, highlighting its potential as a viable option for knee OA detection, albeit with room for improvement. The EfficientNet V2B3 model demonstrated competitive performance with an accuracy of 65.78 percent, further emphasizing the versatility of different EfficientNet variants in handling medical image classification tasks. This result positions EfficientNet as a robust choice for knee OA diagnosis. On the other hand, the Conv2D and MaxPooling2D models, repre-
179
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
frameworks that facilitate seamless integration, addressing privacy, security, and interoperability concerns within healthcare systems. Furthermore, continuous exploration and incorporation of the latest advancements in ML and medical imaging are essential to staying at the forefront of technological progress, ensuring sustained improvement in the performance and efficiency of knee OA detection models.
Figure 8. Confusion Matrix KNN
Figure 9. Confusion Matrix Random Forest senting traditional CNN architectures, achieved an accuracy of 39.56 percent, providing insights into the effectiveness of these established approaches. The discussions center around the remarkable accuracy achieved by the EfficientNet B5 model, showcasing its potential for accurate knee OA severity grading. The moderate accuracies of Random Forest, KNN, EfficientNet V2B3, Conv2D, and MaxPooling2D models offer a nuanced understanding of their capabilities. Future considerations may involve refining model parameters, exploring ensemble approaches, and addressing potential limitations to further enhance the practical implementation of these models in clinical settings.
5. Conclusion and Future Work
180
Our research contributes to the scientific discourse surrounding knee OA detection and underscores the transformative potential of ML in revolutionizing healthcare practices. The journey towards a more technologically advanced and patient-centric healthcare system is underway, and the successful fusion of ML with medical imaging marks a significant step forward in this transformative process. Future research should focus on developing guidelines and
AUTHORS R. Gokulapriya – Department of Computer Science and Engineering, School of Engineering and Technology, CHRIST (Deemed to be University), Bangalore Kengeri Campus, Bangalore, India, email: r.gokulapriya@ christuniversity.in. M. Balamurugan – Department of Computer Science and Engineering, School of Engineering and Technology, CHRIST (Deemed to be University), Bangalore Kengeri Campus, Bangalore, India, email: balamurugan.m@ christuniversity.in. Rakoth Kandan Sambandam – Department of Computer Science and Engineering, School of Engineering and Technology, CHRIST (Deemed to be University), Bangalore Kengeri Campus, Bangalore, India, email: rakoth.kandan@christuniversity.in. Divya Vetriveeran* – Department of Computer Science and Engineering, University College of Engineering Thirukkuvalai, Anna University, Thirukkuvalai, Tamil Nadu, India, email: vdivya.3793@gmail.com *Corresponding author
References
[1] J.Y. Jang et al., “Study of the Efficacy of Artificial Intelligence Algorithm- Based Analysis of the Functional and Anatomical Improvement in Polynucleotide Treatment in Knee Osteoarthritis Patients: A Prospective Case Series,” Journal of Clinical Medicine, 2022; doi:10.3390/ jcm11102845. [2] S. Mohammadi et al., “Artificial intelligence in osteoarthritis detection: A systematic review and meta-analysis,” Osteoarthritis and Cartilage, 2023, pp. 241-253. [3] H.A. Alshamrani, M. Rashid, S.S. Alshamrani, A.H.D. Alshehri, “Osteo-NeT: An Automated System for Predicting Knee Osteoarthritis from X-ray Images Using Transfer-Learning-Based Neural Networks Approach,” Healthcare, 2023; doi:10.3390/healthcare11091206. [4] P. Chen et al., “Fully automatic knee osteoarthritis severity grading using deep neural networks with a novel ordinal loss,” Computerized Medical Imaging and Graphics, vol 75, 2019, pp. 84-92. [5] B. Abdelbasset et al., “A decision support tool for early detection of knee OsteoArthritis using X-ray imaging and machine learning: Data from the Osteoarthritis Initiative,” Computerized Medical Imaging and Graphics, vol. 73, 2019, pp. 11-18.
Journal of Automation, Mobile Robotics and Intelligent Systems
[6] J. Song and R. Zhang, “A novel computer-assisted diagnosis method of knee osteoarthritis based on multivariate information and deep learning model,” Digital Signal Processing, 2023; doi:10.1016/j.dsp.2022.103863. [7] N. Lazzarini et al., “A machine learning approach for the identification of new biomarkers for knee osteoarthritis development in overweight and obese women,” Osteoarthritis and Cartilage 2014-2021 (epub 2017); doi:10.1016/j. joca.2017.09.001. [8] S.J.O. Rytky et al., “Automating three-dimensional osteoarthritis histopathological grading of human osteochondral tissue using machine learning on contrast-enhanced micro-computed tomography,” Osteoarthritis and Cartilage vol. 28, no. 8, 2020, pp. 1133-1144. [9] V. Pedroia et al., “Diagnosing osteoarthritis from T2 maps using deep learning: an analysis of the entire Osteoarthritis Initiative baseline cohort,” Osteoarthritis and Cartilage vol. 27, no. 7, 2019, pp. 1002-1010. [10] A. Haseeb, et al., “Knee osteoarthritis classification using x-ray images based on optimal deep neural network,” Computer Systems Science and Engineering, vol. 47, no. 2, 2023, pp. 2397–2415. [11] A. Rehman et al., “Transfer Learning- Based Smart Features Engineering for Osteoarthritis Diagnosis From Knee X-Ray Images,” IEEE Access, vol. 11, 2023, pp. 71,326-71,338. [12] Y.X. Teoh et al. “Stratifying knee osteoarthritis features through multi-task deep hybrid learning: Data from the osteoarthritis initiative,” Computer Methods and Programs in Biomedicine, 2023; doi:10.1016/j.cmpb.2023.107807. [13] Y. Du, J. Shan, and M. Zhang, “Knee osteoarthritis prediction on MR images using cartilage damage index and machine learning methods,” 2017 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2017, pp. 671-677; doi: 10.1109/BIBM.2017.8217734. [14] G.B. Joseph et al., “Machine learning to predict incident radiographic knee osteoarthritis over 8 Years using combined MR imaging features, demographics, and clinical factors: data from the Osteoarthritis Initiative,” Osteoarthritis and Cartilage, vol. 30, no. 2, 2022, pp. 270-279. [15] M. Kotti et al., “Detecting knee osteoarthritis and its discriminating parameters using random forests.” Medical Engineering & Physics, vol. 43, 2017 pp. 19-29. [16] D.N.A. Ningrum et al., “A Deep Learning Model to Predict Knee Osteoarthritis Based on Nonimage Longitudinal Medical Record,” Journal of Multidisciplinary Healthcare, vol. 14, 2021, pp. 2477-2485. [17] L.S. Lee et al., “Artificial intelligence in diagnosis of knee osteoarthritis and prediction of arthroplasty outcomes: a review,” Arthroplasty, 5 Mar. 2022; doi:10.1186/s42836-022-00118-7.
VOLUME 20,
N° 3
2026
[18] A.J. Prabhakar et al.,“Use of Machine Learning for Early Detection of Knee Osteoarthritis and Quantifying Effectiveness of Treatment Using Force Platform,” Journal of Sensor and Actuator Networks, 2022; doi:10.3390/jsan11030048 [19] A. Tiulpin et al., “Multimodal Machine Learning-based Knee Osteoarthritis Progression Prediction from Plain Radio-graphs and Clinical Data,” Sci Rep, 2019; doi:10.1038/s41598-01956527-3. [20] P.S.Q. Yeoh et al., “Emergence of Deep Learning in Knee Osteoarthritis Diagnosis”, Computational Intelligence and Neuroscience, 2021; doi:10.1155/2021/4931437. [21] C. Ntakolia et al., “A machine learning pipeline for predicting joint space narrowing in knee osteoarthritis patients,” 2020 IEEE 20th International Conference on Bioinformatics and Bioengineering (BIBE), 2020, pp. 934-941. [22] M. Ganesh Kumar and A.D. Goswami, “Automatic Classification of the Severity of Knee Osteoarthritis Using Enhanced Image Sharpening and CNN,” Applied Sciences, , 2023; doi:10.3390/ app13031658. [23] V.K. V, V. Kalpana, G.H. Kumar, “Evaluating the efficiency of deep learning models for knee osteoarthritis prediction based on Kellgren-Lawrence grading system,” e-Prime-Advances in Electrical Engineering, Electronics and Energy, , 2023; doi:10.1016/j.prime.2023.100266. [24] Y.X. Teoh et al., “Discovering Knee Osteoarthritis Imaging Features for Diagnosis and Prognosis: Review of Manual Imaging Grading and Machine Learning Approaches,” Journal of Healthcare Engineering, 2022; doi:10.1155/2022/4138666. [25] O. Cigdem and C.M. Deniz, “Artificial intelligence in knee osteoarthritis: A comprehensive review for 2022,” Osteoarthritis Imaging, 2023; doi:10.1016/j.ostima.2023.100161. [26] A.E. Nelson et al., “A machine learning approach to knee osteoarthritis phenotyping: data from the FNIH Biomarkers Consortium,” Osteoarthritis and Cartilage, vol. 27, no. 7, 2019, pp. 9941001. [27] M. Tschaikowsky et al., “Proof-of-concept for the detection of early osteoarthritis pathology by clinically applicable endomicroscopy and quantitative AI-supported optical biopsy,” Osteoarthritis and Cartilage, vol 29, no. 2, 2021, pp. 269-279. [28] M. Neubauer M et al., “Artificial-Intelligence-Aided Radiographic Diagnostic of Knee Osteoarthritis Leads to a Higher Association of Clinical Findings with Diagnostic Ratings,” Journal of Clinical Medicine, 2023; doi:10.3390/ jcm12030744. [29] J.J. Lee et al., “Can AI Predict Pain Progression in Knee Osteoarthritis Subjects From Structural
181
Journal of Automation, Mobile Robotics and Intelligent Systems
MRI,” Osteoarthritis and Cartilage, April 2019;doi:10.1016/j.joca.2019.02.036. [30] N. Pongsakonpruttikul et al., “Artificial intelligence assistance in radiographic detection and classification of knee osteoarthritis and its severity: a cross-sectional diagnostic study,” European Review for Medical and Pharmacological Sciences, vol. 26, no. 5, 2022, pp. 1549-1558. [31] W.M. Talaat et al., “An artificial intelligence model for the radiographic diagnosis of osteoarthritis of the temporomandibular joint,” Scientific Reports, 2023;doi:10.1038/s41598-023-43277-6. [32] M.W. Brejnebøl et al., “External validation of an artificial intelligence tool for radiographic knee osteoarthritis severity classification,” European Journal of Radiology, 2022; doi:10.1016/j. ejrad.2022.110249. [33] S. Olsson, et al., “Automating classification of osteoarthritis according to Kellgren-Lawrence in the knee using deep learning in an unfiltered adult population,” BMC Musculoskeletal Disorders, 2021; doi:10.1186/s12891-021-04722-7. [34] W. Jung et al., “Deep learning for osteoarthritis classification in temporomandibular joint,” Oral Diseases, vol. 29, no. 3, 2023, pp. 1050-1059. [35] S.M. Ahmed and R.J. Mstafa, “Identifying Severity Grading of Knee Osteoarthritis from X-ray
182
VOLUME 20,
N° 3
2026
Images Using an Efficient Mixture of Deep Learning and Machine Learning Models,” Diagnostics (Basel), , 2022; doi:10.3390/diagnostics12122939. [36] H. Bonakdari et al., “A warning machine learning algorithm for early knee osteoarthritis structural progressor patient screening,” Therapeutic Advances in Musculoskeletal Disease, 23 Feb. 2021, doi:10.1177/1759720X21993254. [37] S.B. Kwon et al., “A machine earning-based diagnostic model associated with knee osteoarthritis severity,” Scientific Reports, 2020; doi:10.1038/s41598-020-72941-4. [38] A. S. Mohammed et al., “Knee Osteoarthritis Detection and Severity Classification Using Residual Neural Networks on Preprocessed X-ray Images,” Diagnostics (Basel), 2023; doi:10.3390/ diagnostics13081380 [39] A. Tiulpin, et al., “Automatic Knee Osteoarthritis Diagnosis from Plain Radiographs: A Deep Learning-Based Approach,” Scientific Reports, 2018; doi:10.1038/s41598-018-20132-7 [40] S. Abd El-Ghany, M. Elmogy, A. A. Abd El-Aziz, “A fully automatic fine tuned deep learning model for knee osteoarthritis detection and progression analysis,” Egyptian Informatics Journal, vol 24, no. 2, 2023, pp. 229-240.
VOLUME 20, N° 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
Swinnetplus: A Novel Deep Learning Approach Using Swin Transformer For Enhancing Brain Tumor Segmentation Submitted: 6th June 2025: accepted: 2nd February 2026
Vikash Verma, Pritaj Yadav DOI: 10.14313/jamris-2026-049
1. Introduction
Abstract The segmentation of brain tumors based on three-dimensional (3D) magnetic resonance imaging (MRI) is a basic but difficult challenge when analyzing medical images due to the large differences in size, shapes, and appearances of tumors, and the high visual similarity of tumor with the normal tissues. To overcome such difficulties, the paper suggests SwinNet Plus, the new hybrid architecture of deep learning that combines the complementary advantages of Swin Transformers and Convolutional Neural Networks (CNNs) to strengthen accurate brain tumor localization. The SwinNetPlus is made to capture both large-scale contextual features and small-scale spatial features with an encoder-decoder architecture. The encoder uses the Swin Transformer based on Enhanced Local Self-Attention (ELSA) to model long-range contextual features and maintain local structural features. In order to further improve feature discriminability, the Swin Transformer adds a channel squeeze-and-excitation block to an adaptive recalibration channel-blockwise responses to multimodal MRI input. Moreover, two special modules are provided to enhance the accuracy in segmentation. The Spatial Local Attention Module (SLAM) has a focus on small local details and tumor edges, which allows precise definition small and irregular tumor areas. The Dense Cross-Multiplication Link is an enhancement of a cross-multiplication link that encourages consistent multiscale fusion, implying that encoder features and decoder features are densely integrated to eliminate irrelevant activation and noise transmission. The suggested SwinNetPlus model is highly tested on BraTS 2021 dataset, with 1,251 multimodal 3D brain MRI scans. The experimental outcomes also indicate that SwinNetPlus has a Dice similarity coefficient of 92.69 percent and Hausdorff distance of 5.34 mm which is superior to the performance of some current methods of 3D brain tumor segmentation. These findings underscore the strength and efficacy of SwinNetPlus in the correct segmentation of brain tumors under magneticity tough circumstances, thus it becomes an attractive instrument in clinical determination support systems.
Segmentation of brain tumors using magnetic resonance imaging (MRI) is an important aspect of clinical diagnosis, treatment planning and monitoring of disease progression. Radiotherapy planning, surgery, and treatment strategies need to be precise in sub regions of tumors, including the entire tumor, tumor core and enhancing tumor. Nevertheless, manual brain tumor segmentation is tedious, subjective, and prone to inter-observer variability, which has led to the development of automated and reliable brain tumor segmentation methods. In the last ten years, there has been a spectacular success of the use of deep learning-based approaches , especially convolutional neural networks (CNNs), in medical image segmentation.With their encoder-decoder design and skip connections which tend to combine low-level detail of the spatial features with higher-level semantic features, architectures such as U-Net and its variants have become standard baselines. Although CNN-based models have been quite successful, they have the perilous limitation of having local receptive fields that limit the extent to which they can capture long range dependencies and global relationships. This drawback is particularly problematic in brain tumor segmentation, where tumors differ greatly in shape, size, location and appearance and where tumor boundaries tend to be similar to the neighboring healthy tissues. Figure 1 shows the sub-regions that help healthcare professionals judge differences in a tumor and choose the best treatment approach to solve these problems, medical image analysis has recently seen the introduction of transformer-based architectures. Transformers, which were initially designed to work with natural languages, use selfattention to model global contextual relationships. Vision Transformers (ViTs) and their variants have demonstrated promising performance in allowing long-range feature interactions. Nevertheless, regular ViTs are computationally costly and need extensive training information, which is not feasible with high-resolution 3D medical images. In response to this, the Swin Transformer provides a more efficient alternative with shifted window-based self-attention hierarchy. The design enables the model to capture both the local and global contexts without expending
Keywords: SwinNetPlus, Brain Tumor Segmentation, Swin Transformer, Spatial Local Attention Module (SLAM), Dense Cross-Multiplication Link (DCML), Medical Image Analysis
Open Access. © 2026 Vikash Verma and Pritaj Yadav, published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP. This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 License.
183
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
and a Hausdorff distance of 5.34 mm. These findings indicate that the model is capable of splitting tumors in the brain even in difficult imaging conditions. In general, SwinNetPlus offers a powerful and efficient platform to segment brain tumors automatically, which has great potential to be applied in practice and contribute to the development of state-of-the-art medical image processing.
2. Related Work Figure 1. Illustration depicting various sub-regions for brain tumor segmentation [1]
184
too much computational complexity. As a result, medical image segmentation has been widely studied, with Swin Transformer-based models gaining growing popularity. Still, current Swin-based methods tend to not maintain fine-spatial resolution, particularly when small tumor objects are involved, and can exhibit inconsistency across features in multiscale fusion when coupled with encoder decoders. In this article, we present SwinNetPlus, a new hybrid deep learning model capable of combining the benefits of Swin Transformers and CNNs in the powerful segmentation of brain tumors in 3D MRI scans. The proposed model is aimed at overcoming the major issues of existing models, which include the problem of changes in scale, unclear delineation of tumors, and the visual resemblance of tumors and healthy tissues. The encoder-decoder model adopted by SwinNetPlus is capable of integrating a global contextual representation with high spatial localization. One of the biggest contributions of SwinNetPlus is the dedicated introduction of two new modules, namely the Spatial Local Attention Module (SLAM) and the Dense Cross-Multiplication Link (DCML). SLAM increases the sensitivity of the network to local spatial characteristics as it focuses on boundary and texture information, which are important in precise tumor delineation. This module allows the model to pay attention to minute changes that are likely to be missed by traditional attention systems. DCML, favors consistency in features and an overall healthy flow of information by blurring multiscale features through cross-multiplicative interactions, thus, minimizing noise and eliminating unhelpful activation. The SwinNetPlus encoder is designed based on Enhanced Local Self-Attention (ELSA) transformer blocks that enhance local feature extraction and maintain global contextual features. These features include channel squeeze-and-excitation blocks to re-calibrate channel-wise feature responses and improve the discriminative feature representation of various MRI modalities. Therefore, the proposed SwinNetPlus architecture is effective on the BraTS 2021 dataset that includes 1,251 multimodal 3D brain MRI scans. Experimental findings show that SwinNetPlus performs better than various state-of-the-art 3D segmentation algorithms with a Dice score of 92.69%
The automated segmentation of brain tumors is a promising research topic because it has been found to be essential in clinical decision-making and treatment planning. First generation research used early computational methods used manual features and classical machine learning models. These methods, however were very sensitive to noise, changes in intensity and anatomical variability of the MRI scans. The introduction of deep learning in the field has superseded most traditional methods with data-driven ones as they are more effective at learning features.
2.1. CNN-Based Segmentation Techniques Fully Convolutional Networks (FCNs) were a breakthrough in the field, because they allowed endto-end semantic segmentation without fully connected layers [2]. U-Net (building on this concept) proposed a symmetric encoder-decoder network with skip connections, which maintain spatial resolution and learn high-level semantic features [3]. U-Net and its variants have emerged as hallmarks of biomedical image segmentation. Several variants include UNet++ [4], which added internal skip connections meaning the encoder and decoder features lessened the semantic disparity, leading to more accurate segmentation. In the case of volumetric medical data, 3D CNNs were designed to better utilize spatial continuity. The 3D U-Net was an extension of the U-Net framework to three dimensions and showed better results in volumetric segmentation tasks [5]. H-DenseUNet was a hybrid CNN architecture that used dense connectivity alongside encoder decoder architectures, and improved feature reuse and gradient flow [6]. Although effective, CNN based models have a small receptive field, which does not allow them to capture long range dependencies that are important in the process of contrasting tumors against visually similar normal tissues.
2.2. Transformer-Based Vision Models Transformers, which were proposed as sequence models in natural language processing [7], were subsequently applied to visual tasks with the establishment of the Vision Transformer (ViT), which estimates the relationships globally via self-attention mechanisms [8]. Despite the excellent representational abilities of ViTs, the high cost of computation and the amount of data needed challenged the applicability to medical imaging. To address these drawbacks, hierarchical transformers like the Swin Transformer were proposed and utilized. These transformers use shifted window-based self-attention to balance both local and global context in an efficient way [9]. The Swin Transformer has since
Journal of Automation, Mobile Robotics and Intelligent Systems
become a standard in image fusion models [10] as well as medical image segmentation models. SwinPA-Net [11] used multiscale feature pyramids to enhance the output of segmentation, whereas Swin-Unet used a U-Net-like design with pure transformer blocks [12]. Pure transformer designs, however, tend to have limited fine-grained localization of boundaries because they lack inductive bias on localizing spatial detail.
2.3. CNN-Transformer Hybrid Architectures Hybrid architectures have been suggested in order to integrate the advantages of CNNs and transformers. TransUNet merged ViT encoders with CNNbased decoders showing better results than the pure CNN models [13]. UNETR used transformers as 3D medical image segmentation encoders, which utilized global context and volumetric consistency [14]. UNETR++ improved on this concept by making the 3D segmentation tasks more efficient and more accurate [15]. SwinBTS presented Swin Transformer blocks to analyze multiple modalities of MRI in the field of brain tumor segmentation, which brought positive results on the BraTS dataset [16]. NestedFormer was a proposed model of more attentive mechanisms that could be used to better integrate multimodal information by using modality [17]. Swin UNETR used Swin Transformer encoders with UNETR-style decoding to enhance the tumor boundary delineation [18]. Although these techniques enhanced contextual modeling, in many cases, they depended on global attention processes that can miss local structures important to accurately segmenting the boundaries of tumors. 2.4. Focusing Machinery and Character Strengthening Both segmentation and robustness of segmentation have been highly exploited by attention-based feature refinement. The Squeezeand-Excitation (SE) networks proposed channelwise attention to refine the feature responses, and enhance the representational power with minimum overhead [19]. Enhanced Local Self-Attention (ELSA) was suggested to improve local context modeling in transformer architectures, fixing the loss of spatial detail using global attention [20]. Recent studies that incorporated ELSA in to Swin Transformers showed better performance in segmentation but still had the issue of multiscale feature consistency and noise suppression [21].
2.5. Research Gap Although there have been major advancements, there are still a number of shortcomings in the methods for segmenting brain tumors. CNN based methods are not powerful in long-range dependency modeling, and transformer based methods usually do not have local sensitivity to features. Hybrid models enhance contextual awareness but often are affected by feature inconsistency when performing multiscale fusion and inadequate boundary refinement. In addition, Swin-based architectures do not sufficiently support the contemporaneous
VOLUME 20,
N° 3
2026
requirement of local spatial attention, strong interaction of multiscale features or noise reduction in complicated MRI scenarios. In order to close these gaps, a coherent framework that clearly improves local feature discrimination, dense and consistent interaction of features across scales, and effective global context integration is needed. The mentioned constraints drive the proposed SwinNetPlus, which use the Spatial Local Attention Module (SLAM) to model the fine-grained boundaries and a Dense Cross-Multiplication Link (DCML) to combine features and decrease noise.
3. Proposed Method
3.1. Overview The proposed SwinNetPlus architecture is formed by an encoder and a decoder, as you can see in Figure 2. The main structure of the model uses the Swin Transformer to identify different levels of features from multimodal MRI scans. To make the feature representation stronger, we place the Enhanced Local Self-Attention (ELSA) blocks in the Swin Transformer. Additionally, the two new modules DCML and SLAM are included to help the model recognize both small-scale and wide-scale information. The input for SwinNetPlus is an MRI image X, where H, W, D and C denote the height, width, depth and number of modalities, respectively. The input is broken into separate patches and then goes through the transformer-based encoder. Swin Transformer blocks are used in the encoder and enhanced with ELSA modules to produce feature representations at different scales. The features from the encoder are passed through skip connections to the decoder path. The decoder contains a spatial Squeeze-and-Excitation (sSE) CNN module that combines multi-resolution features and outputs the final segmentation map. Using transformer-based encoding and convolution-based decoding, SwinNetPlus can use both global attention and precise local details. 3.2. Swin Transformer Encoder The proposed Swin Transformer-based encoder consists of four stages, and each stage is made up of two Swin Transformer blocks. Swin Transformer employs a special self-attention mechanism that helps improve speed and maintain the spatial arrangement. Window-based Multi-Head Self-Attention (W-MSA) units are followed by Shifted Window-based MultiHead Self-Attention (SW-MSA) units in each layer, as shown in Figure 3(a). This progression help ensure that interactions between windows are effective without much extra cost in computation. In the first three stages, ELSA blocks are included in the Swin Transformer. The blocks use Hadamard attention and ghost heads to impove local feature extraction by sharpening the attention inside the windows. This design makes it possible for the model to extract full and localized features from the given image. At this stage, windows are superimposed to improve the features even more. Self-attention is performed within every
185
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 2. SwinNetPlus Architecture
Figure 3. The Swin Transformer architecture comprises distinct components. (a) Self-attention computations within adjacent layers l and l +1 utilize both regular and shifted windows. (b) Further insights into the intricacies of the Swin Transformer are provided
186
window helping to create rich contextual information with less computational cost. By sharing features, this strategy allows nearby windows to help to each other learn new features. Swin Transformer encodes multilevel features by gradually lowering the size of the input and adding more channels. Every time, a patch merging layer cuts the image’s size in half and increases the number of channels by two. At the beginning, C is set to 48 and then doubles at each level: (H/s)× (W/s)× C, (H/2s)× (W/2s)× 2C, (H/4s)× (W/4s)× 4C and (H/8s) × (W/8s) × 8C. In this step, s is the patch size which is the number of patches taken from the input image that do not overlap. After the patches are flattened, they are sent through a linear embedding layer that projects them to C features and then the transformer blocks take over. It takes a lot of computing power for conventional transformers because they use global attention on all the tokens. The Swin Transformer addresses this by using a
window-based system that both simplifies the model and keeps its representation strong. Both W-MSA and SW-MSA help windows share data and keep the system efficient. Because of its good balance between speed and effectiveness, the Swin Transformer is used in our SwinNetPlus model as the main component for extracting features from multimodal MRI data.
3.3. Enhanced Local Self Attention (ELSA) Adding the ELSA block to the Swin Transformer gives it even greater ability to handle long-range connections. As shown in Figure 4, adding the ELSA block to the Swin Transformer’s structure allows it to better perform on dense prediction tasks [20]. The ELSA transformer block enhances the extraction of local features more than traditional Local SelfAttention (LSA) modules do, especially when it is used with dynamic filters in the Swin Transformer [3]. ELSA relies on Hadamard attention which uses Hadamard
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
attention structure and integrating it with the Transformer structure, while Figure 2 shows this implementation (as explained by Eq. (3)). yᵢ = MLP(LN(ŷᵢ)) + ŷᵢ
Figure 4. Swin Transformer with ELSA Block [20] products to preserve high-level relationships and direct attention to nearby locations. Due to this, the model can now show relationships in space more clearly. Higher-order mappings, such as those utilized by Hadamard attention, are widely acknowledged to offer better feature fitting capabilities [3]. y i S _ max f(x i ) x i x i (1)
For example, a study by Ronnnberger, Fischer, and Brox [3] demonstrates that Equation (1) defines a second-order attention mechanism, effectively mapping the input through a higher-dimensional space. This higher-order transformation allows for more expressive feature modeling.In contrast lower-order mappings, may result in reduced accuracy in complex tasks. 3.4. ELSA with Hadamard Attention
In this case, the input feature map is represented by xi, the convolution process is shown by f(xi), and the output feature map is yi. Because the second-order term qikj and the value vi are included, the self-attention structure in the Transformer represents the input xi in a third-order manner. This could be one of the reasons why the finished product has to be improved even more. ELSA with Hadamard attention [8], in which the Hadamard product is inserted as the second-order term between queries and keys, is one such approach. The following is an expression of Hadamard attention formulation (2): ŷ� ᵢ = Softmax(f(Hk ⊙ Hq)) Hᵥ + xᵢ
(2)
The symbol ⊙ performs Hadamard multiplication on operations f(xi), which is a convolution while Hk, Hq, and Hv indicate feature maps extracted through the linear transformation input. The design of the ELSA Transformer module includes adding a duplicate multi-layer perceptron (MLP) module behind the
(3)
3.5. Addressing Shallow Feature Noise via the DCML Module Brain tumor segmentation tasks aim to guide the network’s attention toward the relevant tumor regions. However, significant challenges arise due to variations in tumor size—particularly when tumors are very small—making accurate localization difficult. One of the key issues stems from the encoder’s progressive downsampling operations, which produce increasingly coarse feature representations. This process often leads to the loss or compression of crucial boundary information, thereby compromising segmentation accuracy. A number of approaches have been suggested to address this problem. Using larger images as input allows one to create larger feature maps, which helps to keep more information about the image’s position. Yet, using this technique requires a lot more computing power. A further approach is to merge shallow and deep features, using the fine boundaries in the shallow layers, however, it is difficult to successfully combine these elements. Some fusion methods put in additional information that is not needed, as shallow feature maps keep a lot of background noise. But when feature fusion is involved, this noise can reduce the accuracy of segmentation.To handle this issue, we introduce the Dense Cross-Multiplication Link (DCML) module as seen in Figure 5. Subfigure 5(a) presents a simple type of skip connections that do not include fusion across scales. In subfigure 5(b), dense addition is used to add up the features from different scales. Subfigure 5(c) uses dense concatenation to combine the features by channel. In subfigure 5(d), the proposed DCML method is seen, using multiplicative fusion to add semantic information and reduce the influence of minor features. All models use Batch Normalization (BN) to help the training stay stable and make sure the features are always distributed evenly. Because the DCML module brings together features of various scales, it helps produce more accurate and reliable segmentation, especially for small or uncertain targets in medical images. It merges feature maps from different scales which allows high-level information and low-level details to work together. By integrating features at different scales, the model can pay attention to both shallow and deep features. For example, at the fusion stage m, the output is calculated using Em (m=1,2,3,4) as seen in Eq.(4). Fᵐ = ∏₄ᵢ₍m gᵢ(Eᵢ), m ≥ 1
(4)
The feature transformation in the given context is g, which uses upsampling for resolution restoration and a 1×1 convolution for channel reduction. Three alternative fusion strategies and the multiplicative feature fusion approach are contrasted visually in Figure 5. Without feature fusion, Subfigure 5(a) uses
187
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
technique. So, if one branch cannot combine the top features, multiplicative fusion will bring attention to the issue by creating a large gradient. Therefore, while training, the multiplication feature fusion approach creates strict limits on every branch helping each branch collect better features. Because of this connection between branches, each branch improves and the results are more precise.
Figure 5. DCML Module the U-Net’s jump connection method, which introduces noise from shallow information and improves image segmentation only slightly. In Subfigures 5(b) and (c), shallow and deep features are integrated by addition and concatenation fusion techniques, respectively. These two approaches, however, ignore the relationship between multiscale properties. The following equations can be used to express the three previously described methods: F2mul = g2(E2) * g3(E3) * g4(E4)
F2add = g2(E2) + g3(E3) + g4(E4) F2concat = concat(g2(E2), g3(E3), g4(E4))
(5) (6) (7)
The following equations can be used to explain the computations made when the neural network back propagates to determine the gradient:
F2mul° =g 3 (E 3 ) g 4 (E 4 ) (8) g 2 (E 2 ) ∂F2add° = 1 (9) ∂g 2 (E 2 ) ∂F2concat° = concat(1,0,0) (10) ∂g 2 (E 2 )
188
Equations (8) and (9) demonstrate that the gradient of every branch does not depend on other branches and is not affected by the addition or concatenation of other branches. So, because each branch’s output does not influence the others, the network struggles to find connections between branches. Alternatively, with multiplication, the gradient of a branch can respond to changes in other branches, unlike with the initial
3.6. Spatial Local Attention Module (SLAM) Medical image segmentation models using deep learning often rely on spatial attention mechanisms such as the Spatial Local Attention Module (SLAM). SLAM is designed to help the model selectively focus on tumor-specific regions in brain MRI scans, thereby improving segmentation accuracy by enhancing the representation of clinically important spatial information. The input to SLAM is a feature tensor of size C/ N×H×W, where C is the number of channels, N is the number of samples, and H×W is the spatial dimensions. SLAM initially aggregates the spatial contextual information by summing corresponding activations across the feature maps, strengthening the relational understanding between different anatomical regions. To obtain robust and discriminative spatial cues, SLAM employs two complementary pooling operations: Average Feature Extraction (AFE), which captures global contextual responses, and Maximum Feature Extraction (MFE), which preserves the most salient tumor-related activations. The pooled features are concatenated along the channel axis to generate a fused representation Se∈R2×H×W, enabling the module to encode both dominant and supporting spatial structures. A convolution operation is then applied to Se to learn an attention map Sconv∈R1×H×W, which assigns higher weights to tumor-relevant areas while suppressing less informative regions. This attention map is subsequently multiplied element-wise with the original input feature map to produce the spatially enhanced output Sact, which emphasizes pathological characteristics essential for accurate tumor delineation. During decoding, the SLAM output is fused with features from DCML and other decoder layers using skip connections and progressive upsampling operations. Bilinear interpolation followed by 3×3 convolutions refines structural boundaries, while a final 1×1 convolution restores the feature map dimension to align with the segmentation output space. By integrating localized attention into the decoding pathway, SLAM significantly improves the segmentation of whole tumor (WT), tumor core (TC), and enhancing tumor (ET) regions. In the final step (see Figure 6), the recovered features are passed through a spatial and channel Squeeze-and-Excitation (sSE) Block to refine segmentation accuracy. The sSE Block compresses the feature map U along the channel axis and excites spatially significant regions using a sigmoid activation function and a 1×1×1 convolution. This selective enhancement contributes significantly to finegrained tumor segmentation.
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 6. The architecture of the SLAM module
L DICE =1 2
y p + µ (11) y + p + µ i
i
i
i i
i
i
L BCE (yi ln(pi )+(1 yi ) ln(1 pi )) i
L L DICE L BCE
Figure 7. Channel Squeeze and Spatial Excitation (sSE[21]) Finally towards the end of the encoder pathway, upsampling, and adjustments to the channel count are made with 1x1 convolutions to push the feature map process in the prior stage back to the input image size. After that, as shown in Figure 7 the recovered features are enhanced by applying an sSE Block to give spatial and channel-wise excitation for image segmentation. A feature map U is compressed along channels and spatially excited by the sSE Block, which makes a major contribution to fine-grained image segmentation. A sigmoid activation function and a 1×1×1 convolution layer are used to calculate the final segmentation outputs. 3.7. Loss Function
In medical image segmentation, disparities in lesion sizes and shapes, as well as unbalanced foreground-background distributions, are common issues. To overcome these problems, we propose a novel loss function that blends binary cross-entropy (BCE) loss (Equation 12) with dice loss (Equation 11). This combo approach increases the effectiveness of both model training and optimization. The formulation of this integrated loss function is described below:
(12)
In this case, y is the image’s actual label, p is the expected outcome, and ε is a parameter that is set to 1 for increased stability. Using this aggregated loss function enables the network to converge quickly and steadily, producing good results on a range of medical image datasets. Furthermore, we used this loss function in our comparative experiments for training all the systems under comparison for various segmentation tasks.
4. Experiments and Results 4.1. Dataset
We trained and validated our built model using the BraTS2021 dataset, which we acquired from the BraTS competition. This training dataset includes 1251 brain imaging cases with matching segmentation labels for each case (as seen in Figure 8). The four different 3D MRI modalities included in each case are T1-weighted (T1), T1-enhanced contrast (T1ce), T2-weighted (T2), and T2 fluid-attenuated inversion recovery (FLAIR). To attain a uniform isotropic resolution of 1 × 1 x 1 mm, the image dimensions for each modality are standardized to 240×240×155. They are then aligned, resampled, and skull-stripped. 4.2. Assessment metrics
The 95% Hausdorff distance (HD) and the Dice score coefficient (DSC) are two metrics used to evaluate the precision of brain tumor segmentation. The mathematical formulae (13) and (14) provide the DSC and HD readings, respectively.
189
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 8. Illustration depicting various modalities of BraTS 2021[1]
Figure 9. Segmentation results of the proposed method: a. FLAIR, b. T1ce, c. proposed method, and d. ground truth segmentation DSC(G, P)
2 iN 1G i Pi
G P N i 1
ii
N i 1 i
(13)
HD(G', P') = max max min ||g '- p ' ||, max min ||g '- p ' || (14) g ' G' p ' P ' p ' P ' g ' G '
Gi and Pi stand for voxel I’s actual and expected values, respectively, and G’i and P’i for the sets of ground truth and predicted surface points. While the greatest Hausdorff Distance (or HD) is one thing, the 95th percentile of distances between border points in G’ and P’ is another. This measure is used to reduce the impact that a small number of outliers may have on the evaluation as a whole. 4.3. Results
190
We randomly partitioned the BraTS 2021 training dataset into training and validation subsets, consisting of 1,188 and 63 subjects, respectively. This partitioning was necessary because semantic segmentation labels for the official BraTS 2021 validation set are not publicly accessible. The results presented in Table 1 reflect the performance of our model on the internal
validation set derived from the BraTS 2021 training data. We obtained an average Dice score of 92.69 percent and an average Hausdorff Distance (HD) of 5.34 mm for all three subregions of tumors. In particular, the Enhancing Tumor (ET), Tumor Core (TC) and Whole Tumor (WT) scored 90.12 percent, 94.19 percent and 93.77 percent on Dice and 7.32 mm, 4.40 mm and 4.31 mm on Hausdorff. Figure 9 shows the qualitative results of our method which prove its effectiveness. WT is shown in green and orange, TC in green and gray and ET in red in the annotations. The model clearly separates the tumor regions in the first three rows of Figure 9. Our method is shown to be correct and reliable in locating tumor areas. At the same time, the fourth row reveals that the model has challenges when trying to identify the exact regions of the tumors. This may happen because of different levels of intensity or the challenging structure of the tumor. These findings show that it is difficult to accurately identify tumors with unusual shapes or mixed features. Although there are some gaps, our model does well in most cases with exceptions present in brain tumor segmentation tasks. The last row highlights instances where segmentation could be improved. The segmentation performance on the BraTs 2021 validation dataset is compared with that of six state-ofthe‐art methods in Table 2. To measure performance, Dice Score (percent) and 95 percent Hausdorff Distance (mm) are used for ET, TC and WT. When compared to other models, our method attains the highest average Dice accuracy (92.69 percent) and the best Hausdorff distance (5.34 mm). It is most accurate for WT (93.77 percent) and TC (94.19 percent). The robustness and effectiveness of our model in these results correctly splitting up different parts of tumors in MRI images. 4.4. Ablation Study
To evaluate the contribution of each component within our proposed SwinNetPlus architecture, we conducted a detailed ablation study on the BraTS 2021 validation dataset. Specifically, we assessed the impact of the Dense Cross-Multiplication Link (DCML), the Spatial Local Attention Module (SLAM), and the multiplicative fusion strategy. Table 2 summarizes the performance of different ablated vari-
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Table 1. Summary of Key Related Works and Limitations Author(s)
Method
Key Contributione
Limitation
Long et al. [2]
FCN
End-to-end semantic segmentation
Poor boundary localization
Zhou et al. [4]
UNet++
Nested skip connections
Increased complexity
Ronneberger et al. [3]
Encoder–decoder with skip connections
U-Net
Çiçek et al. [5]
3D U-Net
Dosovitskiy et al. [8]
Volumetric segmentation
ViT
Liu et al. [9]
Global self-attention
Chen et al. [13]
TransUNet
Hatamizadeh et al. [14]
Hybrid CNN–Transformer
Transformer encoder for 3D data
UNETR
Jiang et al. [16]
Swin-based brain tumor segmentation
SwinBTS
Ghazouani et al. [21]
Swin + ELSA
Xing et al. [17]
Enhanced local attention
NestedFormer
High memory usage
Computationally expensive
Efficient hierarchical attention
Swin Transformer
Limited global context
Modality-aware transformer
Weak local boundary sensitivity
Feature inconsistency
Requires large datasets Limited local attention
Inadequate multiscale fusion High computational cost
Table 2. Segmentation outcomes observed on the BraTS 2021 validation dataset Method UNETR[14]
Dice Score (%) ET
58.5
TC
95% Hausdorff Dist. (mm)
WT
76.1
71.16
9.35
91.83
86.59
16.03
92.69
7.32
83.93
85.11
89.15
SegTransVAE[22]
85.48
90.52
92.6
SwinBTS[16]
Swin UNETR[18] Our Method
80
83.21 85.8
90.12
86.4
85.97
92
84.75 88.5
94.19
ET
78.9
3DUNet[5]
NestedFormer[17]
AVG.
86.13 89.53
92.6
88.96
93.77
TC
8.84
WT
8.26
AVG. 8.81
4.73
11.92
12.26
10.26
2.89
3.57
5.84
4.1
5.26 6.01
5.31
15.51 3.77 4.4
4.56 3.65 5.83
4.31
5.04 8.54 5.2
5.34
Table 3. Ablation study of SwinNetPlus on the BraTS 2021 validation dataset Model Variant Baseline ( without DCML/SLAM) Baseline + DCML
DCML
SLAM
Mult. Fusion
Dice Score (AVG) ↑
95% HD (AVG) ↓
87.15
6.93
90.03
5.59
88.76
Baseline + SLAM
89.21
Full Model (SwinNetPlus)
92.69
Baseline + DCML + SLAM (Additive Fusion)
6.15 5.94 5.34
191
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N° 3
2026
Figure 10. Ablation Study Results of SwinNetPlus on the BraTS 2021 Validation Dataset ants using Dice Score (percent) and 95 percent Hausdorff Distance (mm). In Fig 10 visual comparison of the ablation study highlights the progressive improvement in model performance with the addition of DCML, SLAM, and Multiscale Fusion. The Dice Score (AVG) increases steadily from 87.15% in the baseline model to 92.69% in the full SwinNetPlus model, indicating better segmentation accuracy. Simultaneously, the 95% Hausdorff Distance (AVG) decreases from 6.93 mm to 5.34 mm, demonstrating enhanced boundary precision. These results confirm that each added component contributes to performance gains, with the combined use of DCML, SLAM, and Multiscale Fusion achieving the most accurate and precise segmentation on the BraTS 2021 validation dataset.
5. Discussion
192
The results obtained from the evaluation of SwinNetPlus on the BraTS 2021 dataset demonstrate the model’s effectiveness in addressing the primary challenges of brain tumor segmentation, such as lesion heterogeneity, varying tumor sizes, and the need for accurate boundary delineation. Compared to other state-of-the-art methods like UNETR, 3DUNet, and Swin UNETR, SwinNetPlus consistently achieved superior performance across all key metrics, including Dice score and Hausdorff Distance. This improvement can be attributed to the synergy between the convolutional backbone and the hierarchical attention mechanisms of the Swin Transformer, which together facilitate comprehensive global and local feature extraction. The inclusion of ELSA transformer blocks within the encoder played a crucial role in enhancing long-range dependency modeling while preserving fine-grained local details, a limitation of conventional CNN-based models. Moreover, the DCML module effectively addressed the problem of feature degradation during fusion by leveraging multiplicative interactions that preserve
semantic richness while suppressing noise. Similarly, the SLAM allowed the network to focus selectively on critical spatial regions, further refining the segmentation quality, particularly along tumor boundaries. Despite the strong performance, certain limitations were observed. In some complex cases, particularly in instances with highly irregular tumor morphologies or low contrast between tumor and healthy tissue, the model exhibited reduced segmentation precision. This suggests that even advanced attention modules may struggle to generalize across all imaging conditions, highlighting the need for further enhancements in feature sensitivity and robustness. Additionally, while the current architecture is computationally efficient compared to many 3D models, inference time and resource utilization remain important considerations for clinical deployment. Future optimization of the model for real-time processing and integration into diagnostic workflows would be essential for translational impact. Overall, the proposed SwinNetPlus model represents a significant advancement in automated brain tumor segmentation. The combination of transformer-based attention, dense feature fusion, and spatial refinement modules contributes to improved diagnostic accuracy, supporting the model’s potential application in clinical decision-making. Future directions should explore the generalizability of this architecture to other medical segmentation tasks, as well as investigate architectural modifications and novel attention strategies to further boost performance and interpretability.
6. Conclusion and future work
In summary, this study introduces SwinNetPlus, an innovative deep learning framework designed to accurately segment brain tumors from MRI scans. By effectively integrating Convolutional Neural Networks (CNNs) with Swin Transformers, and incorporating advanced modules such as the Dense Cross-Multiplication Link (DCML) and the Spatial Local Attention Module
Journal of Automation, Mobile Robotics and Intelligent Systems
(SLAM), the architecture demonstrates a strong ability to capture complex spatial and semantic features. The SLAM enhances the model’s attention to tumor-relevant regions, enabling finer discrimination of subtle variations, while the DCML module reduces background noise and improves multiscale feature fusion through an efficient multiplicative strategy. The use of ELSA transformer blocks within the encoder strengthens long-range dependency modeling and detailed local feature extraction, further refined by channel squeeze and spatial excitation blocks for enriched spatial and channel-wise representations. Evaluated on 1,251 annotated MRI volumes from the BraTS 2021 dataset, SwinNetPlus achieved an outstanding average Dice score of 92.69 percent and a Hausdorff distance of 5.34 mm, outperforming several state-of-the-art 3D segmentation models. These results confirm the effectiveness and robustness of SwinNetPlus in handling tumor heterogeneity and complex visual features. This work not only contributes a reliable method for brain tumor segmentation but also lays the foundation for future advancements in medical image analysis. Future research could explore expanding SwinNetPlus to segment other types of tumors or pathologies, as well as enhancing its performance through the integration of novel attention mechanisms, fusion strategies, or transformer architectures, thereby pushing the boundaries of automated medical diagnosis. AUTHORS
Vikash Verma* – Department of Computer Science and Engineering, Rabindranath Tagore University, Bhopal, India; email: vikashverma2005@gmail.com. Pritaj Yadav – Department of Computer Science and Engineering, Rabindranath Tagore University, Bhopal, India; email: yadavpritaj@gmail.com. ∗ Corresponding author
References
[1] Z. Liu et al., “Deep learning based brain tumor segmentation: a survey,” Complex & Intelligent Systems, vol. 9, 2023, pp. 1001–1026; doi: 10.1007/s40747-022-00815-5. [2] J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” Proc. IEEE Conference on Compuer Vision and Pattern Recognition (CVPR 15), Jun. 2015, pp. 3431–3440; doi: 10.1109/ CVPR.2015.7298965. [3] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 Oct. 2015, pp. 234–241; doi: 10.1007/978-3319-24574-4_28. [4] Z. Zhou et al., “UNet++: A nested U-Net architecture for medical image segmentation,” Deep Learning in Medical Image Analysis and Multimodal Learn-
VOLUME 20,
N° 3
2026
ing for Clinical Decision Support, Sep. 2018, pp. 3–11; doi: 10.1007/978-3-030-00889-5_1. [5] Ö� . Çiçek et al., “3D U-Net: Learning dense volumetric segmentation from sparse annotation,” Medical Image Computing and Computer-Assisted Intervention – MICCAI 2016, Oct. 2016, pp. 424–432; doi: 10.1007/978-3-319-467238_49. [6] X. Li et al., “H-DenseUNet: Hybrid densely connected UNet for liver and tumor segmentation from CT volumes,” IEEE Transactions on Medical Imaging, vol. 37, no. 12, Dec. 2018, pp. 2663–2674. [7] A. Vaswani et al., “Attention is all you need,” Proc. 31st Conference on Neural Information Processing Systems (NIPS 17), Dec. 2017, pp. 5998– 6008; https://arxiv.org/abs/1706.03762 [8] A. Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” Proc. International Conference on Learning Representations (ICLR 21), May 2021. Available: https://arxiv.org/abs/2010.11929 [9] Z. Liu et al., “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE/CVF International Conference on Computer Vision (ICCV 21), Oct. 2021, pp. 10012–10022; doi: 10.1109/ ICCV48922.2021.00986. [10] J. Ma et al., “SwinFusion: Cross-domain longrange learning for general image fusion via swin transformer,” IEEE/CAA Journal of Automatica Sinica, vol. 9, no. 7, Jul. 2022, pp. 1200–1217. [11] H. Du et al., “SwinPA-Net: Swin transformerbased multiscale feature pyramid aggregation network for medical image segmentation,” IEEE Transactions on Neural Networks and Learning Systems, 2024; doi: 10.1109/ TNNLS.2022.3204090 [12] H. Cao et al., “Swin-Unet: Unet-like pure transformer for medical image segmentation,” Proc. Eur. Conference Computer Vision Workshops (ECCVW 22), Oct. 2022, pp. 205–218; doi: 10.1007/978-3-031-25066-8_9. [13] J. Chen et al., “TransUNet: Transformers make strong encoders for medical image segmentation,” IEEE Trans. Med. Imag., vol. 41, no. 10,Oct. 2022, pp. 2863–2874.. [14] A. Hatamizadeh et al., “UNETR: Transformers for 3D medical image segmentation,” Proc. IEEE/CVF Winter Conference on Applications of Computer Visions (WACV 22), Jan. 2022, pp. 574–584; doi: 10.1109/ WACV51458.2022.00181. [15] A. Shaker et al., “UNETR++: Delving into efficient and accurate 3D medical image segmentation,” IEEE Transactions on Medical Imaging, vol. 43, no. 9, Sept. 2024, pp. 3377–3390.. [16] Y. Jiang et al., “SwinBTS: A method for 3D multimodal brain tumor segmentation using swin 193
Journal of Automation, Mobile Robotics and Intelligent Systems
transformer,” Brain Sciences, vol. 12, no. 6, Jun. 2022; doi: 10.3390/brainsci12060797. [17] Z. Xing, L. Yu, L. Wan, T. Han, and L. Zhu, “NestedFormer: Nested modality-aware transformer for brain tumor segmentation,” Medical Image Computing and Computer Assisted Intervention – MICCAI 2022, Sep. 2022, pp. 140–150; doi: 10.1007/978-3-031-16443-9_14. [18] A. Hatamizadeh et al., “Swin UNETR: Swin transformers for semantic segmentation of brain tumors in MRI images,” Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries, MICCAI 2021 Workshops, Sep. 2021, pp. 272–284; doi: 10.1007/978-3-03108999-2_22. [19] J. Hu, L. Shen, and G. Sun, “Squeeze-andexcitation networks,” Proc. IEEE Conference
194
VOLUME 20,
N° 3
2026
on Computer Vision and Pattern Recognition (CVPR 18), Jun. 2018, pp. 7132–7141; doi: 10.1109/CVPR.2018.00745. [20] J. Zhou et al., “ELSA: Enhanced local self-attention for vision transformer,” ArXiv,2021; https://arxiv.org/abs/2112.12786 [21] F. Ghazouani, P. Vera, and S. Ruan, “Efficient brain tumor segmentation using Swin transformer and enhanced local self-attention,” International Journal of Computer Assisted Radiology and Surgery, vol. 19, no. 2, Feb. 2024, pp. 273–281. [22] Q. -D. Pham et al., “Segtransvae: Hybrid Cnn Transformer with Regularization for Medical Image Segmentation,” 2022 IEEE 19th International Symposium on Biomedical Imaging (ISBI), Kolkata, India, 2022, pp. 1-5, doi: 10.1109/ISBI52829.2022.9761417.
VOLUME 20, N∘ 3 2026 Journal of Automation, Mobile Robotics and Intelligent Systems
APPLICATION OF BOTH STATIC AND DYNAMIC ANALYSIS METHODOLOGIES FOR MALWARE DETECTION Submitted: 5th December 2024; accepted: 2nd April 2025
Andrzej Mycek, Mirosław Roszkowski DOI: 10.14313/jamris-2026-050 Abstract: Currently, the two most rapidly developing fields in IT are Cybersecurity and Artificial Intelligence. Both of these areas intersect significantly—advancements in AI contribute to the growth of the security sector. This is crucial because cybercriminals are also advancing, devising new methods to cause harm, such as data theft, hijacking infrastructure, turning off IT systems, or espionage. To prevent such incidents, security specialists are developing increasingly sophisticated tools that enable the detection and neutralization of attacks. Among these tools are two techniques: static analysis and dynamic analysis. In this thesis, both methods have been thoroughly examined, and the benefits of combining these techniques in the context of protection against malware have been demonstrated using specific malicious software. Keywords: cybersecurity, dynamic malware analysis, malware detection, ransomware, static malware analysis
1. Introduction The Cybersecurity market is continually expanding. According to research by [1], the Cybersecurity market has grown by over 30% in the past three years, and it is projected to increase by 58.5% over the next five years, reaching nearly $350 billion by 2026. The ongoing digital transformation of businesses drives this growth. The advancement of global economies necessitates the development and deployment of increasingly sophisticated IT infrastructures, with a rising number of new devices used daily by individuals, often referred to as the Internet of Things (IoT). Additionally, new regulations and standards concerning data protection and information security are being continuously introduced, with the European Union leading in this domain [2]. A positive development is the growing awareness among companies and ordinary users regarding cyber threats. This awareness is critical given that many companies, including those that produce software, adopt a relatively lax approach to security. The ratio of employees in programming, administrative, and security departments is approximately 100:10:1. This approach, even among IT companies, is ruthlessly exploited by cybercriminals. Currently, the most prevalent types of attacks include Man-in-the-Middle attacks, Brute Force attacks, Distributed Denial of Service (DDoS) attacks,
Malware, Phishing, and Social Engineering [3, 4]. Man-in-the-middle attacks aim to intercept communication between two parties, allowing the attacker to capture and modify transmitted data. Such attacks frequently occur in environments with unsecured Wi-Fi networks, such as restaurants, airports, or pubs. [5] is a method that can defend against this type of attack. Brute Force attacks involve guessing a password manually or, more commonly now, automatically testing all possible combinations until the correct one is found. The downside for attackers is the required time, making this method effective mainly for short and straightforward passwords (ideally, dictionary words) without special characters [6]. Distributed Denial of Service (DDoS) attacks are prevalent among experienced and novice cybercriminals. These attacks aim to overload a service provider’s IT infrastructure by sending multiple requests from multiple sources simultaneously. Attackers may use large volumes of TCP and UDP packets or employ applications that send HTTP/HTTPS requests to web applications [7]. Among these threats, Malware is the most diverse. It encompasses any code added to or removed from software with the intent to cause harm or disrupt system functionality. Malware is a broad term that includes various forms of malicious software, such as trojans, ransomware, viruses, spyware, adware, and worms [8]. The article will explore Malware and analyse methods for detecting these threats. The final types on the list of most common attacks are Phishing and Social Engineering. These attacks are similar as they aim to extract sensitive information such as passwords, personal data, or financial details from potential victims. They involve manipulating users to take actions desired by the fraudsters, often leveraging human nature—such as curiosity or gullibility. Many countries currently run awareness campaigns online, in print media, and on television to warn about such attacks, with a particular focus on protecting elderly individuals who are especially vulnerable [9].
2.
Related Work
Research into malware in the IT world has been conducted since the emergence of the first virus. Undoubtedly, the first such threat was the Creeper
Open Access. © 2026 Andrzej Mycek and Mirosław Roszkowski, published by Łukasiewicz Research Network — Industrial Research Institute for Automation and Measurements PIAP. 4.0 License
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives
195
Journal of Automation, Mobile Robotics and Intelligent Systems
virus, created by Bob Thomas. The malicious software he developed in 1971 operated on the TENEX (TOPS-20) system within the ARPANET network. The program’s purpose was to test the concept of selfreplicating software. Interestingly, this virus did not cause significant harm; it merely displayed the message, “I am the creeper, catch me if you can! [10]. Although the virus was not harmful, it changed people’s approach in the IT industry. IT personnel realized that there are no completely secure systems. Moreover, the appearance of the first virus opened a new chapter in the IT industry. This chapter undoubtedly discusses the development of antivirus systems. Creeper contributed to creating the first antivirus program in history, Reaper [11]. This software laid the foundation for developing highly advanced systems capable of detecting and removing malicious software. However, it also encouraged others to create malware, and the battle between cyber criminals and IT security specialists continues to this day and will persist indefinitely. One of the first researchers to take on the fight against malware was Fred Cohen. His academic research provided the groundwork for future methods of malware detection [12]. One of the first antivirus software producers was the American company McAfee Corp., founded by John McAfee in 1987. This company quickly became a leader in antivirus solutions, and its software used signature-based technology for many years, as discussed later in this work. McAfee’s success was mirrored by companies like Symantec with its Norton AntiVirus product, F-Secure, and Kaspersky Lab—all tools above utilized signature-based technology during that period. Just a few years ago, cybersecurity researchers primarily used two methods for detecting and analyzing malware: Signature-based Detection and Heuristicbased Detection. The former was the most commonly used. The signature-based detection method involves using pre-established signature databases and specific fragments of malicious code. The advantage of this approach is its high effectiveness in detecting known threats; however, its drawback is its inefficiency in dealing with new malware variants [13]. The second prevalent method used in malware detection was the Heuristic-based Detection method. This method involves analyzing the code of malicious software for suspicious behavior that may indicate the presence of malware. Unlike Signature-based Detection, this technique allows for the detection of new and unknown viruses and so-called zero-day attacks. Despite its advantages, this method has drawbacks, the most significant being the risk of generating false alerts (false positives) [14]. Over time, new concepts for detecting malicious software began to emerge. The late 1990s and early 2000s marked a period when behavioral analysis methods gained significant popularity. The rapid expansion of the Internet, along with the widespread 196
VOLUME 20,
N∘ 3
2026
sharing and distribution of malicious software, drastically increased the rate of malware propagation. Despite displaying very similar behaviors, these types of software often have varying structures due to different code modification techniques, such as polymorphism and metamorphism, which pose a substantial challenge for signature-based detection. It is also important to note that malware frequently relies on various API calls provided by operating systems, which makes behavioral analysis particularly effective in identifying such threats [15]. In addition to behavioral analysis, the beginning of the 20th century saw the development of several other techniques. Among the most important of these, which should be highlighted, are undoubtedly sandboxing, memory analysis, and reputation-based analysis. Sandboxing is an isolated environment in which malicious software is executed and thoroughly analyzed. This allows for the verification of software without impacting critical infrastructure. Sandboxing has now become an integral part of corporate infrastructure, enabling the testing of malicious code without risk [16]. Unfortunately, even sandboxing is currently being ‘bypassed’ by cyber criminals. There are now malware strains capable of detecting whether they operate within a sandbox environment. They only activate once they find themselves outside the sandbox’s perimeter. Attackers assess whether their application runs in a sandbox primarily by analyzing input data. In a typical environment, users often click the mouse or type on the keyboard. However, when malware operates within a sandbox, the attacker receives no input data. Fortunately, there are methods to ‘hide’ the sandbox environment from the malware. In article [17], the authors demonstrate how this can be achieved using Delphix. Delphix is an AI tool designed to mask data, making it more difficult for malware to analyze. Delphix replaces any sensitive or confidential data with realistic-looking alternatives that do not contain critical information. Cybercriminals in the current century are increasingly focusing on concealing their malicious software. Malicious software has been developed that leaves no traces on the disk, operating solely in the system’s RAM. The growth of the Internet, in addition to techniques like sandboxing, has significantly contributed to the emergence of reputation-based analysis. This method involves evaluating the trustworthiness of websites, domains, or files. It enables quick blocking of access to dangerous sources based on globally gathered data. The most famous example of such a service is VirusTotal, which we also utilized in our work [18]. The recent history undoubtedly belongs to machine learning and artificial intelligence. Thanks to the remarkable advancement in computational power and machine learning techniques, significant progress has been made in malware detection driven by these technologies. The analysis of vast datasets and the extraction of patterns based on specific behaviors enable the detection of increasingly sophisticated cyberattacks.
Journal of Automation, Mobile Robotics and Intelligent Systems
Furthermore, anomalies in abnormal behavior in systems, users, or networks have become crucial to modern security systems [19].
3. Malware Characteristics Cybercriminals create malicious software to cause harm, steal data, or gain access to IT systems. The “malware family” encompasses various malicious software types, with the most common examples being viruses, trojans, worms, spyware, adware, ransomware, and rootkits. Each type behaves and operates differently, carrying out unique tasks. Viruses aim to replicate and modify other computer programs by inserting their infected code. They are most commonly spread through email attachments and files shared on hosting services. Trojans, in contrast, disguise themselves as legitimate software and rely on deceiving users to disrupt systems, perform harmful actions like data theft, or create backdoors. Worms are similar to viruses primarily in their ability to replicate, but unlike traditional computer viruses, they do not attach themselves to existing programs. Spyware, as the name suggests, monitors user activity, gaining access to emails and even capturing keystrokes. Adware behaves similarly regarding user analysis, but its primary goal is to display unwanted ads to users, generating revenue for cybercriminals. Adware is often bundled with spyware. In the 1990s, rootkits began gaining popularity. Rootkits aim to access the root (administrator) account while remaining undetected. They are capable of modifying software while staying hidden. Nowadays, rootkits are primarily used by cybercriminal groups and in Advanced Persistent Threat (APT) attacks, where the key objective is to conceal presence and maintain long-term control over the victim’s system [20]. Another highly prevalent type of malware is ransomware. The goal of ransomware is to block access to a computer system by encrypting its files. In theory, the victim may receive a decryption key upon ransom. The most common currency in these cases is cryptocurrency, as transactions using cryptocurrencies are difficult to trace, making it challenging to identify the attacker [21]. We have examined ransomware and trojans in more depth later in this work, as these two types of malware were the primary focus of our research. We aimed to demonstrate how applying static and dynamic analysis methods enables the identification of malware and how these methods can contribute to the development of Intrusion Detection Systems [22].
4. Problems and Research Methods Used in the Project The ever-increasing volume of threats in cyberspace has led to a growing demand for cybersecurity professionals year on year. Threat detection and response have become critical components of every organisation’s operational framework. Due to the evolving characteristics of
VOLUME 20,
N∘ 3
2026
modern malware variants, traditional signaturebased methods are proving to be progressively less effective with each passing month. In light of these challenges, static and dynamic analysis techniques are gaining prominence and recognition, offering more accurate—and often profoundly comprehensive—insights into the behavior of executable files, thereby enabling the effective identification of malicious activities. Static analysis allows us to examine malicious code without executing it, which facilitates the rapid and safe acquisition of valuable information, such as utilised libraries and the structural layout of the file. Conversely, dynamic analysis involves executing the code within a fully controlled environment, providing detailed visibility into the malware’s runtime behavior—such as registry modifications, file operations, and attempts to establish connections with commandand-control (C2) servers. This paper presents a comprehensive analysis of both methodologies, illustrated by an investigation of a well-known piece of malware. With the rapid evolution of cyber threats, increasingly sophisticated attack methods are emerging. The malware examined in this work, specifically ransomware and Remcos RAT, represent significant risks to individual users, critical infrastructure, and financial institutions, often leading to substantial financial losses, operational downtime, and privacy breaches. This study thoroughly analyzes the behavior of these two threats using two essential techniques: static analysis and dynamic analysis. Static analysis enabled an assessment of the malicious software’s code, facilitating the identification of specific attack patterns and vectors without executing the malware. On the other hand, dynamic analysis allows for observing malware behavior in a controlled environment. Thanks to this, the characteristic patterns of WannaCry and Remcos RAT were observed in a practical example, and the effectiveness of static and dynamic analysis was assessed in the context of their potential use in IDS-type systems. We began the study of the selected malware by setting up a test environment wholly isolated from the systems we use in our lab daily. Two virtual machines were connected to a virtual network: one running Security Onion, a comprehensive Linux distribution designed for network monitoring, traffic analysis, threat detection, and incident response, and the other with a Windows 10 Pro client system. Security Onion has many pre-installed tools for detecting and analyzing network traffic, such as Suricata, Zeek, Wireshark, Network Miner, Sguil and, among others, ELK Stack (Elasticsearch, Logstash, and Kibana) for managing and analyzing logs [23]. The Ghidra tool, created by the National Security Agency (NSA), was installed on a Windows 10 machine and used in reverse engineering as a tool that allows for binary code analysis and source code recovery from binary files. 197
Journal of Automation, Mobile Robotics and Intelligent Systems
Figure 1. Security Onion Desktop This study examined two types of malware: the WannaCry ransomware and Remcos RAT. WannaCry, a form of ransomware that emerged in 2017, spread rapidly by exploiting a vulnerability in the Microsoft Windows operating system known as EternalBlue. The NSA initially identified this vulnerability and intended to use it by U.S. intelligence agencies as a cyber intelligence tool. However, a hacker group called “The Shadow Brokers” publicly released a collection of NSA tools, including EternalBlue. As a result, WannaCry quickly spread worldwide, infecting and encrypting data on over 200,000 computers across more than 150 countries [24]. The second type analyzed in the work was Remcos RAT. This Trojan allows unauthorized remote access to the victim’s computer in order to spy and steal data. The most common method of spreading is via infected email attachments [25].
5. Laboratory Environment and Malware Analysis The malware research was preceded by preparing a comprehensive lab environment in controlled conditions that allowed for simulated attacks on an endpoint running Microsoft Windows 10 (malware infection) and subsequent analysis of the malicious software. Two virtual machines were connected to the simulated internal network: one running Security Onion for threat detection and analysis and another with Windows 10 functioning as the victim machine, equipped with a static analysis tool. After installing Security Onion, the essential monitoring components and a suite of analytical tools were configured. On the Windows 10 machine, the Ghidra tool was installed and configured for static code analysis. Once configuration work was complete, selected malware research samples were downloaded from the MalwareBazaar website for use in the lab environment. This is a public database of malware samples. It offers access to even the latest malware, making this site a beneficial tool for researchers and security analysts [26, 27]. 5.1.
WannaCry Static Analysis
The first step in conducting static analysis is always to check the file’s properties. Our executable file with the .exe extension was imported into the 198
VOLUME 20,
N∘ 3
2026
Ghidra tool. The import report revealed that the analyzed WannaCry is built for the x86 architecture, with 45 data types, 46 functions, and 96 symbols. The file was then passed on for detailed analysis. Ghidra translated the binary code into assembly language. This step is one of the most important in conducting this type of analysis because it allows for understanding the operation of the malware, especially when the source code of the malicious software is not available. Ghidra can quickly identify functions in the analyzed file, allowing us to examine its interactions with other operating system components. Inspecting the function labelled “entry” enabled us to study the entry code generated for executable files in Windows. Thanks to Ghidra’s powerful capabilities, we were able to convert assembly code into C language, significantly simplifying and accelerating the malware analysis process. WinMain() function was identified in the decompiled code, containing critical operations that initiate the WannaCry ransomware. This included functions responsible for establishing network connections to download resources. In WannaCry’s case, a variable named strange_url was declared within the function assigned to a specific URL. This URL served as a “kill switch.” The malware attempted to connect to this address—if the server responded, the malware would terminate its operation; if not, WannaCry would continue spreading the infection. Through the discovery of this kill switch, WannaCry was ultimately defeated. Marcus Hutchins, who registered the domain, first made this discovery, thereby stopping the malware from spreading further. The malware’s creators likely embedded this mechanism to halt the malware if it spiralled out of control and to facilitate testing in controlled environments. In the second variant, WannaCry was analyzed using tools available at Security Onion to verify the malware regarding network communication. After importing the PCAP file, the packets were analyzed using Snort and Zeek. Using the Squert tool, we discovered two types of suspicious software: EternalBlue and DoublePulsar. EternalBlue is an exploit that exploits a vulnerability in the SMB protocol. In this example, WannaCry attempted to exploit the vulnerability using the EternalBlue exploit to spread across the network. 5.2.
Remcos RAT Static Analysis
Extracting and searching the malware’s strings allowed us to determine early on that we were dealing with an application designed for Windows systems, not, for example, MS-DOS. Using Ghidra, we reverseengineered the malware. We showed that the malware is a keylogger that records every keystroke, video, and sound, and it does all of this using the SendInput function from the winuser.h library (see: Figure 5). However, this was only one functionality of the Trojan. The research also discovered that the Trojan acted as a C2 (Command and Control) server, enabling
Journal of Automation, Mobile Robotics and Intelligent Systems
VOLUME 20,
N∘ 3
2026
Figure 2. Controlled Lab Environment and Malware Infection Flow
Figure 3. Fragment of disassembled WannaCry code
Figure 5. Fragment of a function from the winuser.h library Figure 4. Key fragment of the decompiled code it to execute commands in cmd and download files (see: Figure 6). 5.3.
WannaCry Dynamic Analysis
The primary process spawned three child processes and several additional ones that terminated after completing their tasks. The primary process masked itself within the system, appearing as a standard system process. The ransomware operated using an RSA key pair (public and private) and an AES-128CBC symmetric key for file encryption. Since WannaCry propagated across the network through an SMB protocol vulnerability, filters for
Figure 6. Executing commands in cmd.exe
tcp.port == 445, ARP, and DNS were applied during packet analysis in Wireshark. In Wireshark, we observed DNS queries containing a domain name acting as a kill switch, which indicated the presence of WannaCry. 199
Journal of Automation, Mobile Robotics and Intelligent Systems
5.4.
Remcos RAT Dynamic Analysis
During our analysis, we observed that a notess directory containing a file named logs.dat was created in the system. This file records the victim’s activity, providing the attacker with continuous access to data entered by the victim on the computer. During the data exfiltration phase, we discovered that the trojan communicates with an address functioning as a Command and Control (C2) server. Upon verifying this address, we determined that it is a highrisk IP address that cybercriminals have previously used for malicious activities on the network.
6. Conclusion This study thoroughly analyzed two types of malware, the ransomware WannaCry and the trojan Remcos RAT, using both static and dynamic analysis methods. The tests and analyses that were conducted provided valuable insights into the effectiveness of these methodologies and the behavior of the analyzed malware. Static analysis, which focuses on examining binary files, enabled the identification of characteristics such as file structure, string patterns, cryptographic mechanisms, and specific malware behaviors. On the other hand, dynamic analysis allowed real-time observation of the malware’s actions, such as process creation and network communication. The application of both analysis methods offers a comprehensive understanding of malware functionality. It significantly enhances preparedness and the ability to detect, analyze, and neutralize attacks, including known and emerging threats. Combining static and dynamic analysis methods, a hybrid approach significantly enhances malware detection capabilities. The approach presented in this study demonstrates that in today’s threat landscape—where the volume and sophistication of cyber threats are continually on the rise—adequate protection against emerging forms of malicious software is only achievable through a multi-layered defence strategy that leverages a diverse range of techniques. Future research should focus on automating and integrating these methodologies while incorporating advanced machine learning models for detecting new threats and addressing the challenges of an everevolving cyber landscape. Developing a model designed to correlate various metadata attributes with the observed behavior of malicious software in a controlled environment would enable even more effective threat classification. A promising direction for further research would be to expand the dataset to include additional malware families—such as fileless malware, Advanced Persistent Threats (APTs), and a broader range of ransomware variants. Considering these enhancements, the proposed solution could be successfully implemented within Endpoint Detection and Response (EDR) systems or Security Information and Event Management (SIEM) platforms in production environments. This would 200
VOLUME 20,
N∘ 3
2026
also facilitate a more accurate assessment of the solution’s effectiveness under real-world operational conditions. AUTHORS Andrzej Mycek∗ – CUT Doctoral School, Department of Computer Science, Faculty of Computer Science and Mathematics, Cracow University of Technology, Cracow, Poland, ul. Warszawska 24, 31-155 Cracow, Poland, e-mail: andrzej.mycek@pk.edu.pl. Mirosław Roszkowski – Department of Computer Science, Faculty of Computer Science and Mathematics, Cracow University of Technology, Cracow, Poland, ul. Warszawska 24, 31-155 Cracow, Poland, e-mail: miroslaw.roszkowski@pk.edu.pl. ∗
Corresponding author
References [1] V. Srivastava and V. Sharma, “Sandbox technology in a web security environment: a hybrid exploration of proposal and enactment”, International Journal of Scientific Research and Engineering Development, vol. 5, no. 3, pp. 488–494, 2022. [2] C. Hoofnagle, B. Van der Sloot, and F. Borgesius, “The European Union general data protection regulation: what it is and what it means”, Information & Communications Technology Law, vol. 28, no. 1, pp. 65–98, 2019. [3] A. Bendovschi, “Cyber-attacks – trends, patterns and security countermeasures”, Procedia Economics and Finance, vol. 28, no. 1, pp. 24–31, 2015. [4] A. Mycek and M. Łukaczyk, “Security of containerization platforms: threat modelling, vulnerability analysis, and risk mitigation”, ECMS 2024: Proceedings of the 38th ECMS International Conference on Modelling and Simulation, pp. 585–591, 2024. [5] A. Mallik et al., “Man-in-the-middle-attack: understanding in simple words”, International Journal of Data and Network Science, vol. 3, no. 2, pp. 77–92, 2019. [6] L. Boš njak, J. Sreš , and B. Brumen, “Bruteforce and dictionary attack on hashed real-world passwords”, 2018 41st International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO), vol. 41, no. 1, pp. 1161–1166, 2018. [7] A. Singh and B. Gupta, “Distributed denial-ofservice (DDoS) attacks and defense mechanisms in various web-enabled computing platforms: issues, challenges, and future research directions”, International Journal on Semantic Web and Information Systems (IJSWIS), vol. 18, no. 1, pp. 1–43, 2022.
Journal of Automation, Mobile Robotics and Intelligent Systems
[8] A. Namanya et al., “The world of malware: an overview”, 2018 IEEE 6th International Conference on Future Internet of Things and Cloud (FiCloud), vol. 6, no. 1, pp. 420–427, 2018. [9] I. Vayansky and S. Kumar, “Phishing – challenges and solutions”, Computer Fraud & Security, vol. 1, no. 1, pp. 15–20, 2018. [10] T. Chen and J. Robert, The Evolution of Viruses and Worms, CRC Press, 2004. [11] M. Manjramkar and K. Jondhale, Cyber Security Using Machine Learning Techniques, Atlantis Press, 2023. [12] E. Filiol, Computer viruses: from theory to applications, Springer, 2005. [13] D. Venugopal and G. Hu, “Efficient signature based malware detection on mobile devices”, Mobile Information Systems, vol. 4, no. 4, pp. 33–49, 2008. [14] Z. Bazrafshan et al., “A survey on heuristic malware detection techniques”, IKT 2013 - 2013 5th Conference on Information and Knowledge Technology, vol. 5, no. 1, pp. 113–120, 2013.
VOLUME 20,
N∘ 3
2026
[23] R. Heenan and N. Moradpoor, “Introduction to security onion”, PGCS 2016: The First Post Graduate Cyber Security Symposium - The Cyber Academy, Edinburgh Napier University, pp. 1–4, 2016. [24] Z. Liu et al., “Working mechanism of Eternalblue and its application in ransomworm”. In: Cyberspace Safety and Security, pp. 178–191, 2022. [25] V. Valeros and S. Garcia, “Growth and commoditization of remote access trojans”. In: 2020 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW), pp. 454–462, 2020. [26] J. Chen and R. Deng, “Similarity-based malware classification using graph neural networks”, Applied Sciences, vol. 12, no. 21, pp. 10837–10853, 2016. [27] D. Priambodo et al., “Observe-Orient-Decide-Act (OODA) for cyber security education”, International Journal of Advanced Computer Science and Applications, vol. 13, no. 10, pp. 246–255, 2022.
[15] H. Galal, “Behavior-based features model for malware detection”, Journal of Computer Virology and Hacking Techniques, vol. 12, no. 2, pp. 59–67, 2016. [16] V. Fasna and S. Swamy, “Sandbox: a secured testing framework for applications”, Journal of Technology & Engineering Sciences, vol. 4, no. 1, pp. 3–11, 2020. [17] V. Sathya et al., An Obfuscation Technique for Malware Detection and Protection in Sandboxing, volume 972, Springer, 2021. [18] S. Megira, R. Pangesti, and F. Wibowo, “Malware analysis and detection using reverse engineering technique”, Journal of Physics: Conference Series, vol. 1140, no. 1, pp. 1–12, 2018. [19] M. Hossain Faruk et al., “Malware detection and prevention using artificial intelligence techniques”, 2021 IEEE International Conference on Big Data (Big Data), pp. 5369–5377, 2021. [20] L. Hughes and G. DeLone, “Viruses, worms, and Trojan horses: serious crimes, nuisance, or both?”, Social Science Computer Review, vol. 25, no. 1, pp. 78–98, 2007. [21] M. Akbanov and V. Vassilakis, “WannaCry ransomware: analysis of infection, persistence, recovery prevention and propagation mechanisms”, Journal of Telecommunications and Information Technology, vol. 1, no. 1, pp. 113–124, 2019. [22] A. Mycek, “Monitoring, management, and analysis of security aspects of IaaS environments”, Journal of Telecommunications and Information Technology, vol. 94, no. 4, pp. 108–116, 2023. 201