Vasko O. Mathematical modeling and non-gradient optimization of convolutional networks based on multivalued neurons

Українська версія

Thesis for the degree of Doctor of Philosophy (PhD)

State registration number

0825U004180

Applicant for

Specialization

  • 111 - Математика

15-01-2026

Specialized Academic Board

PhD 11382

Uzhhorod National University State Higher Educational Institution

Essay

This paper is devoted to the development and theoretical justification of a mathematical model of a convolutional neural network based on multi-valued neurons (CNNMVN), designed for pattern recognition and image classification tasks. A comprehensive mathematical framework is formulated, providing a formalized description of convolutional layers and their functional properties. Two types of convolutional kernels: single-channel and multi-channel are proposed, and their differences and impact on model complexity and performance are examined. It is shown that multi-channel kernels provide more stable behaviour, consistent with classical CNN architectures. Subsampling layers adapted for complex-valued data are analysed. Both average and max pooling approaches are considered, and it is shown that averaging is more natural for complex numbers, whereas max pooling requires prior selection of a characteristic (amplitude or phase), which introduces subjectivity into the method. Special attention is given to the mechanisms of error backpropagation. A non-gradient optimization approach in the complex domain is proposed, based on the discrepancy between the desired and actual outputs of the model. Error propagation between fully connected and convolutional layers is described, including cases where subsampling layers are placed between them. To generalize the learning algorithm, the error-sharing principle originally introduced for MLMVN is adapted for CNNMVN, ensuring weight updates that properly account for the multi-valued nature of the neurons. It is demonstrated that convolutional layers in CNNMVN generalize fully connected layers of MLMVN: when the dimensions of the convolutional kernel match those of the input, the convolution operation reproduces the behaviour of a fully connected architecture. The specifics of error correction in convolutional layers are analysed, where a single kernel produces multiple outputs and thus requires aggregation of the resulting errors. A universal correction algorithm based on batch learning in MLMVN is proposed. A separate contribution is a novel approach in which a fully connected MLMVN network is used as a convolutional operator in the frequency domain. The theoretical foundation relies on the convolution theorem, which establishes the equivalence between spatial and frequency-domain representations via the Fourier transform. This allows convolution to be interpreted as element-wise multiplication of the Fourier spectra of the kernel and the signal. Although the method cannot directly support cascaded convolutional layers, it shows high effectiveness in tasks where phase components are derived from pixel intensities. To reduce computational complexity, frequency-domain subsampling is introduced, which uses only a subset of Fourier coefficients and simultaneously improves recognition capability. To assess generalization performance, MATLAB-based software was implemented, and experiments were conducted on the MNIST and Fashion-MNIST datasets. Network topologies with one and two convolutional layers were studied, as well as the behaviour of the adapted subsampling algorithms. The results confirm the efficiency of the models, the convergence of the learning algorithms, and the stable performance of CNNMVN. Experiments further show that at later stages of training, when classification accuracy becomes relatively high, the error level gradually decreases, producing an effect known as “error attenuation.” To mitigate the impact of excessive normalization, a dedicated adaptive learning mechanism is introduced for CNNMVN. This mechanism accounts for both the attenuation of errors in deeper layers and the effect of over-normalization, which often slows learning and reduces generalization ability. It combines scaling of normalization coefficients with adaptive adjustment of weight updates based on current classification accuracy, resulting in more flexible optimization. To deepen understanding of convolutional behaviour in multi-valued neuron networks and their generalization capability, convolution operations and trained kernels were analysed using spectral methods. This analysis enabled the study of kernel behaviour in both spatial and frequency domains. The results show that CNNMVN can extract key spatial features: contours, edges, and structural elements and can also filter frequency components, expanding its applicability to tasks involving spectral data analysis.

Research papers

1. I. Aizenberg and A. Vasko, “Frequency-Domain and Spatial-Domain MLMVN-Based Convolutional Neural Networks”, Algorithms, vol. 17, no. 8, Art. no. 8, Aug. 2024, doi: 10.3390/a17080361.

2. I. Aizenberg and A. Vasko, “Comparative analysis of CNNMVN and MLMVN as frequency domain CNN convolutions”, Bulletin of Taras Shevchenko National University of Kyiv. Physical and Mathematical Sciences, vol. 80, no. 1, pp. 89–96, July 2025, doi: 10.17721/1812-5409.2025/1.12.

3. A. Y. Vasko and A. Y. Bryla, “Adaptive learning rate for CNNMVN”, Науковий вісник Ужгородського університету. Серія «Математика і інформатика», vol. 46, no. 1, pp. 166–177, June 2025, doi: 10.24144/2616-7700.2025.46(1).166-177.

Files

Similar theses