Image Details
Caption: Figure 2.
Architecture of a multimodal fusion model using Bayesian neural networks. The abbreviations are defined as follows: “BBConv1D” and “BBConv2D” denote Bayesian 1D and 2D convolutional layers, respectively, where “BB” stands for Bayes-by-Backprop, the variational inference method used for weight sampling. “BN” refers to batch normalization layers. “ReLU” is the rectified linear unit activation function. “CAM” represents the channel attention module, which uses adaptive average pooling followed by linear layers with sigmoid activation to compute channel-wise attention weights. “BBLinear” denotes a Bayesian fully connected (linear) layer. The notation “16 × 128” and “16 × 124” indicates the number of filters and the feature map size at that layer. The “(3 × layer) × 3” notation at the bottom indicates that the 2D branch consists of three convolutional blocks, each containing three convolutional layers. The ⊕ symbol represents feature concatenation across modalities before the final classification layers.
© 2026. The Author(s). Published by the American Astronomical Society.