
Optical coherence tomography (OCT) is widely used for retinal disease diagnosis, motivating extensive research into deep-learning-based classification. Recent studies increasingly propose hybrid detection–classification architectures, assuming that explicit region-of-interest (ROI) localization improves diagnostic performance. However, the effectiveness of such hybrid pipelines for OCT disease classification has not been systematically examined. In this work, a controlled empirical investigation of solo and hybrid deep learning architectures, including convolutional neural networks (CNNs), Vision Transformers (ViTs), YOLO-based models, and hybrid detection–classification pipelines, is conducted for OCT image classification. Using multiple publicly available OCT datasets, performance, convergence behavior, stability, and cross-dataset robustness are evaluated under consistent preprocessing, optimization, and computational constraints. The results show that hybrid pipelines do not consistently outperform pretrained end-to-end models and often exhibit increased variance, training instability, or performance degradation. These findings highlight the importance of empirically validating architectural assumptions and suggest that increased complexity does not inherently improve OCT disease classification.

Event-based Vision Sensors (EVSs) exhibit advantages such as high temporal resolution and low data redundancy, rendering them promising for applications in high-speed vision and edge computing scenarios. However, the sparsity and irregularity of event data in both spatial and temporal domains, coupled with the presence of noise, make it difficult for traditional processing methods based on full-pixel sampling and regular memory access to be efficiently adapted to such data, which restricts the real-time performance of EVS-based systems. To address this problem, this paper proposes a software–hardware co-design scheme for event image processing. On the software side, event streams are generated from frame-by-frame luminance differences, and a combination of density clustering, multi-level denoising, and convolution kernel weighting processing is adopted to effectively compress the event scale and enhance structural stability. On the hardware side, a sparse convolution acceleration architecture is designed, where a sparse controller performs block-wise scheduling of multiple processing element units to execute convolution operations only on valid data regions. Experimental results demonstrate that the software-based sparsification processing reduces the number of pixels involved in computation by approximately 95% compared with the unprocessed case; under the same convolution kernel settings, the overall computation achieves a speedup of about 15–20 times relative to hardware acceleration-free schemes. In the FPGA implementation, the system attains a maximum operating frequency of 260 MHz, which verifies the feasibility of the proposed software–hardware co-design architecture for real-time event image processing.

Traditional CNN-based bearing fault diagnosis methods are constrained by single-scale convolution kernels, which cannot extract multi-scale fault features, leading to low feature distinguishability, weak generalization ability, and poor noise resistance. To address these issues, this paper proposes a hybrid intelligent fault diagnosis framework named MSCNN–BiGRU–DLSSVM. The model adopts a multi-scale spatiotemporal fusion network for multi-scale spatiotemporal feature complementary learning: the first branch converts 1D vibration signals into 2D images via Markov transition field encoding and feeds them into a customized Multi-Scale Convolutional Neural Network (MSCNN) to extract multi-scale discriminative spatial fault features; the second branch takes raw temporal sequences as input and uses a Bidirectional Gated Recurrent Unit (BiGRU) to capture bidirectional multi-scale temporal dependencies of signals. The complementary multi-scale spatiotemporal features from the two branches are fused and fed into a Differentiable Least-Squares Support Vector Machine (DLSSVM) for fault classification. The DLSSVM achieves end-to-end joint optimization with the dual-branch backbone through reparameterization and differentiable loss design. Comprehensive validation on the public CWRU bearing dataset shows that the proposed framework provides a robust and practical approach for bearing fault diagnosis.

As an immersive imaging modality, point clouds (PCs) sustain compression distortions that degrade visual quality and impair the perceptual experience. To mitigate these distortions and enhance the visual quality of PCs, a synergistic 2D–3D fusion for distortion restoration in G-PCC Trisoup encoded PCs is proposed. To address the challenge of detecting multi-scale and irregularly shaped triangular hole distortions caused by G-PCC Trisoup encoding, which is difficult to perform directly in 3D, a residual-based detection method is proposed. Here, the 3D data is projected onto 2D to meet the detection requirements for triangular holes with varying distortion levels. Considering that screened Poisson surface reconstruction provides more accurate 3D reconstruction results and that estimating depth information of triangular hole regions solely through restoring 2D geometry projection maps may incur significant errors, a joint 2D and 3D geometry estimation method is proposed. This method utilizes the reconstructed 3D information to filter errors in the 2D restoration and combines 2D and 3D estimates by averaging, thereby achieving accurate geometry information estimation for triangular hole regions in PCs. Moreover, relying solely on the restoration of 2D texture projection maps to achieve 3D color correction leads to accuracy limitations; hence a joint 2D and 3D color estimation method is proposed to realize smooth color transitions between triangular hole regions and non-hole regions. Experimental results on multiple PC models from three different datasets demonstrate that the proposed method achieves superior performance in restoring both geometry and texture distortions of G-PCC Trisoup encoded PCs. Compared to the distorted inputs, the restored results exhibit an average improvement of 1.36 dB in both geometry and color quality.

Material segmentation hinges on pixel-level cues in hue, chroma, and lightness, yet most pipelines remain in device-referred sRGB and use ad hoc photometric tweaks. The authors disentangle the effects of color encoding versus principled, opponent-axis augmentations by training the same FCN–ResNet101 across a 3 × 4 grid: three encodings (sRGB, CIELAB, YCbCr) and three colorimetric perturbations—luminance translation (brightness), luminance scaling (contrast), and chroma-magnitude scaling (saturation)—plus a no-augmentation baseline. Using overlap-aware metrics (intersection over union; F1) and confirmatory mixed-effects analyses, they find that augmentation—not encoding—is the primary driver of accuracy. Contrast (luminance scaling) delivers the most reliable global gains; brightness is especially effective for specular materials (metal, glass) while saturation benefits chroma-dominant diffuse classes (paper, fruits). Color encoding acts as a secondary modulator: YCbCr’s luminance–chrominance separation sharpens neutral or composite boundaries (e.g., plastic), whereas CIELAB’s approximate perceptual uniformity aids materials with broader chroma spread. Qualitative overlays and backbone-feature embeddings via t-distributed stochastic neighbor embedding (t-SNE) corroborate these trends, showing cleaner boundaries and more separable latent clusters under luminance-scaled training. Practically, the authors recommend contrast as the safe default, favoring YCbCr when specular cues dominate and CIELAB when chroma carries the signal.

Magnetic particle imaging (MPI) is an emerging preclinical tomography technology with high resolution and sensitivity. However, MPI reconstruction artifacts affect the imaging accuracy and quality, which may lead to misdiagnosis or missed diagnosis of disease. In this study, a modified reconstruction method of MPI is proposed, which is called non-negative weighted least absolute shrinkage and selection operator (LASSO) regularization (NWLR). After the reconstruction based on NWLR, the image post-processing of calibration based on normalized cross-correlation (NCC) and artifact suppression based on the deep image prior (DIP) are conducted to further improve the reconstruction quality. To demonstrate the advantages of the proposed method, the reconstruction result is compared with other three popular methods with respect to both visualization and performance indexes. The results show that the proposed NWLR method yields fewer artifacts and higher accuracy than the other three methods. Furthermore, the PSNR and SSIM are improved by 13.73% and 20.10% through post-processing, respectively. The NRMSE of the proposed NWLR–NCC–DIP method is decreased by 32.07%. The conducted studies demonstrate that the NWLR–NCC–DIP method proposed in this study achieves better reconstruction performance than the other popular methods. The reconstruction quality and accuracy are improved rapidly, which is of great practical importance for its preclinical applications.

This study proposes a human dynamic behavior recognition method based on joint point extraction and deep learning algorithms. Human skeletal information is collected using a Kinect camera to obtain three-dimensional (3D) joint coordinates, completing the extraction of skeletal joint data. Based on the characteristics of bones in 3D space, processing is performed using 3D skeletal joint point cloud data. An improved PointNet++ method is employed to process the point cloud data: a dual-scale feature extraction strategy is adopted to enhance the network’s multiscale feature capture capability, and the farthest point sampling algorithm is optimized to better preserve behavioral details. Finally, a graph convolutional network approach incorporating a channel attention mechanism and graph topology optimization is used to achieve human dynamic behavior recognition based on 3D skeletons. Experimental results show that the method achieves a maximum F1 score of 0.98, a misrecognition rate below 0.4%, and a coefficient of variation of 0, demonstrating high recognition accuracy. However, the performance of this research method may be limited in complex scenes with severe occlusion or non-standard perspectives. Future work will focus on exploring multimodal data fusion and real-time optimization to further enhance its robustness and practicality in open environments.

Hyperspectral super-resolution fusion technology aims to fuse hyperspectral images with multispectral images in the same scene for the super-resolution reconstruction of hyperspectral data. Current deep learning methods are usually trained by data augmentation or constructing complex encoding–decoding networks, often neglecting the physical characteristics of hyperspectral data involving width, height, and three-dimensional channel information. Traditional methods continue to play a role in super-resolution reconstruction although there are deficiencies in the fusion result. For this reason, this paper proposes a Quadratic Optimization Model (QOM) that combines deep learning and traditional mathematical methods. The model first utilizes a three-module neural network for initial fusion; designs corresponding modules for the recovery of dimensional, spatial, and spectral information; and introduces spatial and channel attention mechanisms to enhance feature extraction capability. Subsequently, the preliminary fusion results are optimized by secondary super-resolution through the traditional matrix decomposition method to further improve fusion quality. The experimental results demonstrate that the QOM achieves excellent performance on all seven datasets, exhibiting strong fusion quality while maintaining favorable computational complexity (TFLOPS: 8.6024, Params: 2.9307). Noise experiments verify its high robustness.

Monocular depth estimation (MDE) is a widely used technique in autonomous driving and 3D reconstruction. However, inconsistent and fragmented depth outputs can significantly undermine the reliability of MDE applications in practice. To address this issue, the authors introduce MonoHybrid, a novel self-supervised hybrid network that effectively integrates Transformer and dilated convolutional architectures. This design enables the extraction of both global and local features, enhancing the receptive field and ensuring robust and continuous depth estimation. Additionally, the authors present a new Feature Fusion Module that fuses convolutional and Transformer features, resulting in improved depth estimation performance. Through comprehensive experiments, the proposed network demonstrates notable accuracy and generalization compared to other advanced methods in the field.

Unsupervised visible–infrared person re-identification (USVI-ReID) is a very important and challenging task in machine vision. The key challenge of USVI-ReID is to effectively mine weak class-wise supervision and establish cross-modal correspondences without using any manual annotations. In this paper, the authors propose a soft prototype contrastive learning and instance discrimination method for USVI-ReID. Specifically, soft prototype contrastive learning selects the nearest neighbors with high similarity to the soft prototypes to mine accurate information and guide the model to learn more discriminative features. On this basis, a soft weighting strategy is used to quantitatively measure the relevance of the selected soft prototypes relative to the current centroid prototype, thus further eliminating the interference of the wrong prototype in the model training. To overcome the problems of image noise and complex backgrounds in visible and infrared images, instance discriminative learning is first integrated into USVI-ReID to explore the potential similarity relationship between instances from the bottom up and learn discriminative representations. Finally, the authors propose a progressive training strategy, which enables the model to learn the similarities between instances in the early stage of training and gradually shift its attention to more discriminative categories in the later stage. Extensive experiments are conducted on two public datasets, and quantitative results prove the effectiveness of the proposed method.