
Optical coherence tomography (OCT) is widely used for retinal disease diagnosis, motivating extensive research into deep-learning-based classification. Recent studies increasingly propose hybrid detection–classification architectures, assuming that explicit region-of-interest (ROI) localization improves diagnostic performance. However, the effectiveness of such hybrid pipelines for OCT disease classification has not been systematically examined. In this work, a controlled empirical investigation of solo and hybrid deep learning architectures, including convolutional neural networks (CNNs), Vision Transformers (ViTs), YOLO-based models, and hybrid detection–classification pipelines, is conducted for OCT image classification. Using multiple publicly available OCT datasets, performance, convergence behavior, stability, and cross-dataset robustness are evaluated under consistent preprocessing, optimization, and computational constraints. The results show that hybrid pipelines do not consistently outperform pretrained end-to-end models and often exhibit increased variance, training instability, or performance degradation. These findings highlight the importance of empirically validating architectural assumptions and suggest that increased complexity does not inherently improve OCT disease classification.
Ning Yu, James Finn, Arnulfo Torres, "An Empirical Investigation of Hybrid Detection–Classification Architectures for OCT Image Classification" in Journal of Imaging Science and Technology, 2026, pp 1 - 12, https://doi.org/10.2352/J.ImagingSci.Technol.2026.70.5.050501