
The color consistency between left and right eye images impacts users’ overall visual experience when viewing near-eye displays (NEDs). This has received insufficient attention in the literature. In this study, experiments were conducted to measure the binocular just-noticeable color difference (JNCD) from the CIE 1976 u′v′ chromaticity diagram. A custom binocular viewing apparatus, which contains no lenses, prisms, or mirrors, was developed to allow for accurate and reliable measurement of the binocular JNCD. Five color stimuli (neutral, magenta, green, blue, and yellow) at three luminance levels (100, 150, and 200 cd/m2) were adopted as the reference stimuli in a series of adjustment experiments. During these experiments, observers were required to determine the binocular JNCD across eight color-change directions under three background luminance conditions (0, 40, and 80 cd/m2). It was found that the binocular JNCD, as measured on the u′v′ diagram, was significantly influenced by the direction of color change and the luminance contrast between the stimulus and the background. Specified by the one-standard-deviation ellipse, the experimental data indicated that the binocular JNCD values correspond to 0.0024, 0.0044, 0.0043, 0.0040, and 0.0038 u′v′ units for neutral, magenta, green, blue, and yellow colors, respectively. These findings should contribute to a deeper understanding of binocular color perception and provide a useful reference for NED color calibration and industrial optical design.

Cardiovascular disease (CVD) remains a major global health burden and the leading cause of mortality worldwide. Although machine learning has shown promise for CVD risk prediction, centralized training on sensitive medical data from multiple hospitals raises serious privacy and security concerns. To address this issue, we propose FedCH, a novel federated learning method for collaborative CVD prediction. Specifically, FedCH quantifies client heterogeneity by measuring the similarity between local and global models together with client-specific prediction accuracy. These metrics are used to dynamically weight the global aggregation process, thereby improving the overall model performance while promoting fairness across clients. In addition, by incorporating historical global model parameters into the update process, FedCH mitigates local model fluctuations and enhances the generalization and robustness of the global model. Comprehensive evaluations on four public datasets and a real-world cardiovascular dataset from a tertiary hospital in China show that FedCH consistently outperforms five state-of-the-art federated learning baselines under heterogeneous client settings.

In this study, the color performance of a polarization camera was investigated and compared with that of a color camera. Three experiments were conducted using a well-calibrated imaging system with a 45∘∕0∘ CIE configuration. First, among the seven preset equations, the 3 × 11 characterization matrix exhibited a small regression error. Second, the use of unpolarized and polarized light led to different chromatic shifts, particularly for sections with 90∘ and 0∘ polarization. Lightness (L∗) was found to be more closely related to the degree of linear polarization (DoLP), whereas chromaticity showed less correlation. Finally, the impact of the two levels of glossy samples was examined using both the cameras. The surface characteristics of the objects were observed and correlated with variations in the DoLP and chromatic shifts. This work analyzes the color characteristics of a polarization camera from FLIR. These findings contribute to understanding the influence of polarized light sources, color characterization conversions, and sample variations on camera color accuracy.

Facial expression recognition (FER) under unconstrained environments faces two severe challenges: class imbalance and the loss of high-frequency texture information. To address these challenges, this study proposes TexGeo-Net, a dual-stream deep architecture that effectively rebalances the training data distribution to mitigate the long-tail effect by employing a landmark-guided geometric augmentation strategy while simultaneously utilizing local binary patterns to precisely preserve micro-texture details. In terms of model design, the system adopts the Densely Connected Convolutional Network as the backbone and deeply integrates the Squeeze-and-Excitation channel attention mechanism. Through adaptive feature fusion, this design compensates for the spatial information loss caused by downsampling in deep networks. Experimental results on three benchmark datasets—CK+, RAF-DB, and FER2013—demonstrate that TexGeo-Net performs competitively with existing methods in terms of recognition accuracy and robustness across these heterogeneous datasets. This validates the complementary advantages of geometric structures and texture features, demonstrating the feasibility of TexGeo-Net in real-world applications.

Event detection is a key task in information extraction that aims to identify event trigger words from text and classify them into event types. Existing methods focus more on the trigger word itself, overlooking the dependencies between contexts and lacking comprehensive feature extraction. Therefore, this paper proposes the model TBEF (TAP-BERT + Event Fusion Layer + CRF). First, the TAP-BERT pretrained model for word vector embedding—which incorporates word embedding matrix decomposition, feed-forward neural network pruning, and self-attention mechanism pooling in BERT—is utilized to obtain richer contextual semantic information. Then an event fusion layer consisting of attention and gate mechanisms is employed to better compute the correlation between events and trigger words. Finally, the sequences are labeled by Conditional Random Field (CRF) integrating BIO labeling. Experimental results show that the proposed method yields an F1 value of 68.55% on the MAVEN dataset, achieving advanced performance for event detection.

Current intelligent inspection methods often lack robustness when identifying urban road defects under poor lighting and blurry conditions. Based on the CRDDC2022_Japan dataset, this study introduces an integrated road-distress detection framework that combines Retinex-S image enhancement with an improved YOLOv8 detector. The enhancement stage performs white balance correction, gamma preprocessing, RGB-to-HSI conversion, SelfDeblur-based restoration on the Intensity channel, wavelet-based Retinex enhancement with MSE-guided reflectance refinement, adaptive saturation adjustment, and HSI-to-RGB reconstruction to recover visibility while preserving details. The detection stage replaces the original YOLOv8s backbone with MobileNetV3, reconstructs the neck with GSConv and CSPHet, and redesigns the head with RepConv while adding a small-object detection branch. On the CRDDC2022_Japan dataset, the proposed method improves mAP@0.5 from 0.643 to 0.692 and reduces the parameter count from 11.1M to 9.3M compared with YOLOv8s while also lowering GFLOPS from 28.8 to 15.7. These results indicate that the proposed framework offers an effective accuracy–efficiency trade-off for road-distress detection under low-illumination and blurred conditions.

Recently, deep learning-based models have been widely applied in medical image segmentation, while those methods often focus on creating deeper and more intricate networks to enhance accuracy, ignoring the inference speed, parameter size, and computational complexity, thus these models were difficult to deploy on resource-constrained devices. In this paper, we build a novel architecture named ECPNet that trade-off well between them by using conditionally parameterized convolutions and factorized convolutions. In the encoding process, the residual module of the decomposed convolutional kernel is used to reduce the number of parameters required for extracting target features, supplement the local or global feature information that may be lost in the network layer during the feature learning process, and alleviate the gradient problem caused by stacking network layers. Moreover, the utilization of conditionally parameterized convolution, consisting of multiple expert knowledge, enables the learning of unique convolutional kernels for each channel in the encoding network, which enhances the feature representation capability of neurons. With a parameter size of only 2.02 M, this model is well suited for deployment on a variety of resource-constrained devices. Finally, extensive experiments conducted on three widely used medical datasets demonstrate that the proposed model achieves a superior balance among accuracy, parameter efficiency, memory usage, and inference speed compared to existing models.

This paper presents a novel approach to enhance unmanned aerial vehicle (UAV) based simultaneous localization and mapping (SLAM) performance in dynamic environments using 3D Gaussian Splatting (3DGS) technology. By leveraging Gaussian kernel functions to represent 3D points, 3DGS dynamically adapts kernel sizes based on local geometric structures, enabling smooth rendering in sparse regions while preserving critical details in high-curvature areas. The Oriented FAST and Rotated BRIEF (ORB) feature fusion integrates image texture into point clouds by mapping 2D ORB keypoints to 3D space using depth maps, significantly enhancing geometric and texture detail representation, particularly at object edges and corners. Robustness is improved through adaptive resizing and statistical outlier removal, ensuring accuracy in repetitive or low-texture regions. Adaptive point sizing refines precision using depth variations while eigenvalue-based curvature metrics extract high-curvature features to enhance geometric clarity. An optimized Gaussian rendering framework prioritizes critical regions via a score-based mechanism, balancing fidelity and resource efficiency. Extensive experiments validate significant improvements. On the TUM dataset, the proposed method achieves a monocular tracking root mean square error (RMSE) of 0.0287 m (vs. 0.0396 m for MonoGS) and an RGB-D tracking RMSE of 0.0134 m (vs. 0.0158 m for MonoGS), demonstrating superior pose estimation accuracy. On the Replica dataset, it attains an average peak signal-to-noise ratio of 39.38 dB (vs. 37.50 dB for MonoGS) and a trajectory RMSE of 0.469 cm (vs. 0.58 cm for MonoGS), confirming enhanced map quality and geometric consistency under dynamic conditions. Real-world UAV tests using DJI Mini 4 data showcase improved detail preservation, structural optimization, and real-time capability in complex indoor environments. Visual comparisons demonstrate that the proposed method produces sharper textures, more accurate edge reconstruction, and better illumination consistency than baseline methods, with particularly notable improvements in challenging regions such as object boundaries and low-texture areas. Our approach consistently outperforms state-of-the-art SLAM methods in accuracy, robustness, and reconstruction fidelity.

Endoscopy images are often constrained by poor color representation. To obtain results that do not alter the image color and quality during effective enhancement, this work proposes a color enhancement algorithm based on color intensity using convolutional neural networks (CNNs) on endoscopic images. The overall process of the algorithm is constructed based on the color loss and uneven color loss phenomena during the endoscopic imaging process. First, to address the issue of color uniformity, this work generates a training set through Gaussian blurring, random attenuation, and image partitioning. It proposes a CNN with four layers that have different training structures and mapping output structures for generating the color intensity map of an image. Second, to tackle the problem of color intensity, a channel loss function is introduced to restore the color intensity of the green and blue channels of the endoscopic image. To address the issue of poor color contrast, histogram equalization under limited contrast is introduced to enhance the contrast of each channel of the endoscopic image. Finally, an image fusion strategy is proposed to merge the channel intensity restoration and contrast enhancement maps with the original image based on the color intensity map.

This paper proposes a multi-stage, feature-aware enhancement and restoration framework tailored for AI-generated traditional Chinese painting images. Addressing prevalent issues in current artificial intelligence generated content (AIGC) systems, including blurred brushwork, missing texture details, color distortion, and noise interference, the method integrates adaptive luminance correction, structure-preserving denoising, and multi-scale detail reconstruction. The framework operates in three sequential stages: (1) an adaptive gamma correction module that normalizes luminance distribution to enhance ink-wash clarity; (2) a novel texture-feature weighted bilateral filtering strategy that dynamically suppresses noise while preserving and enhancing stroke integrity and artistic style; and (3) a multi-scale wavelet fusion module that reconstructs fine textures and hierarchical details. Comprehensive experiments demonstrate that the proposed method significantly improves the clarity, structural integrity, and artistic expressiveness of AIGC-generated Chinese paintings. It outperforms both conventional image processing methods and state-of-the-art learning-based models in objective metrics (PSNR, SSIM, LPIPS) and subjective visual fidelity assessments. Our framework is training-free with zero learnable parameters and real-time inference (0.85 s per 512 × 512 image), making it lightweight and practical for deployment. This work provides a practical and effective solution for enhancing the quality of AI-generated art, facilitating its application in digital cultural heritage.