
VR research rarely fails in dramatic ways. It frays in the margins. It slips. It drifts. A timestamp is off. A controller desyncs for a moment. A study crashes right as a participant reaches the final task. These moments feel small and forgettable, yet they accumulate. In this paper, we present an exploration into the fusion of Fulcrum, a generalized support system for user-facing studies, and a new and improved ScryVR, a Unity package built to support user-facing VR studies. The combined system offers clear templates, organized study structures, and simple building blocks that help researchers understand and adjust their experiments without getting lost in technical details. To evaluate this solution, two case studies were conducted with this setup as the main tool in their experiment. Ultimately, the system effectiveness greatly streamlined VR setup and automated many crucial aspects of experiment facilitation. However, limitations were made clear primarily in robust error reporting systems, which report in real time and a more diverse solution for analytical tasks.

Advanced technologies not only produce external transformations in society but also elicit significant emotional responses in individuals. Immersive psychological technology, which leverages the cognitive effects of Extended Reality, Artificial Intelligence, and Virtual Twin Avatars, is the methodology adopted by SITh (Self-Insight Therapy), based on Gestalt psychotherapy—a therapeutic approach emphasizing ‘here-and-now’ awareness and ‘human-in-the-loop’—serves as a prominent example. This study presents a case study of XR-based mourning ritual service leveraging immersive psychological technology to support bereaved family members in coping with grief after a tragic aircraft disaster in South Korea in December 2024. The findings illustrate how technology can facilitate safe third-party intervention in the grief recovery process for individuals who have lost family members. The XR-based mourning ritual case study, combined with an interview with a bereaved family member, indicates a promising direction for advancing traditional healing methods through immersive psychological technology integration.

Virtual Reality (VR) has recently attracted more attention in mental health applications due to its ability to immerse users in controlled and interactive environments. This paper presents a VR-Based AI Mental Health Companion, a multimodal system designed to support therapy, mindfulness, and real-time stress detection within immersive virtual reality (VR) environments. The system integrates artificial intelligence (AI) techniques by including natural language processing, emotion recognition, and physiological signal analysis by creating personalized mindfulness experiences and interactive meditation coaching. GPT-powered non-player characters (NPCs) are designed with specific therapeutic roles in mind, such as guided mindfulness facilitation, emotional support, and stress-aware conversational therapy. The system employs pose estimation to identify key body points and apply rule-based logic to assess posture accuracy during guided yoga exercises, providing real-time feedback to support correct movement execution. The work also includes biometric integration, such as EEG monitoring for enhanced emotional sensing. Expanding language support, increasing the diversity of pose datasets, and incorporating feedback from clinical professionals help refine the system. By combining immersive VR environments, GPT-driven therapeutic NPCs, and real-time posture validation, the proposed VR-Based AI Mental Health Companion demonstrates the potential of AI–VR convergence as a scalable approach to mental health care, with promising applications in stress management and preventative therapy.

This study introduces a simulation framework designed to examine epidemic communication and behavioral interventions utilizing AI-driven non-player characters (NPCs) within a 3D environment created in Unity. The framework rectifies the limitations of conventional epidemiological models by integrating various agents that exhibit adaptive and context sensitive decision-making capabilities. Agents employ large language models (LLMs) and behavior trees to facilitate realistic conversations and responses in epidemic scenarios, contrasting with static rule-based systems. This results in interactions that closely resemble real-world human communication. The simulation enables real-time communication between agents and users in natural language. There are different ways that public health interventions, like social distance measures and communication attempts, can be used and evaluated. The technology enables agents to know what’s going on around them and how far away other people are, so they can act in the right way. The rendering engine in Unity makes the game more realistic, which makes it more interesting and useful. This study shows that agents were able to take part in COVID-19-related conversations using GPT and Convai and give appropriate answers to user questions. The framework makes it easy to do experiments on a large scale and can be used in many different public health settings. Future improvements will include simulating emotional states, making agents more diverse, and adding visual health indicators. This study introduces a scalable, ethical, and interactive instrument intended for researchers to examine human behavior, decision making, and intervention outcomes in simulated epidemic scenarios.

Typing remains a primary mode of interaction in extended reality (XR), yet most XR systems still depend on physical keyboards or specialized sensing hardware. We present a visiononly keystroke tracking pipeline for surface typing that combines visible-key detection, recovery of hand-occluded keys, and finger keypress event detection on monocular video. The system first detects the keyboard and visible keys with a YOLO-based detector trained on oriented key boxes and synthetic hand-occlusion augmentation. Occluded or missed keys are then recovered with lightweight per-key regressors driven by an efficient 3–2 recovery-key selection strategy. For keypress detection, we use monocular 3D hand reconstruction with HaMeR and an autoregressive transformer that processes temporal image sequences, recovered hand meshes, finger keypoint masks, and camera poses to estimate press-down probabilities for each finger. On the Keyboard Key Detection dataset, our visible-key detector achieves 97% mAP@75 with mean IoU around 0.9, while occluded-key recovery sustains 30 fps. On typing videos from the MSU Typing Behavior Database, the keypress module reaches up to 90% event detection accuracy with 28 ms latency. Together, these components form a practical end-to-end vision system for hardware-free keystroke tracking on printed, projected, or rendered keyboard layouts.

Monolingual learners frequently encounter barriers to language acquisition ranging from financial constraints to a lack of situational confidence. Virtual Reality (VR) offers a promising solution by providing a ”safe” digital environment for immersive learning experiences. This paper evaluates a comparative study between two distinct delivery methods of a language lecture within VR: a traditional video presentation and a 3D-modeled experience utilizing consumer-grade motion capture hardware. In addition, this paper provides a solution with cheap consumer motion capture hardware, addressing the financial block above, to create educational content. Overall, results were mixed; while both the video and motion capture versions yielded positive engagement, the video format demonstrated a more significant quantitative increase in results. As part of the evaluation, we analyze the performance of low-cost motion capture in educational content creation and propose design iterations to better isolate the variables influencing these learning gains.

Extended Reality (XR) virtual environment (VE) development is heavily focused on visual models, artifacts and interactions. This reliance on visuals in augmented, virtual, or mixed reality (i.e., AR, VR, or MR) can reduce clarity and overload cognitive resources, especially for uses in education and training, for learners with diverse sensory needs. The next most straightforward sensory modality to add is sound, which is often limited to a passive background effect. However, with research minded implementation from the field of psychoacoustics, sound can be used as an effective interaction method in a VE, while simultaneously increasing effects such as presence and immersion. This research describes the development challenges and solutions encountered in building spatial audio as an active, instructional feedback tool in a STEM focused VR environment. The VE was designed to introduce middle and high school students to basic concepts of how the Internet functions (e.g., IP addresses, data packet routing). By embedding meaningful auditory cues and feedback, the system supports learners by reinforcing task completion and guiding attention. The sound design leverages psychoacoustic principles and spatialization techniques, which were iteratively refined with learner and expert input.

This study aims to clarify the role of sound in evoking the sublime experience within a virtual reality (VR) environment. The sublime is a complex emotion combining awe and fear, arising from vast objects or overwhelming forces. VR is considered an effective medium for safely inducing this experience. However, existing research has predominantly focused on visual factors, and the influence of auditory stimuli—essential for immersion—remains insufficiently explored. In this study, a 3D 360° video of a volcanic crater was presented via a head-mounted display (HMD) under three auditory conditions: Silent, Normal (natural environmental sound), and Reverbed (processed sound). We evaluated the experience using subjective measures (Awe Experience Scale) and objective physiological measures (Electrodermal Activity, pupil diameter, gaze data). The results demonstrated that the presence of sound significantly amplified the sublime experience across both subjective and objective indices. Specifically, the Normal condition showed high integration with visual information, eliciting the strongest emotional arousal and significant pupil dilation. Conversely, while the Reverbed condition induced spatial exploratory behavior (gaze dispersion), it caused a sense of incongruence between sight and sound, tending to lower the quality of the experience. These findings suggest that audio-visual congruence is critical in designing sublime experiences in VR.