INTRODUCTION Reach-to-grasp actions are one of the most fundamental forms of interaction with the environment in primates and constitute the primary experimental model for studying visuomotor integration. In visuomotor integration, the visual properties of objects and the visual feedback provided by observing our own movements are combined with proprioceptive signals and continuously transformed into appropriate motor commands(1). Specialized parieto-frontal circuits mediate these transformations during reach-to-grasp. The superior parietal lobule (SPL) integrates visual information from the extrastriate cortex with somatosensory signals regarding hand position, thereby supporting the coordinate transformations necessary for accurate reaching(2). The anterior intraparietal sulcus (aIPS) encodes object properties and grip types, projecting to ventral premotor cortex (PMv), which is particularly sensitive to the hand visual feedback (HVF)(3). Neuroimaging provides the appropriate spatial scale to identify these large-scale networks, but dissociating the contributions of visual and motor processes has proven challenging because of the physical constraints imposed by magnetic resonance. Most fMRI studies of reach-to-grasp have relied on simplified experimental setups, requiring participants to perform pantomimes or imagine moving in the absence of HVF. For these reasons, a comprehensive understanding of how visual feedback of our own movement and motor execution interacts within the parieto-frontal circuits remains elusive. We addressed this challenge using MOTUM (Motion Online Tracking Under MRI)(4), a novel system combining real-time kinematic tracking of arm and hand movements with virtual reality during fMRI. MOTUM enables the provision of online visual feedback of movements in the virtual world, allowing for the systematic manipulation of HVF during fMRI acquisition. This provides unprecedented insight into how visual and motor systems interact during online action control. METHODS 24 right-handed volunteers (15 females; age: 26.6 ± 3.2 years), with normal or corrected-to-normal vision participated in the study. We employed MOTUM at the Neuroimaging Laboratory of the Santa Lucia Foundation in Rome, Italy, using a Magnetom Prisma 3T scanner (Siemens, Erlangen, Germany). Motion capture was achieved using electromagnetically shielded cameras, which tracked hand and forearm position, complemented by an MRI-compatible kinematic glove that captured finger flexion/extension. Kinematics were streamed in real-time into a virtual environment rendered in Unity 6 (Unity Technologies, San Francisco, USA), displayed through MR-compatible goggles (1920×1080 at 60 Hz). For further details on MOTUM, see ref. 4. The virtual environment simulated a scanner room with a table and a custom 3D object which sections could be highlighted in yellow to cue target locations for reach-to-grasp. A humanoid avatar's arm was visible from a first-person perspective, with movements rendered in real-time based on tracking data. We employed a 2×2 factorial block design with movement (executed or not) and hand visual feedback (present or absent) as factors, yielding four conditions: Move Visible (MV), Move Invisible (MI), Stay Visible (SV), and Stay Invisible (SI). Each block began with an auditory instruction ("Move" or "Observe", 0.4s), followed by three trials with jittered interstimulus intervals (5-13s, mean: 7s). Each trial started with a 2.5 s light cue to either a narrow or wide section of the virtual object. When instructed to “Move”, participants performed a precision grip (narrow sections) or a power grip (wide sections), then returned to the starting position. In MV blocks, the virtual hand followed the participant’s movements in real-time, displaying online visual feedback; in MI blocks, the virtual hand was not visible. When instructed to “Observe”, participants remained still. In SV blocks, the virtual hand replayed movements from randomly selected previous MV trials, allowing participants to observe their own hand kinematics without executing movements; in the SI condition, only the cue appeared. The experiment consisted of 4 runs, each comprising 20 blocks arranged in 5 clusters in a different pseudorandom order. Each run lasted approximately 7 minutes and 20 seconds. Kinematic data of the wrist were analyzed. Instantaneous velocity was calculated, smoothed with a 5-frame moving average, and used in a k-means clustering algorithm to identify transition periods between reaching, grip, and back-movement phases. We extracted movement duration, distance traveled and trajectory curvature separately for each phase to assess whether hand visual feedback influenced reach-to-grasp execution. Functional images were acquired using T2*-weighted gradient-echo EPI (multiband factor 6, isotropic 2.4 mm³ voxels, 60 slices, TR=0.8s, TE=0.03s). Preprocessing was performed using fMRIPrep 25.0.0(5), including motion correction, susceptibility distortion correction, coregistration to the T1w reference anatomy, normalization to MNI space, cortical surface reconstruction, volume-to-surface projection, and resampling to fsLR surface space with a 6 mm FWHM smoothing. Statistical analyses were performed in the fsLR surface space using a standard random-effects approach, with voxel-wise general linear models as implemented in SPM12. Each trial was modeled as a boxcar function of 2.5 s (cue duration) convolved with the canonical HRF. For trials with performed or observed movement (MV, MI, SV) we added a linear parametric modulation by the overall duration of the movement. Framewise displacement and five aCompCor regressors were added as confounds. Contrasts examined main effects of movement and visual feedback, and movement by visual feedback interactions. Statistical parametric maps were thresholded at p < 0.05 cluster-level FDR-corrected (cluster-forming threshold p < 0.001). RESULTS Kinematic analyses revealed limited differences between Move Visible and Move Invisible conditions: the arm followed more curved trajectories when hand visual feedback was provided (trajectory ratio relative to an ideal straight-line path: MV 1.34 ± 0.24, MI 1.26 ± 0.19; paired t₂₃ = -2.72, p = 0.012), and the grip phase duration was significantly longer in MV (1032 ± 529 ms) compared to MI (868 ± 469 ms; paired t₂₃ = -3.92, p < 0.001). More curved trajectories suggest continuous online adjustments guided by visual information about hand-object spatial relations, whereas a prolonged grip phase indicates more careful, visually guided hand-object interaction. Reach-to-grasp movements (main effect of movement; [MV+MI]>[SV+SI]) activated a very well-known network of motor regions, including the bilateral primary motor and somatosensory cortex (M1-S1), the dorsal (PMd) and ventral (PMv) premotor cortex, the supplementary motor area (SMA) extending into the anterior cingulum (AAC), the anterior insula (aINS), the superior parietal lobule (SPL) and the intraparietal sulcus (IPS), the putamen, the thalamus, and the cerebellum. Vision of the own moving hand, irrespective of whether participants moved or not (main effect of visual feedback; [MV+SV]>[MI+SI]) yielded extensive bilateral activation in the motion-sensitive area hMT+, PMd, PMv, aIPS, and SPL, and in the right inferior frontal gyrus (IFG). Interestingly, the bilateral fusiform body area (FBA) was also active, reflecting processing of visual body-part information during hand observation. Subcortical activation was prominent throughout the bilateral cerebellum and thalamus. Interaction maps were masked by the set of regions showing significant activation in either MI or SV relative to SI. Superadditive interactions ([MV-MI]>[SV-SI]) — where motor execution combined with visual feedback (MV) resulted in more activity than that predicted by motor execution and visual feedback alone — were observed in three sets of regions showing different activity profiles. Large bilateral swaths in the lateral occipital cortex, including hMT+ and the intra-occipital sulcus, were comparably activated by the presence of visual feedback irrespective of whether participants moved (MV = SV), but were negatively modulated during movement in the absence of visual feedback (MI < SI). Small clusters in the anterior portions of the hemispheres (including left M1-S1, bilateral aINS, bilateral AAC), on the other hand, were comparably activated during movement irrespective of visual feedback (MV = MI), but were negatively modulated when presented with visual feedback alone (SV < SI). Notably, the left posterior cerebellum (lobule VI/Crus I border) was the only region showing activity during MV that significantly exceeded both MI and SV conditions. Subadditive interactions ([MV-MI]<[SV-SI]) — where motor execution combined with visual feedback (MV) resulted in less activity than that predicted by motor execution and visual feedback alone — were observed in a wide set of regions including the bilateral anterior intraparietal sulcus (aIPS), extending into the supramarginal gyrus (SMG) and the postcentral sulcus in the left hemisphere, and the bilateral posterior superior temporal sulcus (pSTS). These regions were activated both by movement alone (MI) and by visual feedback alone (SV) relative to SI, but the effect of movement was weaker in the presence than in the absence of visual feedback. In SMG and aIPS movement with visual feedback (MV) yielded more activity than MI and SV, while in pSTS MV yielded weaker activity than SV. CONCLUSIONS By leveraging MOTUM to independently manipulate motor execution and hand visual feedback within the same paradigm, we provided the first comprehensive mapping of how visual and motor signals are independently processed and jointly integrated during reach-to-grasp actions. Notably, our paradigm used participants’ own movements, recorded online during the task, as visual stimuli for action observation, enabling direct comparison between executing and observing identical self-generated actions. The main effect of hand visual feedback revealed bilateral activation across a distributed visuomotor network encompassing multiple regions that support action execution. This pattern is consistent with previous findings on action observation and execution mirroring(3). Nonetheless, hand visual feedback provides critical guidance for accurate reach-to-grasp. The spatial overlap across parieto-frontal coupled regions is ideally suited to facilitate forward and feedback loops, enabling continuous comparison between predicted and actual hand trajectories(1), allowing visual monitoring to be integrated within motor control circuits. Importantly, we described distinctive ways in which motor and visual information interacted in distinct brain regions. One pattern resulting in a superadditive interaction consisted in an enhancement of activity when both movement and visual feedback were present and was observed in the left cerebellum. Another pattern resulting in a superadditive interaction consisted in the suppression of the response when the non-preferred modality was presented alone (i.e., when purely observing in the sensory-motor cortex, and when moving in the absence of visual feedback in the extrastriate cortex). Finally, other regions showed more activity when movement and visual feedback were both present, but less than expected by the linear combination of the two main effects. This pattern was observed in the association areas in the posterior parietal and temporal cortex and suggests that they are fine-tuned to the predicted sensory consequence of the reach-to-grasp actions, potentially constituting the core mechanisms of forward models and visuomotor integration of reach-to-grasp in humans. Taken together, all these regions are likely to represent the optimized neural circuitry that supports fluent visually-guided reach-to-grasp behavior in everyday contexts. It is important to acknowledge that our study involved interaction with virtual rather than physical objects, without haptic feedback upon object contact. To overcome this limitation, future studies may employ transparent 3D-printed objects that allow arm visibility to motion tracking cameras. In conclusion, by independently manipulating motor execution and hand visual feedback, we precisely characterized the network supporting the observation of self-generated actions during reach-to-grasp. Moreover, we provide evidence for distinct computational strategies and predictive mechanisms operating across distributed brain networks, ranging from model updating to the reliance on optimized neural pathways for visuomotor integration. These results advance our understanding of the neural architecture supporting skilled environment-directed behavior and demonstrate the potential of the MOTUM system for investigating sensorimotor integration with unprecedented ecological validity and experimental control. REFERENCES 1 Kawato M. Internal models for motor control and trajectory planning. Curr Opin Neurobiol. 1999;9(6):718-727. doi:10.1016/s0959-4388(99)00028-8 2 Hadjidimitrakis K, Bertozzi F, Breveglieri R, Bosco A, Galletti C, Fattori P. Common neural substrate for processing depth and direction signals for reaching in the monkey medial posterior parietal cortex. Cereb Cortex. 2014;24(6):1645-1657. doi:10.1093/cercor/bht021 3 Maranesi M, Livi A, Bonini L. Processing of Own Hand Visual Feedback during Object Grasping in Ventral Premotor Mirror Neurons. J Neurosci. 2015;35(34):11824-11829. doi:10.1523/JNEUROSCI.0301-15.2015 4 Bencivenga F, Tani M, Vyas K, Giove F, Gazzitano S, Galati G. MOTUM: a system for Motion Online Tracking Under MRI. Imaging Neuroscience. Published online December 9, 2025. doi:10.1162/IMAG.a.1081 5 Esteban O, Markiewicz CJ, Blair RW, et al. fMRIPrep: a robust preprocessing pipeline for functional MRI. Nat Methods. 2019;16(1):111-116. doi:10.1038/s41592-018-0235-4
The role of hand visual feedback in grasping actions: an fMRI study using motion online tracking / Tani, M., Ferruzzi, S., Costanzo, R., Perrone, M., Galati, G.. - (2026). (Organization for Human Brain Mapping (OHBM) Bordeaux; France ) [10.5281/zenodo.20817576].
The role of hand visual feedback in grasping actions: an fMRI study using motion online tracking
Michelangelo Tani;Martina Perrone;Gaspare Galati
2026
Abstract
INTRODUCTION Reach-to-grasp actions are one of the most fundamental forms of interaction with the environment in primates and constitute the primary experimental model for studying visuomotor integration. In visuomotor integration, the visual properties of objects and the visual feedback provided by observing our own movements are combined with proprioceptive signals and continuously transformed into appropriate motor commands(1). Specialized parieto-frontal circuits mediate these transformations during reach-to-grasp. The superior parietal lobule (SPL) integrates visual information from the extrastriate cortex with somatosensory signals regarding hand position, thereby supporting the coordinate transformations necessary for accurate reaching(2). The anterior intraparietal sulcus (aIPS) encodes object properties and grip types, projecting to ventral premotor cortex (PMv), which is particularly sensitive to the hand visual feedback (HVF)(3). Neuroimaging provides the appropriate spatial scale to identify these large-scale networks, but dissociating the contributions of visual and motor processes has proven challenging because of the physical constraints imposed by magnetic resonance. Most fMRI studies of reach-to-grasp have relied on simplified experimental setups, requiring participants to perform pantomimes or imagine moving in the absence of HVF. For these reasons, a comprehensive understanding of how visual feedback of our own movement and motor execution interacts within the parieto-frontal circuits remains elusive. We addressed this challenge using MOTUM (Motion Online Tracking Under MRI)(4), a novel system combining real-time kinematic tracking of arm and hand movements with virtual reality during fMRI. MOTUM enables the provision of online visual feedback of movements in the virtual world, allowing for the systematic manipulation of HVF during fMRI acquisition. This provides unprecedented insight into how visual and motor systems interact during online action control. METHODS 24 right-handed volunteers (15 females; age: 26.6 ± 3.2 years), with normal or corrected-to-normal vision participated in the study. We employed MOTUM at the Neuroimaging Laboratory of the Santa Lucia Foundation in Rome, Italy, using a Magnetom Prisma 3T scanner (Siemens, Erlangen, Germany). Motion capture was achieved using electromagnetically shielded cameras, which tracked hand and forearm position, complemented by an MRI-compatible kinematic glove that captured finger flexion/extension. Kinematics were streamed in real-time into a virtual environment rendered in Unity 6 (Unity Technologies, San Francisco, USA), displayed through MR-compatible goggles (1920×1080 at 60 Hz). For further details on MOTUM, see ref. 4. The virtual environment simulated a scanner room with a table and a custom 3D object which sections could be highlighted in yellow to cue target locations for reach-to-grasp. A humanoid avatar's arm was visible from a first-person perspective, with movements rendered in real-time based on tracking data. We employed a 2×2 factorial block design with movement (executed or not) and hand visual feedback (present or absent) as factors, yielding four conditions: Move Visible (MV), Move Invisible (MI), Stay Visible (SV), and Stay Invisible (SI). Each block began with an auditory instruction ("Move" or "Observe", 0.4s), followed by three trials with jittered interstimulus intervals (5-13s, mean: 7s). Each trial started with a 2.5 s light cue to either a narrow or wide section of the virtual object. When instructed to “Move”, participants performed a precision grip (narrow sections) or a power grip (wide sections), then returned to the starting position. In MV blocks, the virtual hand followed the participant’s movements in real-time, displaying online visual feedback; in MI blocks, the virtual hand was not visible. When instructed to “Observe”, participants remained still. In SV blocks, the virtual hand replayed movements from randomly selected previous MV trials, allowing participants to observe their own hand kinematics without executing movements; in the SI condition, only the cue appeared. The experiment consisted of 4 runs, each comprising 20 blocks arranged in 5 clusters in a different pseudorandom order. Each run lasted approximately 7 minutes and 20 seconds. Kinematic data of the wrist were analyzed. Instantaneous velocity was calculated, smoothed with a 5-frame moving average, and used in a k-means clustering algorithm to identify transition periods between reaching, grip, and back-movement phases. We extracted movement duration, distance traveled and trajectory curvature separately for each phase to assess whether hand visual feedback influenced reach-to-grasp execution. Functional images were acquired using T2*-weighted gradient-echo EPI (multiband factor 6, isotropic 2.4 mm³ voxels, 60 slices, TR=0.8s, TE=0.03s). Preprocessing was performed using fMRIPrep 25.0.0(5), including motion correction, susceptibility distortion correction, coregistration to the T1w reference anatomy, normalization to MNI space, cortical surface reconstruction, volume-to-surface projection, and resampling to fsLR surface space with a 6 mm FWHM smoothing. Statistical analyses were performed in the fsLR surface space using a standard random-effects approach, with voxel-wise general linear models as implemented in SPM12. Each trial was modeled as a boxcar function of 2.5 s (cue duration) convolved with the canonical HRF. For trials with performed or observed movement (MV, MI, SV) we added a linear parametric modulation by the overall duration of the movement. Framewise displacement and five aCompCor regressors were added as confounds. Contrasts examined main effects of movement and visual feedback, and movement by visual feedback interactions. Statistical parametric maps were thresholded at p < 0.05 cluster-level FDR-corrected (cluster-forming threshold p < 0.001). RESULTS Kinematic analyses revealed limited differences between Move Visible and Move Invisible conditions: the arm followed more curved trajectories when hand visual feedback was provided (trajectory ratio relative to an ideal straight-line path: MV 1.34 ± 0.24, MI 1.26 ± 0.19; paired t₂₃ = -2.72, p = 0.012), and the grip phase duration was significantly longer in MV (1032 ± 529 ms) compared to MI (868 ± 469 ms; paired t₂₃ = -3.92, p < 0.001). More curved trajectories suggest continuous online adjustments guided by visual information about hand-object spatial relations, whereas a prolonged grip phase indicates more careful, visually guided hand-object interaction. Reach-to-grasp movements (main effect of movement; [MV+MI]>[SV+SI]) activated a very well-known network of motor regions, including the bilateral primary motor and somatosensory cortex (M1-S1), the dorsal (PMd) and ventral (PMv) premotor cortex, the supplementary motor area (SMA) extending into the anterior cingulum (AAC), the anterior insula (aINS), the superior parietal lobule (SPL) and the intraparietal sulcus (IPS), the putamen, the thalamus, and the cerebellum. Vision of the own moving hand, irrespective of whether participants moved or not (main effect of visual feedback; [MV+SV]>[MI+SI]) yielded extensive bilateral activation in the motion-sensitive area hMT+, PMd, PMv, aIPS, and SPL, and in the right inferior frontal gyrus (IFG). Interestingly, the bilateral fusiform body area (FBA) was also active, reflecting processing of visual body-part information during hand observation. Subcortical activation was prominent throughout the bilateral cerebellum and thalamus. Interaction maps were masked by the set of regions showing significant activation in either MI or SV relative to SI. Superadditive interactions ([MV-MI]>[SV-SI]) — where motor execution combined with visual feedback (MV) resulted in more activity than that predicted by motor execution and visual feedback alone — were observed in three sets of regions showing different activity profiles. Large bilateral swaths in the lateral occipital cortex, including hMT+ and the intra-occipital sulcus, were comparably activated by the presence of visual feedback irrespective of whether participants moved (MV = SV), but were negatively modulated during movement in the absence of visual feedback (MI < SI). Small clusters in the anterior portions of the hemispheres (including left M1-S1, bilateral aINS, bilateral AAC), on the other hand, were comparably activated during movement irrespective of visual feedback (MV = MI), but were negatively modulated when presented with visual feedback alone (SV < SI). Notably, the left posterior cerebellum (lobule VI/Crus I border) was the only region showing activity during MV that significantly exceeded both MI and SV conditions. Subadditive interactions ([MV-MI]<[SV-SI]) — where motor execution combined with visual feedback (MV) resulted in less activity than that predicted by motor execution and visual feedback alone — were observed in a wide set of regions including the bilateral anterior intraparietal sulcus (aIPS), extending into the supramarginal gyrus (SMG) and the postcentral sulcus in the left hemisphere, and the bilateral posterior superior temporal sulcus (pSTS). These regions were activated both by movement alone (MI) and by visual feedback alone (SV) relative to SI, but the effect of movement was weaker in the presence than in the absence of visual feedback. In SMG and aIPS movement with visual feedback (MV) yielded more activity than MI and SV, while in pSTS MV yielded weaker activity than SV. CONCLUSIONS By leveraging MOTUM to independently manipulate motor execution and hand visual feedback within the same paradigm, we provided the first comprehensive mapping of how visual and motor signals are independently processed and jointly integrated during reach-to-grasp actions. Notably, our paradigm used participants’ own movements, recorded online during the task, as visual stimuli for action observation, enabling direct comparison between executing and observing identical self-generated actions. The main effect of hand visual feedback revealed bilateral activation across a distributed visuomotor network encompassing multiple regions that support action execution. This pattern is consistent with previous findings on action observation and execution mirroring(3). Nonetheless, hand visual feedback provides critical guidance for accurate reach-to-grasp. The spatial overlap across parieto-frontal coupled regions is ideally suited to facilitate forward and feedback loops, enabling continuous comparison between predicted and actual hand trajectories(1), allowing visual monitoring to be integrated within motor control circuits. Importantly, we described distinctive ways in which motor and visual information interacted in distinct brain regions. One pattern resulting in a superadditive interaction consisted in an enhancement of activity when both movement and visual feedback were present and was observed in the left cerebellum. Another pattern resulting in a superadditive interaction consisted in the suppression of the response when the non-preferred modality was presented alone (i.e., when purely observing in the sensory-motor cortex, and when moving in the absence of visual feedback in the extrastriate cortex). Finally, other regions showed more activity when movement and visual feedback were both present, but less than expected by the linear combination of the two main effects. This pattern was observed in the association areas in the posterior parietal and temporal cortex and suggests that they are fine-tuned to the predicted sensory consequence of the reach-to-grasp actions, potentially constituting the core mechanisms of forward models and visuomotor integration of reach-to-grasp in humans. Taken together, all these regions are likely to represent the optimized neural circuitry that supports fluent visually-guided reach-to-grasp behavior in everyday contexts. It is important to acknowledge that our study involved interaction with virtual rather than physical objects, without haptic feedback upon object contact. To overcome this limitation, future studies may employ transparent 3D-printed objects that allow arm visibility to motion tracking cameras. In conclusion, by independently manipulating motor execution and hand visual feedback, we precisely characterized the network supporting the observation of self-generated actions during reach-to-grasp. Moreover, we provide evidence for distinct computational strategies and predictive mechanisms operating across distributed brain networks, ranging from model updating to the reliance on optimized neural pathways for visuomotor integration. These results advance our understanding of the neural architecture supporting skilled environment-directed behavior and demonstrate the potential of the MOTUM system for investigating sensorimotor integration with unprecedented ecological validity and experimental control. REFERENCES 1 Kawato M. Internal models for motor control and trajectory planning. Curr Opin Neurobiol. 1999;9(6):718-727. doi:10.1016/s0959-4388(99)00028-8 2 Hadjidimitrakis K, Bertozzi F, Breveglieri R, Bosco A, Galletti C, Fattori P. Common neural substrate for processing depth and direction signals for reaching in the monkey medial posterior parietal cortex. Cereb Cortex. 2014;24(6):1645-1657. doi:10.1093/cercor/bht021 3 Maranesi M, Livi A, Bonini L. Processing of Own Hand Visual Feedback during Object Grasping in Ventral Premotor Mirror Neurons. J Neurosci. 2015;35(34):11824-11829. doi:10.1523/JNEUROSCI.0301-15.2015 4 Bencivenga F, Tani M, Vyas K, Giove F, Gazzitano S, Galati G. MOTUM: a system for Motion Online Tracking Under MRI. Imaging Neuroscience. Published online December 9, 2025. doi:10.1162/IMAG.a.1081 5 Esteban O, Markiewicz CJ, Blair RW, et al. fMRIPrep: a robust preprocessing pipeline for functional MRI. Nat Methods. 2019;16(1):111-116. doi:10.1038/s41592-018-0235-4I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


