Repository navigation
Conversation
Co-authored-by: RayStick <50215726+RayStick@users.noreply.github.com> Co-authored-by: isesteves <24605642+isesteves@users.noreply.github.com> Co-authored-by: Sarah Goodale <60117796+goodalse2019@users.noreply.github.com> Co-authored-by: m-miedema <39968233+m-miedema@users.noreply.github.com>
|
Small note - the automatic release on Majormod would bump version to 2025, while we need to release at 2026.0.0. I'd suggest a manual release for this. |
Okay thanks - removed the |
mrikasper
left a comment
There was a problem hiding this comment.
Great work! I have a couple of comments, which I am happy to provide change requests/commits myself, but I didn't want to slow down the process here and was not sure whether you preferred to have comments first or accompany them by change suggestions right away.
None of my comments are major concerns, therefore: approved!
|
|
||
| ## What is QA/QC? | ||
|
|
||
| Physiological quality assessment (QA) refers to the process of evaluating and identifying unreliable signals or signal segments, whereas physiological quality control (QC) encompasses the methods to handle/mitigate poor-quality signals. These procedures are crucial for ensuring the validity of experimental results, for instance if physiological data is used to model sources of noise in neuroimaging modalities or to characterize physiological states and processes in relation to brain function. Despite the consensus of many physiological data users within Physiopy’s community on the importance of QA/QC in physiological data acquisition and application, QA/QC methodologies for physiological data are rarely reported in detail in neuroimaging scientific literature. |
There was a problem hiding this comment.
I think QA typically abbreviates "quality assurance" not assessment. Not sure you changed this intentionally.
|
|
||
| ## When is QA/QC important? | ||
|
|
||
| Some amount of QA/QC is always recommended. The extent to which a researcher makes this a priority, and the thresholds they set to identify unreliable or poor-quality signals will vary with context. For example, in studies developing predictive models from large physiological datasets, a higher degree of signal noise may be tolerable, particularly if the primary aim is prediction and the model is expected to be robust to noisy inputs. On the other hand, QA/QC may be more important when poor signal quality, and associated artefacts, can introduce systematic bias or confounds. A common case is when the condition/state of interest is more prevalent in a population that also happens to exhibit more movement during data collection, and therefore an algorithm could learn to associate movement-related artifacts with the condition rather than the physiological features. In such cases, QA/QC can help reduce these confounds and lead to more accurate and interpretable findings. |
There was a problem hiding this comment.
"signal noise" can read a bit ambiguously, e.g., as "signal to noise" or maybe a placeholder meaning either signal or noise. I think what you mean here is something like "recording noise" or "acquisition noise" to separate it from "physiological noise" referred to in other parts of the documentation.
There was a problem hiding this comment.
Thank you for pointing that out! I think I would go with "acquisition noise" which maybe allows for a broader interpretation (i.e., anything during the acquisition that could add noise to the recordings) compared to "recording noise", which maybe implies more specifically something related to the recording devices/setup itself?
|
|
||
| ## Assessing signal quality in context | ||
|
|
||
| Rather than solely relying on automated signal quality indices (SQIs), it is highly advisable to visually examine the timeseries (Power et al., 2020) to better differentiate physiological events from artifacts and determine the appropriate intervention (i.e., interpolation, rejection of segments, etc.). Different types of artefacts independent of the specific modality can be encountered while QAing physiological signals: clipping, motion, drop or loss of signal. Clipping artefacts can be described as truncated segments (horizontal plateaus), and are related to signal saturation. Motion can introduce high frequency noise in the signal or other waveform irregularities. Finally, drop or loss of signal is characterized by a transient drop in the signal-to-noise ratio or a complete loss of signal. |
There was a problem hiding this comment.
With "(Power et al., 2020)", do you mean https://doi.org/10.1016/j.neuroimage.2019.116234 from the references.md? If so, is there a reason not to link to the referenced paper directly, or at least to the section in the references.md where the references are listed. This would make it easier for the reader to look up the paper/follow the argument.
|
|
||
| ## Assessing signal quality in context | ||
|
|
||
| Rather than solely relying on automated signal quality indices (SQIs), it is highly advisable to visually examine the timeseries (Power et al., 2020) to better differentiate physiological events from artifacts and determine the appropriate intervention (i.e., interpolation, rejection of segments, etc.). Different types of artefacts independent of the specific modality can be encountered while QAing physiological signals: clipping, motion, drop or loss of signal. Clipping artefacts can be described as truncated segments (horizontal plateaus), and are related to signal saturation. Motion can introduce high frequency noise in the signal or other waveform irregularities. Finally, drop or loss of signal is characterized by a transient drop in the signal-to-noise ratio or a complete loss of signal. |
There was a problem hiding this comment.
You use the spelling "artifacts" and then "artefacts" in the next sentence - can we agree on one way of spelling it?
|
|
||
| ## Assessing signal quality in context | ||
|
|
||
| Rather than solely relying on automated signal quality indices (SQIs), it is highly advisable to visually examine the timeseries (Power et al., 2020) to better differentiate physiological events from artifacts and determine the appropriate intervention (i.e., interpolation, rejection of segments, etc.). Different types of artefacts independent of the specific modality can be encountered while QAing physiological signals: clipping, motion, drop or loss of signal. Clipping artefacts can be described as truncated segments (horizontal plateaus), and are related to signal saturation. Motion can introduce high frequency noise in the signal or other waveform irregularities. Finally, drop or loss of signal is characterized by a transient drop in the signal-to-noise ratio or a complete loss of signal. |
There was a problem hiding this comment.
"clipping, motion, drop or loss of signal" - while I would agree that these are the main artifact categories, I wonder whether we could find better labels for some:
"motion" - that is the source, but not the (visual) characteristic of the artifact, which you explain later as "high-frequency noise or other irregularities"; I often hear labels such as "spiking" or "baseline shifts" for these characteristics, maybe one of those or even high-freq noise itself is a better name than motion
"drop or loss of signal" - at first, this sounded as if it was the same to me. If the signal drops, this can mean that it is lost completely (as in: "the call dropped"), and vice versa, a loss can be read as partial, not complete loss. Maybe "signal reduction" is better than "drop" or, if you formulate it from the other angle, "transient noise increase" might be an alternative.
For "loss of signal", you could use "flatlining" or "complete loss of signal" instead.
There was a problem hiding this comment.
I think it's important to find very precise and descriptive/self-explanatory definitions here, but if possible (maybe in a future version), an example figure showcasing these 4 categories would make it even more instructive.
|
|
||
| ## Recovering lost signal | ||
|
|
||
| This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function. |
There was a problem hiding this comment.
| This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function. | |
| This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery of the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal is used to compute metrics within a sliding window or in metrics subsequently convolved with an impulse response function. |
There was a problem hiding this comment.
Check whether I added the right words in the last sentence - the grammar was confusing for me before.
|
|
||
| This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function. | ||
|
|
||
| 1. Manual peak/trough correction |
There was a problem hiding this comment.
Should these be sub-subsections, i.e.,
1. Manual peak/trough correction
|
|
||
| ## Recovering lost signal | ||
|
|
||
| This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function. |
There was a problem hiding this comment.
Since this paragraph advocates for caution using correction methods, I was wondering whether it makes sense to suggest data provenance measures, like keeping a log/label file indicating which time points/segments were altered by which method.
There was a problem hiding this comment.
Totally! Maybe this can be made more explicit in the Reporting data quality section? I think it can also be more detailed in a future version of the documentation since we have also talked at some point of having a more exhaustive reporting section/page, that would not only encompasses QC/QA, but be along the lines of the COBIDAS guidelines.
|
|
||
| 6. Recovering signals from fMRI data | ||
|
|
||
| If the physiological recordings are unusable across an entire run, deep learning can potentially reconstruct or infer these signals directly from the fMRI data. However, the feasibility of this approach depends on the specific signal of interest and the downstream application. For example, Wang and colleagues (2025) have proposed a framework for reconstructing low-frequency respiratory volume and heart rate time series from fMRI data. Other studies have also shown the feasibility of such an approach on different populations and using different machine/deep learning techniques (Addeh et al., 2023; Bayrak et al., 2020; Bayrak et al., 2021; Salas et al., 2020). |
There was a problem hiding this comment.
I think this would also be a relevant paper to add here (related to the rapidtide/happy software):
S. Aslan, L. Hocke, N. Schwarz, and B. Frederick. Extraction of the cardiac waveform from simultaneous multislice fMRI data using slice
sorted averaging and a deep learning reconstruction filter. NeuroImage, 198:303–316, 05 2019. doi:10.1016/j.neuroimage.2019.05.049
|
|
||
| ## When is QA/QC important? | ||
|
|
||
| Some amount of QA/QC is always recommended. The extent to which a researcher makes this a priority, and the thresholds they set to identify unreliable or poor-quality signals will vary with context. For example, in studies developing predictive models from large physiological datasets, a higher degree of signal noise may be tolerable, particularly if the primary aim is prediction and the model is expected to be robust to noisy inputs. On the other hand, QA/QC may be more important when poor signal quality, and associated artefacts, can introduce systematic bias or confounds. A common case is when the condition/state of interest is more prevalent in a population that also happens to exhibit more movement during data collection, and therefore an algorithm could learn to associate movement-related artifacts with the condition rather than the physiological features. In such cases, QA/QC can help reduce these confounds and lead to more accurate and interpretable findings. |
There was a problem hiding this comment.
I wonder if we could be more clear about the movement issue. If I read this line as "movement-related artifacts" in the imaging data, then no amount of physio QA/QC will reduce those confounds, really. Perhaps "could learn to associate movement-related artifacts in the physiological signals with..."
|
|
||
| Other physiological events that look like irregularities may be observed, but can be associated with different physiological states (e.g., increased vigilance) or pathological conditions (e.g., atrial fibrillation–a type of heart arrhythmia). In such cases, irregularities in the signal are reflecting important physiological information, and should not be discarded. However, it is important to note that the identification of irregular physiological events can be used to inform subsequent analyses. For example, ectopic heartbeats should be excluded when calculating heart rate variability, while breathing irregularities have been found to correspond to characteristic changes in fMRI data (Task Force of the European Society of Cardiology and the North American Society of Pacing and Electrophysiology, 1996; Power et al., 2020). | ||
|
|
||
| Furthermore, the interpretation of physiological variability should also take into account the experimental paradigm and the expected physiological response to the task. For example, the timing of observed irregularities should be considered in relation to task events, such as button presses or other movements that may induce changes in the physiological recordings. Similarly, tasks or experimental manipulations may unintentionally or intentionally alter respiration (e.g., through voluntary breathing changes or the use of respiratory masks), which may in turn affect cardiac activity. Thus, variability that is temporally aligned with or physiologically plausible given the task should not necessarily be considered artifactual. Physiological signals can be entrained by task events, which should be considered when distinguishing artifacts from meaningful physiological responses. |
There was a problem hiding this comment.
Within the overall category of "Assessing signal quality in context", it feels remiss not to mention evaluating both individual peaks but also the overall waveform at much longer timescales. Some amplitude changes become more clearly "artifact" or more clearly "true physiological variation" when you scale out.
|
|
||
| Rather than manually inspecting the entire timeseries, automated flagging can be used to identify segments potentially containing poor-quality signals. Many physiological metrics are calculated from peaks and/or troughs (e.g., heart/respiratory rate, respiratory amplitude). Algorithms to automatically detect those features are often impacted by the signal quality, and can also make errors (e.g., misplaced peak, missed peak, extra peak), which can influence the accuracy of the metrics of interest. Segments to validate can be identified on the basis of the rate/inter-peak intervals: very low rate/very high inter-peak intervals can be indicative of a missed peak, whereas very high rate/very low inter-peak intervals can reflect an extra peak. However, relying only on automatically flagged segments can overlook misplaced peaks, as these may not appear as outliers in rate or inter-peak interval metrics, but still impact your metrics of interest. Despite this limitation, this approach serves as an efficient heuristic for initial screening, particularly when working with large datasets. | ||
|
|
||
| Concurrently recording multiple physiological modalities – such as cardiac and respiratory data – during fMRI acquisition enables simultaneous multi-channel evaluation. Given the physiological coupling between the corresponding systems, this cross-modal context can offer valuable insight into the underlying nature of physiological events, guiding QCing decisions (e. g., whether or not to reject some irregularities). For example, large-amplitude fluctuations in a respiratory waveform can be cross-referenced with cardiac inter-beat intervals (IBI), where co-occurring IBI variations would indicate a likely true respiratory event (e.g., respiratory sinus arrhythmia) rather than an isolated artifact. However, this relationship should be interpreted with caution in the MRI environment, as being inside the scanner may itself alter cardiac and respiratory activity, as well as their underlying physiological coupling. In particular, MRI-related anxiety has been associated with changes in the normally expected cardiorespiratory relationship, including the occurrence of negative respiratory sinus arrhythmia (nRSA), characterized by heart-rate deceleration during inspiration (Rassler et al., 2018). MRI-related anxiety has also been associated with the emergence of a prominent respiratory oscillation around 0.32 Hz (approximately 19 breaths/min), described as an “emotional breathing” oscillation (Pfurtscheller et al., 2024). These effects may also change over the course of the acquisition as participants habituate to the scanner environment and become more relaxed. |
There was a problem hiding this comment.
The emphasis is on 'cross-modal' assessment here, and I'm not quite sure the boundaries of that term. Can we also mention how combining respiratory belt with gas recordings can help us evaluate how end-tidal values should be interpreted/trusted
|
|
||
| Concurrently recording multiple physiological modalities – such as cardiac and respiratory data – during fMRI acquisition enables simultaneous multi-channel evaluation. Given the physiological coupling between the corresponding systems, this cross-modal context can offer valuable insight into the underlying nature of physiological events, guiding QCing decisions (e. g., whether or not to reject some irregularities). For example, large-amplitude fluctuations in a respiratory waveform can be cross-referenced with cardiac inter-beat intervals (IBI), where co-occurring IBI variations would indicate a likely true respiratory event (e.g., respiratory sinus arrhythmia) rather than an isolated artifact. However, this relationship should be interpreted with caution in the MRI environment, as being inside the scanner may itself alter cardiac and respiratory activity, as well as their underlying physiological coupling. In particular, MRI-related anxiety has been associated with changes in the normally expected cardiorespiratory relationship, including the occurrence of negative respiratory sinus arrhythmia (nRSA), characterized by heart-rate deceleration during inspiration (Rassler et al., 2018). MRI-related anxiety has also been associated with the emergence of a prominent respiratory oscillation around 0.32 Hz (approximately 19 breaths/min), described as an “emotional breathing” oscillation (Pfurtscheller et al., 2024). These effects may also change over the course of the acquisition as participants habituate to the scanner environment and become more relaxed. | ||
|
|
||
| ## Recovering lost signal |
There was a problem hiding this comment.
Does it make sense to have, as an option on this list, something about using a surrogate? As in, we are missing too much CO2 data, so we use RVT from the belt data to capture similar underlying physiology. This isn't EXACTLY "recovering lost signal" so I also understand if it's outside of scope.
|
|
||
| 6. Recovering signals from fMRI data | ||
|
|
||
| If the physiological recordings are unusable across an entire run, deep learning can potentially reconstruct or infer these signals directly from the fMRI data. However, the feasibility of this approach depends on the specific signal of interest and the downstream application. For example, Wang and colleagues (2025) have proposed a framework for reconstructing low-frequency respiratory volume and heart rate time series from fMRI data. Other studies have also shown the feasibility of such an approach on different populations and using different machine/deep learning techniques (Addeh et al., 2023; Bayrak et al., 2020; Bayrak et al., 2021; Salas et al., 2020). |
There was a problem hiding this comment.
Related to my comment about surrogates above, what about estimating a phys recording from another phys recording. Consider Becca's paper using the respiratory belt data to estimate the end-tidal CO2 data when the end-tidal CO2 data are low quality (doi: 10.1162/imag_a_00536) and/or a related earlier paper from Jean Chen's group (https://doi.org/10.3389/FNIMG.2023.1119539)... this concept is very similar to estimating the physiology recording from the fMRI data, so might be a natural partner in this section
Co-authored-by: Lars Kasper <Lars.Kasper@utoronto.ca>
The QA/QC community practices section is ready for a first review round 🎉
Context for reviewers the content of this document was sourced from community discussions on QA/QC during community best practices meetings. People contributing to these meetings (or the subsequent write up of the notes) have been tagged in this first review. A second review will go out to the wider community later. Any questions, don't hesitate to ask.
Closes #17
Proposed Changes
Change Type
bugfix(+0.0.1)minor(+0.1.0)major(+1.0.0)refactoring(no version update)test(no version update)infrastructure(no version update)documentation(no version update)otherChecklist before review