Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .readthedocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
version: 2

build:
os: "ubuntu-20.04"
os: "ubuntu-lts-latest"
tools:
python: "3.10"

Expand Down
1 change: 1 addition & 0 deletions docs/source/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,7 @@ Community Practices for Collecting Physiological Data
Setup <phys_set_up>
Data Collection <collect_phys>
Data Processing <process_phys>
Data Quality <qa_qc>
References <references>

Licence <https://github.com/physiopy/physiopy-community-guidelines?tab=CC-BY-SA-4.0-1-ov-file#readme>
Expand Down
63 changes: 63 additions & 0 deletions docs/source/qa_qc.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# 6. QA/QCing physiological data

In the Data Collection section of this documentation, we have mentioned different procedures to include during data acquisition to assess and optimize the quality of the recorded signals at scan time. In this section, we discuss post-acquisition methods to evaluate the quality of physiological data, how to mitigate certain artefacts, and the overall relevance of physiological data quality for application in neuroimaging contexts.

```{important}
The current version of this section is under active development, and will continue to be expanded. In future, more guidelines will be provided for specific physiological modalities and a specific section on ‘Tools and Resources for QA/QC’ will be added; currently this document is focused on principles and concepts only. If you are interested in helping to develop material for this section, you can submit an issue on our [Github repository](https://github.com/physiopy/physiopy-community-practices/issues), join the Physiopy Slack by emailing physiopy.community@gmail.com, and/or subscribe to [the newsletter](https://physiopy-community.beehiiv.com/) to get details on the next QA/QC community meeting, alongside other Physiopy updates.
```

## What is QA/QC?

Physiological quality assessment (QA) refers to the process of evaluating and identifying unreliable signals or signal segments, whereas physiological quality control (QC) encompasses the methods to handle/mitigate poor-quality signals. These procedures are crucial for ensuring the validity of experimental results, for instance if physiological data is used to model sources of noise in neuroimaging modalities or to characterize physiological states and processes in relation to brain function. Despite the consensus of many physiological data users within Physiopy’s community on the importance of QA/QC in physiological data acquisition and application, QA/QC methodologies for physiological data are rarely reported in detail in neuroimaging scientific literature.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think QA typically abbreviates "quality assurance" not assessment. Not sure you changed this intentionally.


## When is QA/QC important?

Some amount of QA/QC is always recommended. The extent to which a researcher makes this a priority, and the thresholds they set to identify unreliable or poor-quality signals will vary with context. For example, in studies developing predictive models from large physiological datasets, a higher degree of signal noise may be tolerable, particularly if the primary aim is prediction and the model is expected to be robust to noisy inputs. On the other hand, QA/QC may be more important when poor signal quality, and associated artefacts, can introduce systematic bias or confounds. A common case is when the condition/state of interest is more prevalent in a population that also happens to exhibit more movement during data collection, and therefore an algorithm could learn to associate movement-related artifacts with the condition rather than the physiological features. In such cases, QA/QC can help reduce these confounds and lead to more accurate and interpretable findings.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"signal noise" can read a bit ambiguously, e.g., as "signal to noise" or maybe a placeholder meaning either signal or noise. I think what you mean here is something like "recording noise" or "acquisition noise" to separate it from "physiological noise" referred to in other parts of the documentation.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for pointing that out! I think I would go with "acquisition noise" which maybe allows for a broader interpretation (i.e., anything during the acquisition that could add noise to the recordings) compared to "recording noise", which maybe implies more specifically something related to the recording devices/setup itself?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if we could be more clear about the movement issue. If I read this line as "movement-related artifacts" in the imaging data, then no amount of physio QA/QC will reduce those confounds, really. Perhaps "could learn to associate movement-related artifacts in the physiological signals with..."


## Assessing signal quality in context

Rather than solely relying on automated signal quality indices (SQIs), it is highly advisable to visually examine the timeseries (Power et al., 2020) to better differentiate physiological events from artifacts and determine the appropriate intervention (i.e., interpolation, rejection of segments, etc.). Different types of artefacts independent of the specific modality can be encountered while QAing physiological signals: clipping, motion, drop or loss of signal. Clipping artefacts can be described as truncated segments (horizontal plateaus), and are related to signal saturation. Motion can introduce high frequency noise in the signal or other waveform irregularities. Finally, drop or loss of signal is characterized by a transient drop in the signal-to-noise ratio or a complete loss of signal.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With "(Power et al., 2020)", do you mean https://doi.org/10.1016/j.neuroimage.2019.116234 from the references.md? If so, is there a reason not to link to the referenced paper directly, or at least to the section in the references.md where the references are listed. This would make it easier for the reader to look up the paper/follow the argument.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You use the spelling "artifacts" and then "artefacts" in the next sentence - can we agree on one way of spelling it?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"clipping, motion, drop or loss of signal" - while I would agree that these are the main artifact categories, I wonder whether we could find better labels for some:

"motion" - that is the source, but not the (visual) characteristic of the artifact, which you explain later as "high-frequency noise or other irregularities"; I often hear labels such as "spiking" or "baseline shifts" for these characteristics, maybe one of those or even high-freq noise itself is a better name than motion

"drop or loss of signal" - at first, this sounded as if it was the same to me. If the signal drops, this can mean that it is lost completely (as in: "the call dropped"), and vice versa, a loss can be read as partial, not complete loss. Maybe "signal reduction" is better than "drop" or, if you formulate it from the other angle, "transient noise increase" might be an alternative.

For "loss of signal", you could use "flatlining" or "complete loss of signal" instead.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's important to find very precise and descriptive/self-explanatory definitions here, but if possible (maybe in a future version), an example figure showcasing these 4 categories would make it even more instructive.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Rather than solely relying on automated signal quality indices (SQIs), it is highly advisable to visually examine the timeseries (Power et al., 2020) to better differentiate physiological events from artifacts and determine the appropriate intervention (i.e., interpolation, rejection of segments, etc.). Different types of artefacts independent of the specific modality can be encountered while QAing physiological signals: clipping, motion, drop or loss of signal. Clipping artefacts can be described as truncated segments (horizontal plateaus), and are related to signal saturation. Motion can introduce high frequency noise in the signal or other waveform irregularities. Finally, drop or loss of signal is characterized by a transient drop in the signal-to-noise ratio or a complete loss of signal.
Rather than solely relying on automated signal quality indices (SQIs), it is highly advisable to visually examine the timeseries (Power et al., 2020) to better differentiate physiological events from artifacts and determine the appropriate intervention (i.e., interpolation, rejection of segments, etc.). Different types of artifacts independent of the specific modality can be encountered while QAing physiological signals: clipping, motion, drop or loss of signal. Clipping artefacts can be described as truncated segments (horizontal plateaus), and are related to signal saturation. Motion can introduce high frequency noise in the signal or other waveform irregularities. Finally, drop or loss of signal is characterized by a transient drop in the signal-to-noise ratio or a complete loss of signal.


Other physiological events that look like irregularities may be observed, but can be associated with different physiological states (e.g., increased vigilance) or pathological conditions (e.g., atrial fibrillation–a type of heart arrhythmia). In such cases, irregularities in the signal are reflecting important physiological information, and should not be discarded. However, it is important to note that the identification of irregular physiological events can be used to inform subsequent analyses. For example, ectopic heartbeats should be excluded when calculating heart rate variability, while breathing irregularities have been found to correspond to characteristic changes in fMRI data (Task Force of the European Society of Cardiology and the North American Society of Pacing and Electrophysiology, 1996; Power et al., 2020).

Furthermore, the interpretation of physiological variability should also take into account the experimental paradigm and the expected physiological response to the task. For example, the timing of observed irregularities should be considered in relation to task events, such as button presses or other movements that may induce changes in the physiological recordings. Similarly, tasks or experimental manipulations may unintentionally or intentionally alter respiration (e.g., through voluntary breathing changes or the use of respiratory masks), which may in turn affect cardiac activity. Thus, variability that is temporally aligned with or physiologically plausible given the task should not necessarily be considered artifactual. Physiological signals can be entrained by task events, which should be considered when distinguishing artifacts from meaningful physiological responses.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"alter respiration [...], which may in turn affect cardiac activity" - do you mean respiratory sinus arrythmia - if so, it can be mentioned in a bracket, together with other phenomena that you were thinking off (this was the first one that came to my mind).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Within the overall category of "Assessing signal quality in context", it feels remiss not to mention evaluating both individual peaks but also the overall waveform at much longer timescales. Some amplitude changes become more clearly "artifact" or more clearly "true physiological variation" when you scale out.


Rather than manually inspecting the entire timeseries, automated flagging can be used to identify segments potentially containing poor-quality signals. Many physiological metrics are calculated from peaks and/or troughs (e.g., heart/respiratory rate, respiratory amplitude). Algorithms to automatically detect those features are often impacted by the signal quality, and can also make errors (e.g., misplaced peak, missed peak, extra peak), which can influence the accuracy of the metrics of interest. Segments to validate can be identified on the basis of the rate/inter-peak intervals: very low rate/very high inter-peak intervals can be indicative of a missed peak, whereas very high rate/very low inter-peak intervals can reflect an extra peak. However, relying only on automatically flagged segments can overlook misplaced peaks, as these may not appear as outliers in rate or inter-peak interval metrics, but still impact your metrics of interest. Despite this limitation, this approach serves as an efficient heuristic for initial screening, particularly when working with large datasets.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"very low rate/very high inter-peak intervals" - maybe add a qualification, e.g., "compared to typical physiological ranges / the statistical distribution/histogram of the individual"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After reading the whole section, I wonder whether the juxtaposition of manual scrutiny above and automatic assessment here ("Rather than manually inspecting the entire timeseries, automated flagging") is the main one? I think both the 4 artifact categories mentioned in the beginning can be assessed manually/automatically, while also peak-to-peak interval time courses here, once computed, can be inspected both manually or thresholded automatically. The differentiation is rather: assessing the raw time series (above) vs. assessing a derived measure (like IBI) here.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In general, it might make sense to have sub-subsections for the "Assessing signal quality in context" subsection, e.g., "Assessing the raw physiological trace", "Artifacts vs physiological events", "Correlations with Behavior", "Integrating multi-modal/channel" information"


Concurrently recording multiple physiological modalities – such as cardiac and respiratory data – during fMRI acquisition enables simultaneous multi-channel evaluation. Given the physiological coupling between the corresponding systems, this cross-modal context can offer valuable insight into the underlying nature of physiological events, guiding QCing decisions (e. g., whether or not to reject some irregularities). For example, large-amplitude fluctuations in a respiratory waveform can be cross-referenced with cardiac inter-beat intervals (IBI), where co-occurring IBI variations would indicate a likely true respiratory event (e.g., respiratory sinus arrhythmia) rather than an isolated artifact. However, this relationship should be interpreted with caution in the MRI environment, as being inside the scanner may itself alter cardiac and respiratory activity, as well as their underlying physiological coupling. In particular, MRI-related anxiety has been associated with changes in the normally expected cardiorespiratory relationship, including the occurrence of negative respiratory sinus arrhythmia (nRSA), characterized by heart-rate deceleration during inspiration (Rassler et al., 2018). MRI-related anxiety has also been associated with the emergence of a prominent respiratory oscillation around 0.32 Hz (approximately 19 breaths/min), described as an “emotional breathing” oscillation (Pfurtscheller et al., 2024). These effects may also change over the course of the acquisition as participants habituate to the scanner environment and become more relaxed.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"e. g." has an extra space here, should be "e.g.", I guess.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Concurrently recording multiple physiological modalities – such as cardiac and respiratory data – during fMRI acquisition enables simultaneous multi-channel evaluation. Given the physiological coupling between the corresponding systems, this cross-modal context can offer valuable insight into the underlying nature of physiological events, guiding QCing decisions (e. g., whether or not to reject some irregularities). For example, large-amplitude fluctuations in a respiratory waveform can be cross-referenced with cardiac inter-beat intervals (IBI), where co-occurring IBI variations would indicate a likely true respiratory event (e.g., respiratory sinus arrhythmia) rather than an isolated artifact. However, this relationship should be interpreted with caution in the MRI environment, as being inside the scanner may itself alter cardiac and respiratory activity, as well as their underlying physiological coupling. In particular, MRI-related anxiety has been associated with changes in the normally expected cardiorespiratory relationship, including the occurrence of negative respiratory sinus arrhythmia (nRSA), characterized by heart-rate deceleration during inspiration (Rassler et al., 2018). MRI-related anxiety has also been associated with the emergence of a prominent respiratory oscillation around 0.32 Hz (approximately 19 breaths/min), described as an “emotional breathing” oscillation (Pfurtscheller et al., 2024). These effects may also change over the course of the acquisition as participants habituate to the scanner environment and become more relaxed.
Concurrently recording multiple physiological modalities – such as cardiac and respiratory data – during fMRI acquisition enables simultaneous multi-channel evaluation. Given the physiological coupling between the corresponding systems, this cross-modal context can offer valuable insight into the underlying nature of physiological events, guiding QCing decisions (e.g., whether or not to reject some irregularities). For example, large-amplitude fluctuations in a respiratory waveform can be cross-referenced with cardiac inter-beat intervals (IBI), where co-occurring IBI variations would indicate a likely true respiratory event (e.g., respiratory sinus arrhythmia) rather than an isolated artifact. However, this relationship should be interpreted with caution in the MRI environment, as being inside the scanner may itself alter cardiac and respiratory activity, as well as their underlying physiological coupling. In particular, MRI-related anxiety has been associated with changes in the normally expected cardiorespiratory relationship, including the occurrence of negative respiratory sinus arrhythmia (nRSA), characterized by heart-rate deceleration during inspiration (Rassler et al., 2018). MRI-related anxiety has also been associated with the emergence of a prominent respiratory oscillation around 0.32 Hz (approximately 19 breaths/min), described as an “emotional breathing” oscillation (Pfurtscheller et al., 2024). These effects may also change over the course of the acquisition as participants habituate to the scanner environment and become more relaxed.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The emphasis is on 'cross-modal' assessment here, and I'm not quite sure the boundaries of that term. Can we also mention how combining respiratory belt with gas recordings can help us evaluate how end-tidal values should be interpreted/trusted


## Recovering lost signal

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it make sense to have, as an option on this list, something about using a surrogate? As in, we are missing too much CO2 data, so we use RVT from the belt data to capture similar underlying physiology. This isn't EXACTLY "recovering lost signal" so I also understand if it's outside of scope.


This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function.
This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would later be computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sounded more natural to me to move the "later" before the "be computed"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function.
This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artifacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

artifact/artefact

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery if the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal used to compute metrics within a sliding window or metrics subsequently convolved with an impulse response function.
This section gives a brief overview of the different methods that exist to correct poor-quality physiological data segments. However, regardless of the chosen strategy, it is important to verify its effectiveness since it will influence the features or regressors that would be later computed. For example, if you are interested in heart rate variability, incorrect beat placement outweighs missing values (see [HeartPy’s simulation on the matter](https://python-heart-rate-analysis-toolkit.readthedocs.io/en/latest/heartrateanalysis.html#on-the-accuracy-of-peak-position); van Gent, Farah, Nes, & Arem, 2018a; van Gent, Farah, Nes, & van Arem, 2018b). In such a case, masking the segment might be a better option than interpolating a missed/incorrectly detected beat or manually correcting the peak if in doubt. Alternative metrics that are less affected by outliers could also be considered (e.g., instantaneous heart rate). Furthermore, it is important to consider that the influence of artefacts or deviations introduced during the recovery of the physiological signal may be attenuated in magnitude or temporally extended at later stages of the processing pipeline, particularly when the recovered signal is used to compute metrics within a sliding window or in metrics subsequently convolved with an impulse response function.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Check whether I added the right words in the last sentence - the grammar was confusing for me before.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this paragraph advocates for caution using correction methods, I was wondering whether it makes sense to suggest data provenance measures, like keeping a log/label file indicating which time points/segments were altered by which method.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Totally! Maybe this can be made more explicit in the Reporting data quality section? I think it can also be more detailed in a future version of the documentation since we have also talked at some point of having a more exhaustive reporting section/page, that would not only encompasses QC/QA, but be along the lines of the COBIDAS guidelines.


1. Manual peak/trough correction

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should these be sub-subsections, i.e.,

1. Manual peak/trough correction


As mentioned previously, automatic feature detection algorithms are susceptible to errors. Although it is more time-consuming and does not guarantee complete accuracy, manual feature correction leverages the data curator’s expertise and contextual understanding of the research goals to resolve complex edge cases.

2. Interpolation

Interpolation estimates missing values based on the surrounding data points and the chosen algorithm (e.g., linear, quadratic, cubic, nearest-neighbor). Consequently, interpolation must be applied cautiously when missing peaks occur during periods of high signal variance (irregular rhythms that could be biologically plausible) or rapid fluctuation. This method is also less reliable when applied to long segments. However, defining a ‘long’ segment remains context-dependent, as it depends on the specific metric of interest and whether it targets slow or fast dynamics (i.e., longer segments are more tolerable for lower frequency signals).

3. Imputation

Data imputation replaces missing or unreliable values using estimated or plausible values. These techniques can be useful to reconstruct missing signals when working with quasiperiodic signals. More and more algorithms based on deep learning have been developed in the past years, showing promising results (Xu et al., 2022). Nevertheless, much like interpolation, caution is necessary as imputed data can be speculative, particularly when applied over long gaps.

4. Filtering

The use of filtering strategies depends on the metrics of interest you may want to compute on your signals and the artifact type. For example, when calculating the derivative of the respiratory trace, filtering can be applied a priori if cardiac contamination is present. However, filter selection requires caution, as improper parameters can introduce signal distortion.

5. Masking

The process of masking consists of replacing values in segments with zeros or NaN values. As discussed previously, masking might be preferred over interpolation or imputation depending on the metrics of interest. When working with physiological regressors, slightly padding or extending the masks helps account for temporal delays.

6. Recovering signals from fMRI data

If the physiological recordings are unusable across an entire run, deep learning can potentially reconstruct or infer these signals directly from the fMRI data. However, the feasibility of this approach depends on the specific signal of interest and the downstream application. For example, Wang and colleagues (2025) have proposed a framework for reconstructing low-frequency respiratory volume and heart rate time series from fMRI data. Other studies have also shown the feasibility of such an approach on different populations and using different machine/deep learning techniques (Addeh et al., 2023; Bayrak et al., 2020; Bayrak et al., 2021; Salas et al., 2020).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this would also be a relevant paper to add here (related to the rapidtide/happy software):

S. Aslan, L. Hocke, N. Schwarz, and B. Frederick. Extraction of the cardiac waveform from simultaneous multislice fMRI data using slice
sorted averaging and a deep learning reconstruction filter. NeuroImage, 198:303–316, 05 2019. doi:10.1016/j.neuroimage.2019.05.049

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Related to my comment about surrogates above, what about estimating a phys recording from another phys recording. Consider Becca's paper using the respiratory belt data to estimate the end-tidal CO2 data when the end-tidal CO2 data are low quality (doi: 10.1162/imag_a_00536) and/or a related earlier paper from Jean Chen's group (https://doi.org/10.3389/FNIMG.2023.1119539)... this concept is very similar to estimating the physiology recording from the fMRI data, so might be a natural partner in this section


## Reporting data quality

Given that a QA/QC pipeline could vary depending on the metrics of interest to extract from the signal and their specific use, we encourage transparency in reporting the criteria used to assess the quality of physiological signals. We also encourage researchers to share raw timeseries and not only derivatives when possible, and to provide clearly documented code for data preprocessing (including QA/QCing) and data analysis. Since different quality control pipelines can be applied on the same dataset, researchers should provide clear documentation regarding which data were used, and which were rejected. If this is systematically reported, it would be easier to evaluate the impact of quality control variability on different analyses, which could help provide more standardized QA/QC guidelines for physiological data.

## Concluding remarks

Even though this documentation aims to provide guidelines on QA/QC practices for physiological data, remaining gaps must be acknowledged. Given the current lack of transparency in the application of QA/QC for physiological data in neuroscientific literature, there is a clear need to better document these procedures across different use cases, populations and specific physiological modalities. Reporting standards à la COBIDAS would be a valuable addition to the field. Initiatives toward this goal have started to be undertaken by the Society for Psychophysiological Research for physiological signals (specifically for heart rate and heart rate variability reporting; Quigley et al., 2024) However, recording physiological signals concurrently with fMRI poses its own set of distinct challenges (e.g., different types of artifacts to consider during the preprocessing), and often serves specific applications, such as physiological noise modeling for fMRI data denoising. Thus, efforts to characterize the effects of concurrent neuroimaging on the acquisition and quality of physiological data are crucial (Schumann et al., 2021).
Loading
Loading