Data-driven comparison of EEG preprocessing pipelines in EEGLAB.
PipeCompare is an EEGLAB plugin that chooses how to preprocess an EEG recording, based on the recording itself. It compares complete preprocessing pipelines (filters, reference, bad channels, ICA with ICLabel, epoch rejection) and recommends the one that measures your ERP or band power most precisely, without distorting it.
- What PipeCompare does
- What happens to your data
- Features
- Where PipeCompare fits in an analysis
- Requirements
- Installation
- Quick start
- Scripting interface
- How a pipeline is chosen
- Outputs
- Limitations
- Testing
- Documentation
- Citation
- License
Before an ERP amplitude or a band power can be measured, the raw recording has to be cleaned: it is filtered, re-referenced, bad channels are found and repaired, eye and muscle artifacts are removed with ICA, the data are cut into epochs and noisy epochs are rejected. Every one of these steps has settings, such as the high-pass cutoff, the ICLabel threshold for removing a component or the amplitude limit for rejecting an epoch. One complete set of steps and settings is a pipeline. Published studies use many different pipelines, and the choice changes how much noise is left in the final measurement. In practice the settings are usually copied from an earlier paper or kept out of habit, without checking how well they work on the data at hand.
PipeCompare makes this choice from the data, in these steps:
- It reads the dataset open in EEGLAB, including what has already been done to it (from
EEG.history), so earlier processing is neither repeated nor ignored. - You say what will be measured: an ERP component (mean amplitude, peak amplitude or peak latency in a time window, at the electrodes you choose, time-locked to the events you choose), or the power in a frequency band. Presets cover the ERP CORE components and the classic frequency bands.
- It builds every pipeline from the steps and settings left open. The Standard recipe, for example, compares up to 12 combinations of high-pass and low-pass filters, and in each of them finds the best ICLabel threshold and epoch-rejection limit from the data; you can also decide the steps and values yourself.
- It runs each pipeline with EEGLAB's own functions, on your data.
- It removes pipelines that harm the data. A pipeline is excluded when it keeps too few trials, interpolates too many channels, or changes the brain signal. To check the last point, a known artificial signal is added to a copy of the data and carried through the same pipeline; if it comes out smaller, shifted or reshaped, the pipeline is excluded.
- It ranks the rest by precision. The measure is the standardized measurement error (SME; Luck et al., 2021): how much your averaged value would vary if the experiment were repeated. It is corrected for how much each pipeline shrinks the signal, so a pipeline cannot look good just by making everything smaller. A bootstrap test then finds the pipelines that cannot be told apart from the best.
- It recommends one pipeline, among those as good as the best the one that keeps the most trials and changes the signal least, and explains why. You can adopt it as a new EEGLAB dataset, already epoched and cleaned and ready to average, or save it as a MATLAB script to process your other recordings the same way.
The comparison never uses an experimental effect (condition differences, p-values), so selecting a pipeline does not bias the statistical test you run afterwards.
PipeCompare is useful when you start working with a new dataset, paradigm or recording setup and want preprocessing settings with a documented, data-based reason, or want to check whether your usual settings suit these data. It works on one recording at a time, and covers ERP and band-power measures; time-frequency, connectivity and source analysis are not covered (see Limitations).
With the Standard recipe, every pipeline runs the steps below in this order. Steps marked compared are tried with several settings; the others are done once, in the same way for every pipeline, so that the comparison is about the settings that matter.
- Channels removed earlier are put back (only with the average reference). If you deleted EEG channels before PipeCompare, they are interpolated back first, so that the average is taken over the whole montage.
- Bad channels are found and repaired. A channel is marked bad when its signal is much spikier (kurtosis) or much noisier (joint probability, e.g. a poorly connected electrode) than the other channels, more than 5 standard deviations away. A flat channel (no signal, as when an electrode is disconnected: its spread is zero or under 1 % of the median spread of the channels tested) is always bad; it is set aside before the two measures, which cannot be computed on a constant and would otherwise be thrown off for all the other channels too. The test looks at a copy of the data high-passed at 1 Hz, so slow drifts are not mistaken for bad channels. Bad channels are then interpolated from their neighbours (spherical interpolation), so the dataset keeps all its channels. EOG, ECG and EMG channels (by channel type, or by name such as VEOG or ECG1) are never tested. A pipeline that interpolates more than 20 % of the channels is excluded.
- Re-reference (if you chose the average reference). It comes after the bad channels are repaired, so a bad channel cannot spread its noise into every other channel; non-EEG channels are left out of the average. Data that were already average-referenced are averaged again, which removes the repaired channels' share of the earlier average.
- ICA is computed once (extended Infomax,
pop_runica) on a copy high-passed at 1 Hz, which gives a cleaner decomposition. The copy also leaves out the stretches of recording that are far noisier than the rest (movement, electrode bursts; see ICA on clean data), so that they do not claim components of their own. The resulting decomposition is applied to all of the data. All pipelines share this one decomposition, which saves the largest part of the computing time. - High-pass filter, compared: 0.1, 0.3, 0.5 and 1 Hz.
- Low-pass filter, compared: 20, 30 and 40 Hz.
- Artifact components are removed with ICLabel, threshold chosen from the data: a component is removed when ICLabel gives it a probability above a threshold of being eye, muscle, heart, line noise or channel noise, and the threshold is found in each pipeline from its own data together with the rejection limit of step 9: every threshold from 0.5 up is tried, and the highest one that measures as precisely as the best one is used (see ICLabel threshold chosen from the data). For peak measures, or when epochs are repaired, the thresholds 0.7, 0.8 and 0.9 are compared instead.
- Epochs are cut around the events you chose, and the baseline is subtracted.
- Noisy epochs are rejected, limit chosen from the data: an epoch is dropped when any EEG channel exceeds ±x µV (non-EEG channels are ignored), and x is found in each pipeline from its own data: every limit is tried, not a few, and the highest limit that measures as precisely as the best one is used (see Rejection limit chosen from the data). For peak measures (peak amplitude or latency, advanced panel) the limits 75, 100 and 150 µV are compared instead.
Two more options of the rejection can be ticked; they are not part of Standard. Peak-to-peak limits (ERP CORE's method): instead of testing whether any value is beyond ±x µV, the rejection takes the highest minus the lowest value within a moving 200 ms window; the limit is again chosen from the data (see Peak-to-peak rejection). Repairing epochs instead of rejecting them: an epoch in which only 1 to 3 channels exceed the rejection limit is kept, and those channels are interpolated from the other channels within that epoch only; an epoch with more channels over the limit is still rejected. With repair, the limit is not chosen from the data (that choice assumes the epochs over the limit are removed): the fixed limits 75, 100 and 150 µV (100, 150 and 200 µV peak-to-peak) and the ICLabel thresholds 0.7, 0.8 and 0.9 are compared, which makes nine times as many pipelines. See Repairing epochs with a few bad channels.
That makes 4 × 3 = 12 pipelines for an ERP, each with its own best ICLabel threshold and rejection limit. The Filters only recipe compares only steps 5 and 6. For band power, the data are cut into 2 s segments instead of epochs, and only the filter edges nearest the band, outside it, are used (a filter outside a band does not change its power), which leaves 1 pipeline (its ICLabel threshold and rejection limit still chosen from the data). In the advanced panel or a script you can change every list above, change the order, allow a step to be skipped, and add other steps: line-noise removal (the 50 or 60 Hz mains frequency is detected from the recording), ASR (artifact subspace reconstruction), resampling, rejection by joint probability or kurtosis, a reference to chosen channels, or any operation from EEGLAB's menus, plugins included.
Along the way PipeCompare also takes care of the following, so you do not have to:
-
It names flat channels (no signal, e.g. a disconnected electrode) before the run, in the dialog's first line and the panel's data warnings, and the bad-channel step always marks them bad. Without that step, remove or interpolate them yourself before an average reference.
-
It checks that the channel labels fit the electrode positions before the run, and names channels whose signal does not look like their neighbours' (see Montage check). This is a warning only.
-
It tells you when ICA did nothing useful: when ICLabel recognised almost none of the components, or no pipeline removed any component, and when the recording was too short for ICA (see When ICA does nothing).
-
It tells you when an electrode you measure was rebuilt from its neighbours by the recommended pipeline (as a bad channel). A rebuilt electrode has less noise than a real one, so that pipeline's noise looks smaller than it is (see Reading the result).
-
It tells you when one channel caused most rejected epochs, which usually means it is a bad channel the detection missed (see Reading the result).
-
It respects what was done before. Processing already applied (read from
EEG.history) is listed and not repeated. A filter the data already have is kept as one of the choices (e.g. low-pass as in the data (30 Hz)) instead of forcing a stricter one, and you are warned when that earlier filter alone already distorts the signal you measure. When ICA was already run and components removed, ICA is not compared again. -
It leaves out what the data cannot support, and says why. Without channel locations, interpolation and ICLabel are left out; without ICLabel installed, ICA is left out; on data that are already epoched, filters are not compared (they must run before epoching) and the data keep their own epochs as long as they hold the measurement window. A condition with too few events (fewer than 10) is flagged before the run.
-
It checks units. Data stored in volts are recognised from their amplitude and compared in µV.
-
It measures lateralized components correctly. N2pc and LRP are scored as ERP CORE does, contralateral minus ipsilateral, from the event types you give for each side.
-
It handles long runs. A progress window shows the step running and the time so far and left; Stop keeps the pipelines already finished. A search can be saved to a checkpoint folder and resumed, and can run in parallel with the Parallel Computing Toolbox.
-
It keeps everything reproducible. The adopted dataset's
EEG.historycontains every EEGLAB command that produced it, and Save script… writes the same steps as a MATLAB function that makes the data-driven decisions (bad channels, components, rejected epochs) anew on each recording you run it on.
- Works on the dataset you have open. Reads the current EEGLAB state (continuous or epoched
data, channel locations, ICA decomposition, reference, sampling rate) and parses
EEG.history, so earlier processing is taken into account. - ERP and band-power measures. Mean amplitude, peak amplitude and peak latency for event-related data. Log band power for continuous recordings such as resting state.
- Presets for common analyses. ERP CORE parameters (Kappenman et al., 2021) for N170, MMN, N2pc, N400, P3, LRP and ERN, and standard delta, theta, alpha and beta bands.
- User-defined search space. Choose the steps and their order, fix some values and let others vary, allow steps to be skipped. Any EEGLAB menu operation that takes and returns the dataset, plugins included, can be added as a step and configured in its own EEGLAB dialog.
- Exhaustive search. Every admissible pipeline is run, with shared leading steps computed only once. The search is never sampled or truncated.
- Signal-preservation check. A known signal is carried through every pipeline with the same data-driven decisions as the real data (bad channels, ICA, rejected epochs, removed components, ASR reconstructions); steps whose decisions cannot be replayed are re-run on it and flagged (see Limitations). Pipelines that distort it are excluded.
- Exact best rejection limit and ICLabel threshold. Within each pipeline, the epoch-rejection limit and the ICLabel threshold are chosen from the data together, among all values that give different results, not from a short list.
- Statistically controlled ranking. A paired bootstrap with simultaneous intervals identifies the pipelines that cannot be distinguished from the best, controlling the error rate across all candidates.
- Reproducible output. Each candidate's full list of EEGLAB commands, an exportable script, and
adoption as a new dataset whose
EEG.historyreproduces it. - Practical for long searches. A progress window that can stop a search and keep the finished pipelines, checkpointing and resume, and optional parallel execution.
PipeCompare replaces the preprocessing part of a single-recording EEGLAB analysis:
- Before PipeCompare (done by you): import the recording, add channel locations (Edit > Channel locations), and, if needed, remove channels that are not EEG or that you know are broken.
- PipeCompare: filtering, reference, bad-channel detection and interpolation, ICA with ICLabel, epoching, baseline correction and epoch rejection. These are the steps it compares, so they should not be applied beforehand.
- After PipeCompare: average the epochs into ERPs (or compute spectra for band power) and measure. The adopted dataset is already filtered, re-referenced and cleaned, so these steps are not repeated.
Processing that was already applied to the data is read from EEG.history and taken into account:
the dialog lists it, does not compare it again, and checks whether it already distorts the
measured signal. The cleanest comparison starts from the raw continuous recording, because then
every step is part of the comparison.
PipeCompare compares pipelines on one recording at a time. To process a whole study the same way, choose the pipeline on a representative recording and run the saved script (Save script…) on every recording; the script makes the data-driven decisions (bad channels, components, rejected epochs) from each recording's own data.
| Component | Requirement |
|---|---|
| MATLAB | R2021b or later (see the tested versions below). GNU Octave is not supported: the interface uses uifigure. |
| EEGLAB | 2022.0 or later. The PipeCompare menu and Add EEGLAB menu step… need EEGLAB's main window; scripts also run after eeglab nogui |
| EEGLAB plugins | firfilt (bundled with EEGLAB); ICLabel for IC removal; clean_rawdata for ASR |
| Optional | Parallel Computing Toolbox, for parallel execution; Signal Processing Toolbox, for ASR at sampling rates other than 100, 128, 200, 256, 300, 500 and 512 Hz and for faster low-frequency high-pass filtering |
MATLAB R2026a with EEGLAB 2026.0.0 is the combination run by hand
(docs/COVERAGE.md). Continuous integration runs the full automated test suite
on every MATLAB release from R2021b to R2026a and the latest release with EEGLAB 2026.0.0, and on
every EEGLAB release from 2022.0 to 2025.1.0 with both MATLAB R2022a and the latest MATLAB. It also
runs it on Windows (the latest MATLAB with EEGLAB 2026.0.0, and R2022a with EEGLAB 2022.0) and on
macOS (the latest MATLAB with EEGLAB 2026.0.0 and 2022.0; on GitHub's Mac test machines the windows
of older MATLAB releases close by themselves, so those are not tested there). On
R2022b and R2023a, MATLAB's own message box (uialert) never returns on the test machines'
virtual display, even in an empty window, so there the tests use a stand-in that prints the
message instead of drawing the box.
Two known problems of older EEGLAB releases, which PipeCompare works around or reports:
- EEGLAB 2024.0 and older, on recent MATLAB releases, stop with "Colon operands must be real scalars" when cutting time ranges out of continuous data (EEGLAB's own Edit > Select data). PipeCompare makes these cuts itself when that happens, with the same result.
- EEGLAB 2024.0 comes with ICLabel 1.5, which classifies components wrongly (its scalp maps are rotated, so eye components are mostly missed). When IC removal runs with that ICLabel, the Command Window shows a warning; update ICLabel to 1.6 or later in File > Manage EEGLAB extensions before comparing pipelines with IC removal.
PipeCompare reads the current dataset from EEGLAB's base-workspace variables EEG, ALLEEG
and CURRENTSET, and stores adopted pipelines there.
PipeCompare is listed in EEGLAB's extension manager (list of EEGLAB extensions), so the simplest way to install it is from inside EEGLAB. The other two ways are for computers without internet access in MATLAB and for working with the source code.
Keep only one copy of PipeCompare in eeglab/plugins/. EEGLAB loads every plugin folder it finds,
so two copies (for example one from the extension manager and one unzipped by hand) both add
their menus and their functions shadow each other.
From EEGLAB's extension manager (recommended)
- Start EEGLAB and choose File > Manage EEGLAB extensions.
- Find PipeCompare in the list (typing its name in the search field narrows the list), tick it and press Install/Update.
- The menu Tools > PipeCompare appears; if it does not, restart EEGLAB.
EEGLAB downloads the release zip and places it in eeglab/plugins/PipeCompare<version>/. When a
newer version is listed, the same window offers it, and installing it replaces the old folder.
From a GitHub release
- Download
PipeCompare<version>.zipfrom the Releases page. - Unzip it into
eeglab/plugins/. The folder must be namedPipeCompareorPipeCompare<version>, because EEGLAB takes the plugin's name and version from the folder name. - Start or restart EEGLAB. The menu Tools > PipeCompare appears.
From source
git clone https://github.com/abwoo/PipeCompare.gitEither copy the cloned folder into eeglab/plugins/ as above, or keep it elsewhere and register
it with the running EEGLAB:
eeglab % start EEGLAB first
pipecompare_setup % run from the PipeCompare folderpipecompare_setup adds the menu to the running EEGLAB session only. Run it again after each
eeglab call, because EEGLAB rebuilds its menus.
-
Load a dataset in EEGLAB.
-
Open Tools > PipeCompare > Compare pipelines…
-
Choose what to measure: an ERP component (with the event types it is time-locked to) or a frequency band, or your own time window or band with the electrodes you pick. The list only offers what the data support. The electrodes and time window of each ERP component are not taken from your data but from the published ERP CORE conventions (Kappenman et al., 2021), for example P3 at Pz, 300–600 ms. You can change a component's electrodes with Electrodes… next to its time window (for example Pz, CPz and POz for the P3); the window stays the convention's. N2pc and LRP keep their electrode pair, since they are scored as the difference between the two sides. When several electrodes are chosen, PipeCompare averages them first and scores that average waveform; for band power, the power of each chosen electrode is computed and then averaged. The simple mode scores one measure per run; to compare pipelines on several components at once, each with its own electrodes, use the advanced panel (Add component…).
Then tick the steps you want. They are listed in the order they run, and that order is fixed, so you only decide which steps are done, never in which order:
Order Step What is compared 1 Bad channels: detect and interpolate nothing (one fixed setting, see below) 2 Reference: as recorded or average nothing (your choice, the same in every pipeline) 3 ICA: remove artifact components the ICLabel probability above which a component is removed, chosen from the data in each pipeline (every threshold from 0.5 is tried; 0.7, 0.8, 0.9 compared with 7b) 4 High-pass filter the cutoff (0.1, 0.3, 0.5, 1 Hz) 5 Low-pass filter the cutoff (20, 30, 40 Hz) 6 Epochs and baseline (segments for band power) always done 7 Reject epochs over an amplitude limit the limit, chosen from the data in each pipeline (every limit is tried) 7a Measure the limit peak-to-peak in 200 ms windows (optional) the same, on peak-to-peak values 7b Instead, repair epochs with up to 3 channels over the limit (optional) the limit: 75, 100, 150 µV (100, 150, 200 µV peak-to-peak) Two buttons tick a usual set at once: Standard, preselected, ticks steps 1 to 7 (every step except the options 7a and 7b); Filters only ticks the two filters. Lines 7a and 7b are options of the rejection and can only be ticked together with step 7 (see Peak-to-peak rejection and Repairing epochs with a few bad channels). Ticking 7a does not change the number of pipelines; ticking 7b multiplies it by nine (three fixed limits and, with ICA, three ICLabel thresholds instead of both chosen from the data). An unticked step is not done at all: for example, without the high-pass the data keep whatever high-pass they already had. Steps that these data cannot take are greyed out, with the reason next to them: on epoched data the filters (they must run before epoching), without channel locations bad-channel interpolation and ICA (ICLabel needs the locations), without the ICLabel plugin ICA, and ICA when the history shows that ICA components were already removed. When the average reference is chosen (or the data are already average-referenced), bad-channel detection is always ticked, because a bad channel in the average would spread into every channel. For band power, when ICA or epoch rejection is compared, each filter uses one cutoff, the one nearest the band outside it (outside the band a filter does not change its power), so the standard set gives 1 pipeline instead of 12; tick only the filters to compare their cutoffs.
Bad channels are detected once and ICA is fitted once, before the filters, so every filter setting shares one decomposition. A channel is bad when its kurtosis (spiky) or joint probability (noisy, e.g. poor contact) is more than 5 SD from the other channels', or when it is flat (no signal). Both steps look at a 1 Hz high-passed copy, so slow drifts do not mislead them; the data themselves are filtered only by the settings being compared. EOG, ECG and EMG channels (by type or by name, such as VEOG) are not tested as bad channels and do not count in epoch rejection. The dialog shows how many pipelines will run before you start. Steps the data cannot support are left out, with the reason; for example, interpolation and ICLabel require channel locations, and for band power no filter edge inside the band is compared. Epoched data keep their own epochs when these hold the measurement window. N2pc and LRP are scored as ERP CORE measures them, contralateral minus ipsilateral: choose the event types of each side (target on the left and on the right; left-hand and right-hand responses). To compare ASR (artifact subspace reconstruction) or other steps, use the advanced panel or a script.
Start from the raw continuous data. If your analysis uses the average reference, choose it in the Reference line of the step list rather than re-referencing beforehand: it is then applied in every pipeline after the bad channels are interpolated and before ICA. (Data already average-referenced are averaged again after the interpolation.) Steps already applied to the data are read from
EEG.historyand listed in the dialog; they are not compared again. A filter the data already have is kept as one of the choices (e.g. low-pass as in the data (30 Hz)), and the result warns when such a filter alone already distorts the measured signal. With the average reference, EEG channels removed before PipeCompare are interpolated back first, so the average covers the whole montage. -
Press Run. A progress window shows how many pipelines are done, the step running now (for example fitting ICA), the time so far and the time left. Before the first pipeline is done (it runs the steps all pipelines share, such as ICA, which can take several minutes) the bar moves back and forth instead of filling. Stop ends the search and keeps the pipelines already finished. The Command Window prints each step and each finished pipeline as the search runs, then a summary; the same log is kept in
pipecompare_last_run.login MATLAB'stempdir. -
The result window says which pipeline to use and why, naming each pipeline by its settings (for example high-pass 0.5 Hz, low-pass 30 Hz). Use this pipeline stores it as a new EEGLAB dataset; Save script… writes it as a MATLAB function that you can run on your other recordings; Show all pipelines lists every pipeline with the reason it was excluded. Use this pipeline runs the steps again, except ICA, whose decomposition comes from the comparison; save the new dataset afterwards with File > Save current dataset as. The new dataset is cut into epochs (segments for band power) and cleaned, ready to average into ERPs or to compute band power; do not filter, re-reference or reject epochs again.
The recommended pipeline is marked with *, followed by the best of the others. The columns are:
| Column | Meaning |
|---|---|
| checks | passed: the pipeline met every constraint; excluded: it broke one (the reason is under Show all pipelines); error: a step failed; not run: the search was stopped first |
| noise (SME) | the standardized measurement error of your measure, divided by the share of a known signal the pipeline keeps. Lower is better: the averaged value is measured more precisely |
| trials kept | the smallest share of trials kept in any condition |
| signal change | how much the pipeline changed the amplitude of the known signal |
| settings | the settings that differ between the pipelines |
The line under the table names the steps that every pipeline shares. When several pipelines cannot be told apart from the best, PipeCompare recommends the one that keeps the most trials, so that data are not cleaned more than they need to be. The result then says whether the recommended pipeline's noise is within 5 % of the best one's (a negligible difference) or whether the data are too few to tell, with how much noisier it could be. It also says how the best pipeline does on trials that were not used to choose it: either its advantage is not luck, or its noise there is about X % higher than it looks (see How a pipeline is chosen). When the settings compared make no difference on these data, the result says so. When no pipeline passes because most lost too many epochs to rejection, it names the channels most often over the rejection limit (likely bad channels to remove or interpolate before running again) and, for data that keep their recorded reference, suggests the average reference.
When the recommended pipeline interpolated an electrode you measure (it was detected as a bad channel, or deleted before PipeCompare and put back), the result says so, for example Pipeline 3 rebuilt Pz, which you measure, from the electrodes around it. A rebuilt electrode is a weighted average of its neighbours and has less noise than a real electrode, so that pipeline's noise (SME) looks smaller than it is and it is favoured over pipelines that keep the electrode. The real electrode's noise cannot be known (it was bad), so this is not corrected; if the electrode is really bad, measure at other electrodes (Electrodes…) and run again.
When pipelines do pass, PipeCompare still looks at why the recommended pipeline rejected its epochs. If one channel (or two or three) was over the limit in at least half of the rejected epochs, and in at least 3 of them, the result names it, for example Most rejected epochs are due to one channel: T7 was over the limit in 40 of the 45 epochs rejected in pipeline 2. Such a channel is probably bad for long stretches without being bad enough over the whole recording for the bad-channel test. Look at it in Plot > Channel data (scroll); if it is bad, mark it as bad or interpolate it (Tools > Interpolate electrodes) and run again, which keeps more trials. The note appears after the headline, in the Command Window and in the advanced panel's notes.
The result also says how large a difference this recording can show, for example With this many trials and this noise, two conditions must differ in P3 by about 6 uV to be told apart; smaller differences need more trials. This comes from the recommended pipeline's own measurement error (SME, the error of your measure in its own units, not the gain-corrected value in the table). For two conditions with errors a and b, the difference between them has the error sqrt(a² + b²); a true difference of 2.8 times that (1.96 + 0.84) is found by a two-sided test at p < .05 in 80% of recordings like this one. With more than two conditions, the pair with the largest error is used; with one condition, the value is compared with 0. When the largest difference actually seen between the conditions (or the value itself, with one condition) is smaller than this, the result adds that more trials (more events, or several recordings) would help more than other preprocessing. That comparison is made in the measure's own units, so it works the same for amplitudes (uV), latencies (ms) and band power (log10 uV²). The same sentence appears in the Command Window and in the advanced panel's result notes. It describes this one recording, trial by trial; a group study's power depends on the number of participants instead.
Under the table, the window lists what the recommended pipeline did, step by step in the order the steps ran; select another row to see that pipeline's steps instead. Each line gives the step's settings and what it decided on these data, for example:
1. Bad channels (kurtosis or joint probability over 5 SD, found on a 1 Hz high-passed copy): O1, O2 interpolated (not tested: VEOG, HEOG)
2. Average reference (left out: VEOG, HEOG)
3. ICA (extended runica), fitted on a 1 Hz high-passed copy and applied to the data (the fit leaves out stretches far noisier than the rest)
4. High-pass filter 0.5 Hz
5. Low-pass filter 40 Hz
6. ICLabel: 4 of 60 components removed (Muscle, Eye, Heart, Line Noise, Channel Noise with probability 0.8 or more); 21 look like brain activity, 9 were labelled Other
7. Epochs -200 to 800 ms around event type(s) ...
8. Baseline -200 to 0 ms removed
9. Epochs beyond +/-150 uV on any channel rejected: 12 of 120
Trials kept per condition: ...
When epochs are repaired (line 8 of the step list), the rejection line adds how many epochs were kept that way, for example Epochs beyond +/-150 uV on any channel rejected: 9 of 120; 7 epoch(s) kept by interpolating up to 3 channel(s) within the epoch, and with the average reference a last line Average reference again (after the epochs repaired by interpolation) follows.
When ICA did nothing useful, the result says so after the headline (see When ICA does nothing).
This is a readable summary of the pipeline's EEG.history; the dataset you get with Use this
pipeline keeps the full history (every EEGLAB command) unchanged. The same list is printed in the
Command Window after a run from pop_pipecompare, and the advanced panel shows it under Steps
when you select a result row.
For any EEGLAB step, order search, several components or different constraints, open Advanced… in the dialog, or Tools > PipeCompare > Advanced panel…. The panel defines the same measures as the dialog (ERP components, N2pc and LRP contralateral minus ipsilateral, band power). See docs/PANEL.md for a description of every control.
Electrodes on the earlobes or mastoids are reference sites, not scalp electrodes. PipeCompare
recognises them automatically, in any dataset, by their standard names: A1 and A2 (10-20
earlobes) and M1 and M2 (mastoids), in any upper or lower case and also with the POL
prefix some EDF exports add (POL A1). A channel whose type is set to REF in the channel
locations counts too. These channels:
- stay in the data, unchanged;
- are not tested as bad channels, so they are never interpolated (the scalp cannot predict them);
- do not count in epoch rejection;
- are left out of an average reference, as EOG channels are;
- are not interpolated back when they were removed before PipeCompare.
This also covers data already referenced to linked ears or mastoids, where A1 and A2 are flat or mirror each other and a bad-channel test would wrongly flag them. The dialog (step 1, Data), the Command Window log and the advanced panel's dataset summary list them, e.g. Ear/mastoid channels left out: A1, A2; in the panel they are filled into each step's exclude list, where you can edit them.
Some channels with these names are not ears, and are left as scalp channels:
- caps numbered by letter, such as BioSemi's A1-A32: when the data also have A3 (or M3), A1 and A2 are scalp electrodes there;
- TP9 and TP10, which sit near the mastoids but are scalp electrodes in the 10-10 system;
- a channel you set to type
EEGin Edit > Channel locations while other channels have other types. Use this to keep A1 and A2 as scalp channels. (When every channel has typeEEG, that is the importer's default, not a choice, and the names decide.)
Before the run, PipeCompare checks whether each channel's signal looks like that of the channels next to it. On the scalp, neighbouring electrodes record similar signals, so a channel that does not resemble its neighbours is either a bad channel or not where its label says it is (for example, two labels swapped when the recording was set up or exported). A wrong label matters more than it seems: bad-channel interpolation and ICLabel both use the positions that come from the labels.
How it works:
- It uses the EEG channels that have a location, leaving out EOG, ECG and similar channels and the ear and mastoid channels. It needs at least 8 of them; with fewer, it is skipped.
- It takes the first 10 minutes of a continuous recording (for epoched data, as many epochs as fit in 10 minutes, each with its mean removed), keeps 1-30 Hz and applies an average reference, all on a copy. Your data are not changed.
- For each channel it computes the mean correlation with its 3 nearest channels (by electrode position).
- A channel is flagged when this correlation is far below the other channels': more than 3 robust standard deviations below their median (median and MAD; the spread counts as at least 0.05), or below 0.1 while the median of the montage is at least 0.3. On caps with few, widely spaced electrodes the median is lower, and then only the first rule applies. The worst channel is flagged first and the check runs again without it, so a swapped channel does not also drag its neighbours down. At most a quarter of the channels are flagged.
- For each flagged channel it names the two channels it resembles most. If these are far away on the head, the label is probably wrong; if it resembles no channel (correlation below 0.3), it is more likely a bad channel.
The result appears in the dialog's first line (Data), in the Command Window log, and in the advanced panel's list of data warnings, for example: Montage check: these channels do not resemble their neighbours ... AF3 (r = -0.50 with its neighbours; most like TP7 r = 0.61, CP6 r = 0.55) .... It is a warning only: nothing is changed and the comparison runs as usual. If channels are flagged, check the montage with whoever recorded the data, correct the labels or locations in Edit > Channel locations, and run again.
After the run, PipeCompare looks at what ICA and ICLabel did in the recommended pipeline (or, if it has no ICA, the first pipeline with ICA), and says so when:
- ICLabel recognised almost none of the components: fewer than 2 components have a Brain probability of 50 % or more, or the median probability of Other is above 0.8. The message gives the numbers, for example ICLabel recognised almost none of the 61 ICA components: 0 look like brain activity (Brain 50% or more) and 56 were labelled Other. ICLabel judges components partly by their scalp maps, so this usually means that the channel labels do not match the electrode positions (see Montage check); too little clean recording for ICA can do the same. In that case the ICA step changed nothing useful.
- No pipeline removed any component. If ICLabel did recognise brain components, the message says that no component passed the artifact threshold, which is expected on clean data. When the threshold was chosen from the data and ICLabel did take some components for artifacts (probability over 0.5), it says instead that removing them did not make the measure measurably more precise, so they were kept.
- ICA had too little data. A usual rule is that ICA needs at least 20 × (number of components)² data points for a reliable decomposition, more being better: with 60 components, 72 000 points, that is 4.8 minutes at 250 Hz or 72 seconds at 1000 Hz. With fewer, the components can mix brain activity and artifacts, and ICLabel cannot sort them. The message gives the numbers, for example ICA had too little data: 30000 data points for 61 components, while a reliable decomposition needs about 20 x 61 x 61 = 74420 or more. This is said whatever ICLabel recognised. The fix is a longer recording; fewer channels (fewer components) also lower the need.
The note appears in the result window after the headline, in the Command Window, and in the advanced panel's notes below the results. Each pipeline's step list also gives, on its ICLabel line, how many components look like brain activity and how many were labelled Other.
The usual rejection drops an epoch when any EEG channel goes beyond ±x µV. That test looks at the distance from zero, which after baseline correction is the distance from the baseline. Slow drifts can carry a clean epoch over the limit, while a blink that starts from a low point can stay under it. ERP CORE (Kappenman et al., 2021) therefore uses a moving-window peak-to-peak test, as ERPLAB's moving window peak-to-peak threshold:
- a window of 200 ms moves over the epoch in steps of 50 ms;
- in each window, for each EEG channel, the highest value minus the lowest value is taken;
- the epoch is rejected when, in any window, this exceeds the limit on any EEG channel (EOG and other non-EEG channels, and ear or mastoid channels, are not tested, as in the usual test).
A blink or a fast jump is a large change within 200 ms and is caught; a slow drift changes little
within 200 ms and is not. Because the limit is a range rather than a distance from zero, its
values are larger: simple mode chooses the peak-to-peak limit from the data, as for the usual
test, and compares 100, 150 and 200 µV peak-to-peak when epochs are repaired (tick measure the
limit peak-to-peak in 200 ms windows under the rejection; 'peaktopeak' in pop_pipecompare).
It can be combined with repairing epochs: the channels over the limit are the ones repaired.
In the advanced panel, the amplitude rejection step (reject_threshold) has the parameters
method (absolute, the default, or peaktopeak; both can be compared, as absolute | peaktopeak), window (the window length in ms, default 200) and uv (any limits, or auto). Each
pipeline's EEG.history records the test as pipecompare.run.Steps.markPeakToPeak(...), which
marks the epochs where EEGLAB's own threshold test puts its marks (EEG.reject.rejthresh).
Which rejection limit is best depends on the recording: a limit that is too strict throws away good trials, one that is too loose keeps artifacts, and both make the measure less precise. Instead of trying three limits, PipeCompare finds the best one in each pipeline, from that pipeline's data at the rejection step:
- For every epoch it takes the value the limit is compared with: the largest absolute value over the tested channels and the whole epoch (peak-to-peak: the largest peak-to-peak value in the moving windows). An epoch is rejected when this value is above the limit.
- A limit between two neighbouring values of this list keeps exactly the same epochs, so there are only as many different results as epochs, and all of them are tried: the limit is the exact best, not the best of a few tries. For each one the SME the ranking uses (RMS over the conditions, and over the measures for several components) is computed at once from sums over the kept trials.
- Only limits that keep at least 10 trials and 50 % of the trials of every condition (the
ranking's own limits,
minTrialsandminRetention) count. - Limits whose SME the data cannot tell apart from the best limit's (the same paired bootstrap and simultaneous intervals as the ranking, over up to 60 limits spread over the range, the best included) are equally good, and the highest of them is used: no epoch is rejected without evidence that rejecting it makes the measure more precise. The number used is the roundest one among the limits that keep the same epochs (for example 120 µV rather than 117.43 µV).
- The limit is then applied with EEGLAB's own test (
pop_eegthresh, or the peak-to-peak marks), soEEG.historyrecords the number, and the signal check removes the same epochs from its copy.
The Command Window shows the limit used and the best one, and the steps of each pipeline in the
results say limit chosen from these data with the value. It needs a score per trial, so it works
for mean amplitude and band power; with peak measures, or when epochs are repaired (the limit
then also decides which channels are interpolated), fixed limits are compared. The scores are
those of the data at the rejection step; it is exact when nothing after the rejection changes them
(the baseline is already removed from them). In the advanced panel and in scripts, set the
reject_threshold parameter uv to auto (it can be compared with fixed limits, as
auto | 100). A saved script keeps auto, so on another recording the limit is found from that
recording's data.
Which ICLabel threshold is best also depends on the recording: a low threshold removes components that hold brain signal, a high one leaves eye and muscle activity in, and the epochs it leaves in change which rejection limit is best. Instead of trying three thresholds, PipeCompare finds the best one in each pipeline, together with the rejection limit that follows:
- For every component it takes the largest ICLabel probability among the artifact classes (eye,
muscle, heart, line noise, channel noise). EEGLAB's
pop_icflagremoves a component when this value is above the threshold (and below 1), so a threshold between two neighbouring values removes exactly the same components. There are only as many different results as components with a value over 0.5 (more likely that artifact than anything else), plus removing nothing, and all of them are tried: the threshold is the exact best, not the best of a few tries. - Removing components is a linear operation (the data minus the removed components' part, as
pop_subcompcomputes it), and so are epoching and baseline removal, which come between it and the rejection. So each set of components is taken out of the epoched data directly, without running the steps again. - For each set, the rejection limit is then found as in Rejection limit chosen from the data (every limit; a fixed limit when the step has one), and the SME at the best limit is that set's score. The SME is divided by the signal gain, which removing components can change: the known signal of the signal check goes through the same removal, and a set that fails the signal check (for example because it removes a component that carries the measured signal) does not count, nor does one that keeps too few trials.
- Sets whose SME the data cannot tell apart from the best set's (the same paired bootstrap and simultaneous intervals as the ranking) are equally good, and the one removing the fewest components is used: no component is removed without evidence that removing it makes the measure more precise. The number used is the roundest threshold that removes exactly that set (for example 0.8 rather than 0.8137).
- The threshold is then applied with EEGLAB's own functions (
pop_icflag,pop_subcomp), soEEG.historyrecords the number, and the rejection step that follows finds its limit on the data that are left, as it would after any threshold.
The Command Window shows the threshold used and the best one, and the steps of each pipeline in
the results say threshold chosen from these data with the value. It needs a score per trial, so
it works for mean amplitude and band power; with peak measures, or when epochs are repaired, fixed
thresholds are compared. It is exact when only epoching, baseline removal and an amplitude
rejection that removes epochs come after it, which the plan checks; the signal check of each set
uses all epochs (the rejected ones change it only where epochs overlap), and the full pipeline's
own signal check is the one the ranking uses. In the advanced panel and in scripts, set the
icremove parameter threshold to auto; every pipeline below it must then have the same steps
after it. When the search adopts a pipeline, the threshold it chose is reused; a saved script keeps
auto, so on another recording the threshold is found from that recording's data.
ICA finds the independent sources in the data. Stretches of the recording that are far noisier than the rest, such as movement or a loose electrode for a few seconds, take up components of their own and leave fewer for eye, muscle and brain sources, so the decomposition is worse for the whole recording. PipeCompare therefore fits ICA on a copy without those stretches, and applies the result to all of the data (nothing is cut from your data):
- continuous data are divided into 1 s windows (epoched data: each epoch is one window);
- for each window, the spread (standard deviation) of every EEG channel is taken, and the mean of their logarithms gives the window's noise level (EOG and other non-EEG channels do not count, so blinks stay in and ICA can learn them);
- a window is left out when its noise level is more than 3 robust standard deviations above the median of all windows (median and MAD), at most 20 % of the windows, the noisiest first.
On clean data nothing is left out. The Command Window and each pipeline's EEG.history say how
much was left out. In the advanced panel, the ICA step's parameter fitClean (1 by default)
turns this off with 0, and can be compared (0 | 1).
Sometimes a single electrode loses contact for a moment: in a few epochs one channel goes far over
the rejection limit while all the others are fine. Rejecting those epochs loses trials because of
one channel. With repair epochs with up to 3 channels over the limit ticked (line 8 of the
step list; 'epochinterp' in pop_pipecompare), the rejection step works as follows, in each
epoch:
- It tests every EEG channel against the limit, as usual (non-EEG and ear/mastoid channels are not tested).
- If 1, 2 or 3 channels are over the limit, and they all have locations, those channels are
replaced in that epoch only by spherical-spline interpolation from the other channels of the
same epoch (the same computation as EEGLAB's
eeg_interp), and the epoch is kept. An electrode you measure (for example Pz for the P3) is never repaired: if it is one of the channels over the limit, the epoch is rejected. A rebuilt electrode is an average of its neighbours and has less noise than a real one, so repairing it would make the pipeline look more precise than it is. - If more than 3 channels are over the limit, the epoch is rejected as before.
This is the idea of the epoch-level channel interpolation in FASTER (Nolan, Whelan & Reilly,
2010). It runs after ICA, filtering, epoching and baseline correction, as part of the rejection
step, so it sees the same data the limit is applied to. Repaired epochs are not tested again.
With the average reference, the data are averaged again afterwards: the replaced values were part
of that epoch's average, which would otherwise keep a share of the bad signal in every channel.
The signal check applies exactly the same repairs (same channels, same epochs) to its copy, so any
change to the measured signal is counted. Repaired epochs do not count towards the 20 % limit on
interpolated channels, which is about whole channels. Each pipeline's EEG.history records the
repairs as one command listing the channels and epochs.
In the advanced panel, every rejection step (amplitude, joint probability, kurtosis) has the
parameter interpolate: the largest number of flagged channels an epoch may have and still be
repaired (0, the default, turns it off). Set any number there, or several (for example 0 | 3)
to compare pipelines with and without repairs. For joint probability and kurtosis, the flagged
channels are those over the per-channel limit.
Everything in the menu can also be started from MATLAB's Command Window. PipeCompare works on the
dataset that is current in EEGLAB, so start EEGLAB and load the dataset first; this creates the
variables EEG, ALLEEG and CURRENTSET that PipeCompare reads and writes:
eeglab % start EEGLAB (adds PipeCompare to the path)
EEG = pop_loadset('filename', 'mydata.set', 'filepath', '/path/to/folder');
[ALLEEG, EEG, CURRENTSET] = eeg_store(ALLEEG, EEG, 0); % make it the current EEGLAB dataset
eeglab redrawWhen the dataset is already open in EEGLAB's window, skip these lines. If EEG in the Command
Window is not the dataset selected in EEGLAB (for example after EEG = ... on another dataset),
pop_pipecompare stops with make this dataset current first; EEG = ALLEEG(CURRENTSET); makes
them the same again.
Open the windows of the menu:
[EEG, LASTCOM] = pop_pipecompare(EEG); if ~isempty(LASTCOM), eegh(LASTCOM); end % the simple dialog (Tools > PipeCompare > Compare pipelines...)
pipecompare.PipeCompare.app(); % the advanced panel (Tools > PipeCompare > Advanced panel...)eegh(LASTCOM) adds the command that repeats your choices to EEGLAB's history, as the menu does.
Or skip the dialog and give the choices as options. The menu dialog records an equivalent
command in EEGLAB's command history (ALLCOM):
% ERP: compare pipelines for the P3, one condition per event type
EEG = pop_pipecompare(EEG, 'measure', 'P3', 'events', {'target', 'standard'}, 'recipe', 'standard');
% N2pc, contralateral minus ipsilateral: the event types of each target side
EEG = pop_pipecompare(EEG, 'measure', 'N2pc', 'left', {'111', '112'}, 'right', {'121', '122'});
% Continuous data: compare pipelines for alpha-band power in 2 s segments
EEG = pop_pipecompare(EEG, 'measure', 'alpha', 'recipe', 'filters');
% Only some steps (any of 'badchannels', 'ica', 'highpass', 'lowpass', 'reject', 'epochinterp'),
% always run in that order
EEG = pop_pipecompare(EEG, 'measure', 'P3', 'events', {'target'}, 'steps', {'highpass', 'lowpass', 'reject'});
% Standard plus repairing epochs that have up to 3 channels over the rejection limit
EEG = pop_pipecompare(EEG, 'measure', 'P3', 'events', {'target'}, ...
'steps', {'badchannels', 'ica', 'highpass', 'lowpass', 'reject', 'epochinterp'});
% A component with your own electrodes instead of its ERP CORE site
EEG = pop_pipecompare(EEG, 'measure', 'P3', 'channels', {'Pz', 'CPz', 'POz'}, 'events', {'target'});
% Your own window (s) and electrodes; or 'measure', 'band', 'band', [8 12]
EEG = pop_pipecompare(EEG, 'measure', 'custom', 'window', [0.25 0.5], 'channels', {'Cz', 'CPz'}, ...
'events', {'target'}, 'recipe', 'standard');Replace 'target' and 'standard' with the event types in your dataset. With options, the
progress window (with Stop) and the result window open as from the menu; add 'show', 'off'
to run without any window, for example in a script that runs unattended (the Command Window
still prints each step):
EEG = pop_pipecompare(EEG, 'measure', 'P3', 'events', {'target'}, 'reference', 'average', 'show', 'off');The dataset is returned unchanged. The result is stored in the variable pipecompare_result
(also the third output, [EEG, com, result] = pop_pipecompare(...)), and these commands do what
the result window's buttons do:
pipecompare.PipeCompare.adopt(pipecompare_result); % Use this pipeline: new EEGLAB dataset, now current
pipecompare.PipeCompare.adopt(pipecompare_result, 12); % ... or pipeline 12 (one that passed the checks)
pipecompare.PipeCompare.writeScript(pipecompare_result, [], 'my_pipeline.m'); % Save script...
pipecompare.PipeCompare.script(pipecompare_result); % print the recommended pipeline's EEGLAB commandsAfter adopt, EEG in the Command Window is the new dataset (cut into epochs and cleaned); save
it with pop_saveset. The saved script is a function: run it on any recording with the same
channels and event types, for example EEG = my_pipeline(EEG); (the file must be in the current
folder or on the MATLAB path). The file name is the function's name, so it may contain only
English letters, digits and underscores, must start with a letter, and must not be the name of an
existing MATLAB or EEGLAB function (such as pop_epoch); PipeCompare refuses other names. Type
help pop_pipecompare for all options.
For full control, describe what is measured (an analysis contract) and what may vary (a plan):
pipecompare.PipeCompare.state(); % summary of the current dataset and its history
% What is measured: conditions, epoch, baseline, and the measure (window, channels, type)
c = pipecompare.eval.Contract( ...
'conditions', {'target', {'target'}; 'standard', {'standard'}}, ...
'epoch', [-0.2 1.0], ...
'baseline', [-0.2 0], ...
'components', {'P3', [0.30 0.60], {'Pz', 'CPz', 'POz'}, 'mean'});
% What may vary
p = pipecompare.plan.Plan();
p = p.add('highpass'); % cutoff searched over 0.1, 0.3, 0.5, 1 Hz
p = p.add('lowpass', 'cutoff', 30); % fixed value
p = p.addEeglab('EEG = pop_reref(EEG, []);', 'reref'); % one EEG = pop_*(EEG, ...) call as a step
p = p.add('ica', 'fitHighpass', 1);
p = p.add('icremove', 'threshold', {0.8, 0.9}); % searched over two values
p = p.add('epoch'); % windows come from the contract
p = p.add('baseline');
p = p.add('reject_threshold', 'uv', {100, 150}, 'interpolate', {0, 3}); % with and without repairing epochs
r = pipecompare.PipeCompare.optimize(p, c, struct('checkpoint', 'pc_run1'));
pipecompare.PipeCompare.adopt(r); % recommended pipeline -> new dataset
pipecompare.PipeCompare.writeScript(r, 3, 'pipeline3.m'); % candidate 3 as a function for any recording
r = pipecompare.PipeCompare.resume('pc_run1'); % continue an interrupted searchBand power of continuous data uses a band-power contract and the same plan without the baseline
step (segments have no baseline window, so a baseline step would make every pipeline fail):
c = pipecompare.eval.Contract('analysis', 'bandpower', 'segment', 2, ...
'bands', {'alpha', [8 12], {'O1', 'Oz', 'O2'}});By default, steps run in the order they were added. If that order is invalid, PipeCompare lists
the conflicts instead of rearranging the steps. Set p.OrderMode = 'search' to compare every
valid order; p = p.pin(id) keeps a step in place and p = p.before(a, b) makes step a run
before step b (a plan is a value object, so assign the result).
The available steps are listed in docs/COVERAGE.md. Search options (for
example parallel, checkpoint and maxLeaves) are documented in help pipecompare.run.Executor,
and constraints and ranking options in help pipecompare.eval.Rank.
-
Feasibility. Pipelines that fail or violate a constraint are excluded, each with a stated reason. Default constraints:
- at least 10 trials and 50 % retention per condition;
- at most 20 % of channels interpolated;
- preservation of the known signal: amplitude error ≤ 10 %, peak shift ≤ 10 ms, artifactual deflection ≤ 5 %, waveform correlation ≥ 0.95, topography correlation ≥ 0.90 (for band power only the amplitude error and topography correlation apply).
The filters are checked first. Filters act on a signal the same way whatever the data, so what a pipeline's filters alone (high-pass, low-pass, line-noise filter) do to the known signal is computed before anything runs, on a short stretch of the signal copy around the first events (with more than a filter length on each side). A pipeline whose filters alone break one of these limits (the topography aside, which a filter does not change) is excluded without being run, with a reason that starts with filters alone; for example a 1 Hz high-pass for a broad P3. This saves the filtering of the whole recording and everything after it for those pipelines. It is done for ERP measures on continuous data, when no resampling or epoching comes before the last filter;
'filterCheck', falseturns it off. -
Precision. Feasible pipelines are scored by their SME, divided by the gain each pipeline applies to the known signal. This makes the comparison one of signal-to-noise ratio, so a pipeline cannot appear more precise just by attenuating everything (Zhang, Garrett & Luck, 2024).
-
Uncertainty. A paired bootstrap over trials, with simultaneous intervals across all pairs of candidates (White, 2000; Romano & Wolf, 2005), finds the set of pipelines that the data cannot distinguish from the best.
-
Equivalence. Not distinguished can mean that the pipelines are equally good or that there are too few trials to tell. A pipeline is called equivalent to the best only when the upper bound of its difference from the best is within 5 % of the best's objective (
equivalenceMargin, default 0.05); the bound is simultaneous over the pipelines not shown worse. The results say which of the two applies to the recommended pipeline. -
Selection check. The best of many pipelines looks better than it is, because part of its advantage is chance. PipeCompare therefore also scores the choice on trials that were not used to make it: the trials of each condition are split into two halves at random, the pipeline that is best on one half is scored on the other half (rescaled to all trials), and the other way round, over 20 random splits (
nSplits; consecutive halves for segments of one recording; 0 turns the check off; at least 4 trials per condition are needed). The result reports how much better the best looks than it is on those trials. The pipelines' own decisions (bad channels, ICA, rejected epochs) were made on all trials; only the choice between pipelines is held out. -
Recommendation. Within the set of step 3, PipeCompare recommends the least aggressive pipeline: the one with the highest trial retention, then the least signal distortion.
The ICLabel threshold and the epoch-rejection limit of the Standard recipe are chosen inside each pipeline, before this ranking, with the same SME, constraints and bootstrap (see ICLabel threshold chosen from the data and Rejection limit chosen from the data).
Pipelines that differ in something that changes the measured quantity, such as the reference, are
ranked separately and never compared with each other. With several such groups there is no overall
recommendation: each group has its own, and adopt and writeScript then need a candidate id.
A multiverse summary reports, for each measure and condition, how much its value varies over the
feasible pipelines and which searched choice accounts for most of that variation (η²); it is
descriptive and not used in the ranking. The full derivations are in
docs/METHODS.md.
- A ranked table of all pipelines: status, objective (gain-corrected SME), simultaneous interval of its difference from the best, trial retention, interpolation, signal-check results, and the reason for each exclusion.
- The recommended pipeline and why it was chosen.
- The full EEGLAB command sequence for every candidate.
- A MATLAB function that runs a chosen candidate's steps on any recording, deciding bad channels,
components and rejected epochs from that recording's data (
writeScript). - Adoption of a candidate as a new EEGLAB dataset whose
EEG.historyreplays it (adopt).
- The analysis contract is fixed by you. Event types, epoch and baseline windows, regions of interest and measurement windows define what is measured, and are never searched.
- The signal check is necessary, not sufficient. The known signal has an assumed shape: a smooth bump as wide at half maximum as your measurement window (at least 50 ms), with a topography centred on the region of interest. Real components may be affected differently.
- Some steps are re-run on the signal copy. EEGLAB commands added as steps that make their
own data-driven decisions, other than mark-and-remove workflows, are re-run on the copy carrying
the known signal rather than replayed; this includes ASR added as an EEGLAB command. The built-in
asrstep replays the reconstructions ASR chose and is re-run only when they cannot be reproduced exactly. The result flags every re-run step. - Reference. The dialog offers the recorded reference or the average reference. For another reference (e.g. linked mastoids), add a re-reference step with those channels in the panel or a script.
- Channel locations are needed for spherical interpolation and ICLabel. Without them those steps are excluded before the search, with the reason.
- Peak latency is a weak criterion. Its precision is hard to estimate with few trials, so it rarely separates pipelines; mean-amplitude measures are more informative.
- Computation cost. Every pipeline also runs on the signal copy, roughly doubling the EEGLAB computation. Memory use grows with plan depth; the expected peak is reported when it exceeds 2 GB. Resampling early in the plan reduces both.
- Single-dataset scope. STUDY-level processing, time-frequency measures, connectivity and source analysis are not covered.
docs/COVERAGE.md lists every supported operation and how it has been validated.
eeglab nogui % EEGLAB and its plugins must be on the path
addpath(fullfile(pwd, 'tests')); % run from the PipeCompare folder
results = run_all();The suite runs on synthetic data with known ground truth. It covers history parsing, plan
enumeration (checked against brute force), the statistics (validated by simulation), signal
preservation for every built-in step except resample, end-to-end searches with resume and
parallel execution, and EEGLAB and panel integration. Continuous integration runs every suite,
including those that open windows (on a virtual display); tests that need a toolbox CI does not
install, such as parallel execution (Parallel Computing Toolbox), are skipped there. Steps that
require clicking in EEGLAB's own dialogs cannot be automated and are listed in
tests/MANUAL_GUI_CHECK.md.
Two optional checks use your own data, which is never committed: set the environment variable
PIPECOMPARE_REAL_SET to a .set file, or run tests/realdata_validation.m.
| Document | Contents |
|---|---|
| docs/PANEL.md | The dialog and the advanced panel, control by control |
| docs/METHODS.md | Measures, statistics and signal check, with derivations and references |
| docs/COVERAGE.md | Supported steps and their validation status |
| CHANGELOG.md | Release history |
| CONTRIBUTING.md | Bug reports, documentation fixes and tests |
If you use PipeCompare in published work, please cite it:
@software{pipecompare2026,
author = {abwoo},
title = {PipeCompare: data-driven comparison of EEG preprocessing pipelines in EEGLAB},
year = {2026},
version = {0.9.4},
url = {https://github.com/abwoo/PipeCompare}
}Please also cite the methods it builds on:
- Luck, S. J., Stewart, A. X., Simmons, A. M., & Rhemtulla, M. (2021). Standardized measurement error: A universal metric of data quality for averaged event-related potentials. Psychophysiology, 58(6), e13793.
- Kappenman, E. S., Farrens, J. L., Zhang, W., Stewart, A. X., & Luck, S. J. (2021). ERP CORE: An open resource for human event-related potential research. NeuroImage, 225, 117465.
- Zhang, G., Garrett, D. R., & Luck, S. J. (2024). Optimal filters for ERP research I: A general approach for selecting filter settings. Psychophysiology, 61(6), e14531.
- White, H. (2000). A reality check for data snooping. Econometrica, 68(5), 1097–1126.
- Romano, J. P., & Wolf, M. (2005). Stepwise multiple testing as formalized data snooping. Econometrica, 73(4), 1237–1282.
The pipelines run on EEGLAB (Delorme, A., & Makeig, S. (2004). Journal of Neuroscience Methods, 134(1), 9–21); when the chosen pipeline removes components with ICLabel, cite it too (Pion-Tonachini, L., Kreutz-Delgado, K., & Makeig, S. (2019). NeuroImage, 198, 181–197). The complete reference list is in docs/METHODS.md.
PipeCompare is released under the MIT License.