Medical Image Preprocessing: Techniques, Workflows, and Best Practices
A central reader opens a follow-up brain MRI and the tumor margin looks different from the baseline scan, not because the tumor has changed, but because the scanner, the coil, and the acquisition protocol have. What is easy to miss is that most of what looks like a raw image has already been through a substantial amount of processing before it ever reaches PACS, let alone an analyst's screen.
Medical image preprocessing is often described as a pipeline a research team runs on raw pixel data: denoising, normalization, bias correction, resampling. For images that arrive as DICOMs from a modern scanner, much of that medical image processing has already happened, embedded in the reconstruction pipeline the scanner uses to turn raw signal into the DICOM it sends to PACS. What matters for imaging core labs and image analytics teams downstream is knowing which corrections are already applied, which are still performed afterward, and which distortions no amount of post-processing can remove.
Medical image preprocessing techniques and when to use them
Not every technique applies to every image, and not every one is something a downstream team applies itself. Items 1 through 5, background removal, denoising, normalization, bias correction, and resampling, are usually embedded in the scanner's own reconstruction pipeline and already applied by the time a DICOM reaches PACS. Registration (item 6) is more often a genuine post-processing step, applied later in dedicated workstation software. Artifact handling and contrast (items 7 and 8) work differently again: most MRI artifacts can only be detected, not corrected, and MRI contrast is primarily a property of the acquisition itself.
1. Background removal and region-of-interest extraction
Background removal, also called region of interest (ROI) segmentation, isolates the anatomy that matters from everything else in the frame, scanner bed, air, non-target organs, so that later steps are not skewed by irrelevant pixels. In brain MRI, this usually means skull stripping: removing the skull, scalp, and other non-brain tissue before segmentation or volumetric analysis. Atlas-based tools, such as ANTs' brain extraction workflow, register the raw scan to a labeled template and use the resulting mask to isolate brain tissue automatically.
This step is closely tied to medical image segmentation, which typically consumes the ROI mask as its starting point.
2. Image denoising
Every acquisition introduces some degree of random intensity fluctuation, from thermal noise in MRI receiver coils to quantum mottle in low-dose CT. Denoising reduces that noise while preserving the structural detail a radiologist or a model actually needs to see.
Common methods include Gaussian and median filtering for straightforward noise reduction, wavelet-based denoising for a better balance between smoothing and edge preservation, and deep learning-based denoisers for modality-specific noise patterns, such as speckle in ultrasound. The tradeoff is real: as freeCodeCamp's guide to preprocessing medical images puts it, aggressive denoising can erase the features a diagnosis or a machine learning model actually needs, so filters should be tuned and validated per modality rather than applied with default parameters.
3. Intensity normalization and standardization
Unlike CT, where Hounsfield units give every voxel a physically calibrated meaning, MRI intensities are arbitrary. The same tissue can read as a different gray value on two different scanners, or even on the same scanner a year apart. Intensity normalization rescales pixel or voxel values onto a common scale so that images become comparable across patients, sites, and time points.
Simple approaches clip intensities to a percentile range and rescale to the image's data type range. More robust approaches, such as z-score normalization or the Nyúl method, rely on histogram matching: a standard intensity scale is learned from a reference set of images, and every new scan's histogram is mapped onto that scale, an approach detailed in this histogram-based normalization study on brain MRI. WhiteStripe normalization, anchored to normal-appearing white matter, is a common variant for brain MRI specifically.
A version of this normalization is commonly built into the scanner's own reconstruction pipeline, GE's PURE, Philips' CLEAR, and Siemens' Normalize are examples, though multicenter pipelines often reapply an independent step, since a manufacturer's correction is tuned to a single scanner, not a multi-site dataset.
4. Bias field and intensity inhomogeneity correction
MRI has a preprocessing problem that CT does not: intensity inhomogeneity, often called the bias field. Imperfections in coil sensitivity and the magnetic field itself produce a smooth, low-frequency shift in brightness across the image, so the same tissue can appear brighter on one side of the brain than the other, even though nothing about the anatomy has changed. Left uncorrected, this bias field can distort segmentation, throw off intensity normalization, and introduce false variability into any quantitative analysis.
The N4ITK algorithm, an improvement on the earlier N3 method, is the standard research approach: it fits a smooth B-spline model to the bias field, then divides it out of the image. A study on background parenchymal enhancement in breast MRI found that removing this bias is critical for spatially consistent image interpretation.
Most scanner vendors ship their own proprietary bias correction as part of reconstruction, already burned into the DICOM by the time it reaches PACS. When a research team reapplies a method such as N4ITK on top of that, since vendor implementations vary by scanner model, it should still run before intensity normalization, not after, a sequencing detail that is easy to get wrong.
5. Resampling and resizing
Resampling changes the pixel or voxel size of an image without altering what it represents spatially, standardizing resolution across a dataset collected on different scanners or protocols, a goal MathWorks' overview of medical image preprocessing describes as reducing acquisition artifacts while standardizing a dataset. A 3D volume acquired with 1.2mm slices needs to be resampled to a common isotropic voxel size, such as 1mm cubed, before it can be reliably compared or fed into a model with a fixed input size.
The voxel size a scanner outputs is already fixed at reconstruction, so resampling to a shared grid is usually a downstream step, applied once images from different scanners or sites need to sit in the same dataset. For 2D images, resizing with an appropriate interpolation order and anti-aliasing preserves quality; the same principle applies to voxel spacing in 3D.
6. Image registration and spatial normalization
Registration aligns two or more images to a shared coordinate system, whether that means comparing a follow-up scan to a baseline in the same patient, fusing a PET scan with a CT for anatomical context, or mapping a subject's anatomy onto a standard template for population-level analysis. The alignment can be rigid, affine, or fully deformable, depending on how much the anatomy is expected to differ between the images being aligned.
Tools such as ANTs' antsRegistration use multiscale, mutual-information-based nonlinear registration, a step fMRIPrep's preprocessing workflows apply as standard practice in neuroimaging to make results comparable across subjects and studies.
Unlike the previous five techniques, coregistration is reliably a downstream, post-processing step rather than something baked into the DICOM at the scanner, typically applied in dedicated workstation software when a reader needs to compare a follow-up scan to a baseline or fuse two modalities such as PET and CT.
7. Artifact detection and correction
Not every distortion in a medical image comes from noise or intensity scaling, some come from the acquisition itself. Motion artifacts appear as blurring or ghosting when a patient moves during a scan. Susceptibility artifacts distort MRI near metal implants or air-tissue boundaries. Metal artifacts in CT show up as streaking from dense objects like surgical hardware or dental fillings. Each needs to be identified before quantitative analysis, or it can be mistaken for a clinical finding.
For MRI specifically, most artifacts come from acquisition imperfections rather than anything introduced afterward, and there is often no way to recover the information a later step would need to correct them. In practice, the realistic goal downstream is detection, not correction: flagging a motion-affected scan or a susceptibility-compromised region, rather than assuming it can be processed away. Genuine correction, such as retrospective motion correction or CT metal artifact reduction, is narrow and applies only in limited cases.
8. Contrast enhancement and intensity adjustment
Even a well-normalized, artifact-free image can look like it needs more contrast, but for MRI, contrast is primarily a property of the acquisition rather than something adjusted afterward. MRI signal is relative, not physically calibrated the way CT Hounsfield units are, and a study is typically built from multiple acquisitions rather than one image enhanced in post-processing. Techniques such as Contrast Limited Adaptive Histogram Equalization (CLAHE) exist, but they are not commonly used on clinical MRI data in practice.
For CT, windowing serves a related but distinct purpose: because Hounsfield units are already calibrated, adjusting the window level and width lets a reader emphasize bone, soft tissue, or lung detail from the same scan. This changes how the image is displayed, not the underlying acquired data.
How to build a medical image preprocessing pipeline
Individually, each of these techniques is well understood. What is less obvious is where each one happens, and by the time an image reaches an imaging core lab, a meaningful part of that sequence has usually already run.
Confirming the scanner, protocol, and acquisition parameters through DICOM metadata extraction is a useful first step, since it indicates which modality-specific processing was likely applied, and by whom, before the file reached you. For a typical brain MRI, that means background suppression, denoising, bias correction, intensity normalization, and resampling, tuned to that scanner and protocol and already burned into the pixel values by the time the DICOM reaches PACS, not something a downstream team redoes from raw data, since the raw signal is rarely available outside the scanner console.
Registration works differently: coregistering a follow-up scan to a baseline, or fusing modalities such as PET and CT, is usually a genuine post-processing step, applied later in dedicated workstation software. Artifact handling and contrast follow yet another logic: most MRI artifacts can only be detected, not corrected, and MRI contrast is primarily a function of the acquisition, not something adjusted afterward. CT differs on both counts, since Hounsfield units are already calibrated, so windowing, denoising, and metal artifact correction matter more than bias correction or normalization.
The practical takeaway is knowing which corrections are already embedded in the DICOM you received, which ones your workflow still needs to apply, typically registration, and which distortions no processing step can undo. That awareness is what turns research-ready imaging data into data that can actually be trusted for analysis.
Best practices for reliable medical image preprocessing
A preprocessing pipeline that works once on a handful of test images is not the same as one that holds up across a multicenter trial or a production model. A few practices separate the two:
- Know what has already happened. Check whether the scanner already applied a correction before reapplying it yourself.
- Understand the data first. Know the acquisition protocol, scanner variability, and modality-specific quirks, arbitrary MRI intensities, calibrated CT Hounsfield units, ultrasound speckle, before choosing a technique.
- Preserve the original data. Always keep the DICOMs as received. Preprocessing should be a reproducible transform, never a destructive edit.
- Document every step and parameter. An audit trail of preprocessing decisions is what makes a result defensible on review in a regulated clinical trial context.
- Validate the output. Spot-check preprocessed images against the originals, not just at the end.
- Apply the same pipeline consistently across the full dataset. Inconsistent preprocessing introduces the same kind of variability the process is meant to eliminate.
- Tailor preprocessing to the downstream analysis. A pipeline built for segmentation is not automatically the right one for radiomics or for training a deep learning model.
- Plan for missing or corrupted data. Define upfront how the pipeline handles incomplete series, failed acquisitions, or unreadable files, rather than discovering the gap mid-analysis.
Medical image preprocessing tools and software
The right tool usually depends on the modality, the programming environment, and whether the pipeline plugs into a larger clinical trial imaging workflow. Since items 1 through 5 are largely handled at the scanner, these tools see their heaviest use downstream, for registration and multi-site harmonization.
| Tool | Type | Best for |
| MATLAB Medical Imaging Toolbox | Commercial, GUI and scripting | End-to-end preprocessing across modalities with built-in visualization |
| SimpleITK | Open source, Python / C++ / R | Bias field correction, filtering, and registration via the ITK engine |
| NiBabel | Open source, Python | Reading and writing neuroimaging file formats, including NIfTI and DICOM |
| ANTs | Open source, command line and Python | State-of-the-art deformable registration and brain extraction |
| FSL | Open source, command line and GUI | Brain imaging analysis, tissue segmentation, and motion correction |
| SPM | Open source, MATLAB-based | Statistical parametric mapping for functional and structural brain MRI |
| TorchIO | Open source, Python (PyTorch) | Preprocessing and augmentation pipelines for deep learning on 3D volumes |
Preparing medical images for reliable analysis
Preprocessing rarely gets the attention that segmentation models or reader adjudication does, but knowing what has already happened to an image, inside the scanner, before it reaches PACS, matters just as much as anything a downstream pipeline does to it. Treating a DICOM as fully raw when it has already been denoised, bias corrected, and normalized by the vendor is one of the quieter ways a multicenter study can introduce the inconsistency preprocessing is meant to eliminate.
Building an image analytics workflow that holds up across scanners, sites, and time points starts with knowing which corrections are already baked into the DICOMs you receive. This is the kind of problem imaging core labs and central reading teams solve every day, using imaging data for clinical research that has been standardized before a single reader looks at it, and AI-ready medical imaging datasets that hold up under model training and validation.
If you are evaluating how to standardize image analytics across a multicenter imaging program, talk to the Collective Minds team about how imaging core lab workflows handle it at scale.
Things you might be wondering...
What is medical image preprocessing?
Medical image preprocessing is the set of operations, such as denoising, normalization, bias correction, registration, and resampling, applied to medical images before they are used for diagnosis, research, or machine learning. Many are already applied by the scanner's reconstruction pipeline before the DICOM is created; registration is typically the one still applied downstream.
What are the most common image preprocessing techniques?
The most widely used techniques are background removal and ROI extraction, denoising, intensity normalization, bias field correction, resampling, and image registration. Most of these are usually already applied during scanner reconstruction; registration is the one most often still performed downstream. Which techniques apply, and in what order, depends on the imaging modality and the analysis that follows.
What preprocessing steps are commonly used for MRI images?
MRI preprocessing typically includes bias field correction and intensity normalization, both usually already applied by the scanner, skull stripping and registration to a template or prior scan, which are more commonly applied downstream, and resampling to a common voxel size across scanners. Because MRI intensities are arbitrary, knowing which normalization has already been applied matters as much as applying it correctly.
What is the difference between medical image processing and preprocessing?
Preprocessing prepares images for use, correcting noise, artifacts, and inconsistencies so the data is clean and standardized, whether that happens automatically at the scanner or afterward in a downstream pipeline. Image processing is the broader field that includes preprocessing along with downstream tasks like segmentation, feature extraction, classification, and diagnosis. Preprocessing is a subset of image processing, not a separate discipline.
Reviewed by: Pilar Flores Gastellu on September 22, 2026



