Skip to content

Repository files navigation

Note: This repository is a copy of the original code hosted at gitlab.inria.fr/yassis/DeepAneSeg, developed during my PhD at Inria/LORIA.

DeepAneSeg

CI

Detection of intracranial aneurysms in Time-of-Flight MRA with a 3D U-Net trained on small patches. This is the code of the paper An Efficient Data Strategy for the Detection of Brain Aneurysms from MRA with Deep Learning (MICCAI 2021 DALI workshop), packaged as a set of commands that go from raw scans to detections and evaluation metrics.

Method

The main difficulty is the extreme class imbalance: aneurysms occupy a tiny fraction of a scan. Rather than a new architecture or loss, the paper proposes a data strategy:

  • Fast annotation: each aneurysm is approximated by a sphere, defined by two points on a diameter.
  • Small patches, to increase the number of training samples.
  • Sample selection and synthesis: negative patches are centered half on blood vessels and half in the parenchyma, and patches containing an aneurysm are duplicated and deformed by a random 3D spline transform.

Trained with a binary cross-entropy loss on 111 patients (155 aneurysms), a 5-fold cross-validation reached a sensitivity of 0.72 with 0.14 false positives per patient (ADAM challenge criteria).

Abstract

The detection of intracranial aneurysms from Magnetic Resonance Angiography images is a problem of rapidly growing clinical importance. In the last 3 years, the raise of deep convolutional neural networks has instigated a streak of methods that have shown promising performance. The major issue to address is the very severe class imbalance. Previous authors have focused their efforts on the network architecture and loss function. This paper tackles the data. A rough but fast annotation is considered: each aneurysm is approximated by a sphere defined by two points. Second, a small patch approach is taken so as to increase the number of samples. Third, samples are generated by a combination of data selection (negative patches are centered half on blood vessels and half on parenchyma) and data synthesis (patches containing an aneurysm are duplicated and deformed by a 3D spline transform). This strategy is applied to train a 3D U-net model, with a binary cross entropy loss, on a data set of 111 patients (155 aneurysms, mean size 3.86 mm ± 2.39 mm, min 1.23 mm, max 19.63 mm). A 5-fold cross-validation evaluation provides state of the art results (sensitivity 0.72, false positive count 0.14, as per ADAM challenge criteria). The study also reports a comparison with the focal loss, and Cohen's Kappa coefficient is shown to be a better metric than Dice for this highly unbalanced detection problem.

Quick start

git clone https://github.com/youssefassis/DeepAneSeg && cd DeepAneSeg
uv sync --extra cu126                                                    # or --extra cpu, see Installation
uv run deepaneseg-prepare source_dir=ADAM data_dir=Data 'labels=[1]'     # or lay out your data as in Data
uv run deepaneseg-remove-skull data_dir=Data
uv run deepaneseg-extract-points data_dir=Data
uv run deepaneseg-preprocess data_dir=Data
uv run deepaneseg-train data_dir=Data name=exp1
uv run deepaneseg-predict train_dir=Data/0Work/exp1
uv run deepaneseg-evaluate train_dir=Data/0Work/exp1

Each step is described in Usage, and its results in Outputs.

Installation

The project uses uv and Python 3.10+. From the repository folder, pick the PyTorch build:

uv sync --extra cu126    # NVIDIA GPU (CUDA 12.6)
uv sync --extra cpu      # CPU only, also the build used on macOS

This creates .venv with the exact versions of uv.lock. Commands then run with uv run, or directly once .venv is activated.

Without uv, pip install git+https://github.com/youssefassis/DeepAneSeg installs the package and its commands (with PyTorch's default build from PyPI).

Data

Layout

Each patient is a folder named P followed by four digits:

Data/
  P0001/
    volume.nii.gz    TOF-MRA scan
    F.csv            aneurysms: two points (x, y, z columns, in mm) on a diameter of each aneurysm
    config.json      {"init volume": "volume.nii.gz", "pts aneurysm": "F.csv"}
    noskull.nii.gz   skull-stripped scan (written by deepaneseg-remove-skull)
    points.csv       negative patch centers (written by deepaneseg-extract-points)
  P0002/
  ...
  0Work/             splits, trainings and logs (see Outputs)

Patients without aneurysm simply have no pts aneurysm entry. Coordinates are in the scanner (RAS) space of the NIfTI file, so the points can be placed with any viewer, e.g. as markups in 3D Slicer.

Converting a dataset

deepaneseg-prepare converts a dataset of scans with aneurysm label masks to this layout. Its defaults match the ADAM challenge (<case>/orig/TOF.nii.gz and <case>/aneurysms.nii.gz, where 2 marks treated aneurysms); other datasets only need their own scan and mask patterns. Each aneurysm becomes a two-point annotation, and Data/cases.csv maps the patients to the original cases.

uv run deepaneseg-prepare source_dir=ADAM data_dir=Data 'labels=[1]'
uv run deepaneseg-prepare source_dir=my_data data_dir=Data 'scan=scans/{case}.nii.gz' 'mask=masks/{case}.nii.gz'

Usage

1. Preprocessing: skull-strip the scans, then select the negative patch centers.

uv run deepaneseg-remove-skull data_dir=Data
uv run deepaneseg-extract-points data_dir=Data

2. Split: assign the patients to training, validation and testing (70/20/10 by default).

uv run deepaneseg-preprocess data_dir=Data               # 'split=[0.8,0.1,0.1]' overwrite=true

3. Training: outputs go to Data/0Work/exp1; running the same command again resumes the training.

uv run deepaneseg-train data_dir=Data name=exp1

4. Prediction: writes a probability map and the detections of each test patient.

uv run deepaneseg-predict train_dir=Data/0Work/exp1

Any other scan, or folder of scans, can be predicted the same way. Scans are skull-stripped first, like the training data.

uv run deepaneseg-predict train_dir=Data/0Work/exp1 input=scan.nii.gz output_dir=results

5. Evaluation: detection metrics of the test predictions, with the ADAM challenge criteria (a detection counts when its center lies within an aneurysm): sensitivity and false positives per patient.

uv run deepaneseg-evaluate train_dir=Data/0Work/exp1     # threshold=0.5 min_size=null

Configuration

Settings are Hydra configs in deepaneseg/configs/: one file per command, composed from groups (data, model, optimizer, scheduler, augmentation). Any value can be overridden on the command line:

uv run deepaneseg-train data_dir=Data name=exp2 batch_size=8 optimizer.lr=3e-4 model.f_maps=32
uv run deepaneseg-train --cfg job --resolve data_dir=Data name=exp2      # print the resulting config

To add an alternative, e.g. another scheduler, add a file with a _target_ class to the group, such as deepaneseg/configs/scheduler/step.yaml, and select it with scheduler=step.

Runs are reproducible: seed (0 by default) fixes the split and the training, and each training records its git commit and library versions in run_info.json. For identical GPU runs, also set deterministic=true (slower).

Outputs

Everything is written under Data/0Work:

split_pats.json                  train/valid/test split
logs/                            logs of the preprocessing commands
exp1/                            one training
  .hydra/config.yaml             resolved configuration, read back by deepaneseg-predict and deepaneseg-evaluate
  run_info.json                  git commit and library versions
  last_checkpoint.pytorch        model after the last epoch, used for prediction by default
  best_checkpoint.pytorch        model with the best validation metric
  logs/                          TensorBoard logs: uv run tensorboard --logdir Data/0Work
  train.log
  prediction/test/               for each test patient (prediction/scans/ for input=...):
    <patient>.nii.gz               probability map, on the patient's voxel grid
    <patient>_detections.csv       one row per detection: x, y, z and radius (mm), voxels, probability
    <patient>_detections.fcsv      the same detections as 3D Slicer markups
  evaluation/                    per_patient.csv and summary.json

Requirements

  • Memory: training keeps the training and validation scans in memory, about 4 bytes per voxel: 150 MiB for a 512×512×150 scan, around 16 GiB for 111 patients.
  • GPU: the U-Net has 19.1 M parameters (73 MiB). Its memory use comes mostly from the batch of 20 patches of 48³ voxels; if the GPU runs out of memory, lower batch_size and validation_batch_size.
  • Disk: predictions are float32 probability maps, one per scan.

Development

uv run pytest               # tests
uv run ruff check .         # lint
uv run ruff format .        # formatting

The code is organised as follows:

deepaneseg/
  cli/           commands
  configs/       Hydra configuration
  data/          patient data I/O, dataset conversion, splits, augmentation transforms
  volume/        patch extraction, sphere burning, point selection, skull stripping
  models/        3D U-Net
  training/      dataset, losses, metrics, trainer
  inference/     patch-wise prediction, detections and evaluation
  experimental/  post-paper models (vessel-coupled U-Nets) and their trainer
scripts/experimental/  analysis scripts from the PhD experiments

Differences from the published code

The Kappa metric/loss follows Cohen's definition and the focal loss is computed per voxel (Lin et al.), so trainings using them can give numbers that differ slightly from the published ones.

Citation

If you find this work useful, please cite the paper (also available from "Cite this repository" on GitHub):

@inproceedings{assis2021efficient,
  title={An efficient data strategy for the detection of brain aneurysms from MRA with deep learning},
  author={Assis, Youssef and Liao, Liang and Pierre, Fabien and Anxionnat, Ren{\'e} and Kerrien, Erwan},
  booktitle={Deep Generative Models, and Data Augmentation, Labelling, and Imperfections: First Workshop, DGM4MICCAI 2021, and First Workshop, DALI 2021, Held in Conjunction with MICCAI 2021, Strasbourg, France, October 1, 2021, Proceedings 1},
  pages={226--234},
  year={2021},
  organization={Springer}
}

Acknowledgements

This work was funded by the Grand-Est Region and the University Hospital (CHRU) of Nancy, France.

About

Efficient data strategy for brain aneurysm detection in 3D MRA (MICCAI-DALI 2021)

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Used by

Contributors

Languages