Research
Tracking spatial temporal details in ultrasound long video via wavelet analysis and memory bank
Overview Research area: Computer vision for medical image analysis, specifically spatiotemporal segmentation of long ultrasound videos. Technical level: Advanced. The paper assumes familiarity with en

- arXiv
- 2512.15066
- Published
- 2025-12-17
- Authors
- Chenxiao Zhang, Runshi Zhang, Junchen Wang
AI summary
Overview
Research area: Computer vision for medical image analysis, specifically spatiotemporal segmentation of long ultrasound videos.
Technical level: Advanced. The paper assumes familiarity with encoder-decoder segmentation networks, wavelet transforms (including the lifting scheme and inverse transforms), ConvLSTM, cross-attention memory banks, and channel attention.
Scope: The paper proposes MWNet, a wavelet- and memory-bank-based encoder-decoder network that segments organs and lesions in long ultrasound videos, and evaluates it on four ultrasound video datasets against eleven comparison methods.
What This Paper Is About
Ultrasound video is widely used for screening, diagnosis, and surgical planning, and automatically outlining the target tissue (segmentation) is a key step in computer-assisted surgery. Ultrasound footage is hard to segment automatically because of low contrast, blurred boundaries, confusing look-alike regions, and speckle noise, and these problems get worse for small objects in long video sequences where short-term memory mechanisms lose track of the target.
The paper's goal is a single end-to-end network that keeps fine-grained boundary detail (which the authors associate with high-frequency components of feature maps) while also modeling long-range temporal dependencies across many frames.
Key Contributions
-
Memory-based wavelet convolution (MWConv) backbone. A hierarchical encoder that extracts multiscale spatial and temporal features in the wavelet domain. It uses cascaded wavelet compression to fuse multiscale frequency-domain features inside each convolutional layer, and ConvLSTM to add high-frequency information from adjacent frames.
-
Long short-term memory bank. A memory mechanism using cross-attention and memory compression that combines temporal correlations between both nearby and distant frames and the current frame, to support tracking objects over long videos.
-
High-frequency-aware feature fusion (HFF) module. A decoder module using adaptive wavelet filters that attends to different frequency components in multiscale feature maps, suppresses noise, and highlights high-frequency components.
-
Extensive benchmarking. Evaluation on four ultrasound video datasets covering two thyroid nodule datasets, a thyroid gland dataset, and a heart dataset, with the code released at https://github.com/XiAooZ/MWNet.
Main Findings
-
Best overall segmentation metrics claimed: The paper reports that MWNet achieved the highest DSC and IoU values and the lowest MAE across all four datasets compared with the state-of-the-art methods. The specific MWNet values are not present in the provided content, which truncates Tables 1-4.
-
Comparison methods covered image and video approaches: Six image methods (UNet, Swin-UNet, SETR, SegFormer, SegNeXt, Mask2Former) and five video methods (FLA-Net, Vivim, LGR-Net, OneVOS, MemSAM).
-
Only partial comparison numbers are visible in the provided content: On the Thyroid Nodule test set, UNet scored 75.88 DSC, 61.13 IoU, MAE 0.0184, with 28.99M parameters and 77 FPS, while Swin-UNet scored 75.73 DSC, 60.94 IoU, MAE 0.0208, with 27.17M parameters and 42 FPS. The table content stops there, so the remaining rows are not reported in the provided text.
-
Small thyroid nodules were segmented more accurately: The authors state their method can more accurately segment small thyroid nodules, demonstrating effectiveness for small ultrasound objects in long video.
-
Qualitative improvement at boundaries: On the Thyroid Nodule and VTUS datasets, predicted masks matched the ground truth better and boundary segmentation accuracy improved notably, which the authors link to handling confusing locations, speckle noise, and blurred boundaries.
-
Diagnosed failure modes of competing methods: U-Net and FLA-Net captured insufficient global context; LGR-Net extracted insufficient high-frequency detail, hurting small-object boundaries; SETR and Vivim suffered from insufficient feature fusion, degrading blurred boundaries; SegNeXt used only the last three feature-map layers and overlooked low-level information.
-
Evaluation metrics used: Dice similarity coefficient (DSC), intersection over union (IoU), mean absolute error (MAE), precision, and recall.
Methodology in Plain English
The network has three parts: an encoder, a memory bank, and a decoder.
Encoder. It extends ConvNeXt with three hierarchical stages. In each stage, the input feature maps go through a wavelet convolution block: the maps are decomposed with a cascaded wavelet transform, which repeatedly splits the low-frequency component to build multiresolution representations (recursion continues until the lowest resolution equals 1/4 of the encoder's deepest resolution). Small-kernel convolutions are applied to the downsampled maps, and an inverse wavelet transform progressively upsamples and recombines them, adding in the low-frequency component from the higher resolution at each scale. This enlarges the effective receptive field without the memory cost of large-kernel convolutions. A temporal fusion module based on ConvLSTM then models dependencies between the current frame and its preceding frames at each stage.
Memory bank. At the deepest stage, a long short-term memory bank relates the current frame's features to both long-term and short-term memory via cross-attention (query from the current feature, key and value from the memory). Memory bank lengths are restricted to bound GPU memory use. Short-term memory drops the earliest features when it exceeds its length; long-term memory uses a similarity-driven update that discards highly redundant features on insertion, so early-stage information is retained in compressed form.
Decoder. A series of HF-aware feature fusion modules combines each low-resolution decoder feature with the skip connection from the corresponding encoder layer. The module uses an adaptive wavelet transform to split features into low- and high-frequency subbands, reweights subbands with a squeeze-and-excitation channel attention block (favoring high-frequency subbands for the high-pass filter and low-frequency subbands for the low-pass filter), and reconstructs with an inverse adaptive wavelet transform. Two pairs of these filters are applied, with the results combined by pointwise convolution.
Adaptive wavelets. The classical lifting scheme splits a 1D signal into even and odd samples, predicts the odd samples from the even ones to form high-frequency detail coefficients, and updates the even samples to form a low-pass approximation. The paper replaces the fixed predictor and updater functions with learnable one-dimensional convolutions, making the transform data-driven, and extends it to two dimensions by applying it horizontally then vertically.
Training setup. AdamW optimizer, learning rate 10⁻⁴, PolyLR scheduler, LinearLR warmup, layer-wise learning rate decay, batch size 2, 180,000 iterations, weighted Dice plus binary cross-entropy loss. The WTConv backbone was pretrained on ImageNet-1k for over 300 epochs. Videos were temporally cropped from randomized initial frames into 10-frame sections, with identical augmentations (random blur, random flip, color jitter) applied within a section. Input images were resized to 512×512. Implementation used PyTorch and mmsegmentation on a single NVIDIA GeForce RTX 4090 GPU. Speed is reported at batch size 1 as milliseconds per frame; inference time is single-model, single-scale, without post-processing.
Datasets. The Thyroid Nodule dataset has 64 videos totaling 5611 pixel-annotated frames, resolutions from 333×510 to 523×771, split 40/10/14 into training, validation, and test. VTUS has 100 thyroid ultrasound videos totaling 9342 frames with pixel-wise ground truth from transverse and longitudinal B-mode scans captured with Mindray Resona8 and TOSHIBA Aplio500 vendors, cross-annotated by three experts with more than three years of thyroid diagnosis experience, with videos ranging from 31 to 196 frames and a 7:3 split giving 70 training and 30 testing videos. TG3K has 16 ultrasound thyroid videos with pixel-wise gland annotations; following Gong et al., sequences where the thyroid gland occupies more than 6% of the image were selected, yielding 3585 frames from 16 sequences at resolutions from 277×284 to 316×333, with 12 training, 2 validation, and 2 test sequences. CAMUS contains echocardiography data from 500 patients with pixel-wise left ventricle annotations, acquired at the University Hospital of St Etienne (France), comprising 1000 standardized sequences (500 apical 2-chamber and 500 apical 4-chamber views) split into 700 training, 100 validation, and 200 test.
Why This Matters
Impact on research. The paper pushes back on the common practice of inserting wavelet modules as plug-and-play components into otherwise conventional encoders and decoders. It argues that the decoder's feature fusion stage receives too little attention in existing wavelet-based ultrasound work, and it instead builds wavelet-based layers into both the encoder and the decoder, with fixed-basis wavelets in the encoder (for receptive field expansion and multiscale features) and adaptive wavelets in the decoder (for adaptive low/high-pass filtering). It also targets a known scaling weakness of long-term memory banks: naive accumulation of all frames grows memory consumption with sequence length, which the restricted memory banks with similarity-driven updates are designed to avoid.
Real-world applications.
- Screening and diagnosis of thyroid nodules from ultrasound video, including small nodules that are easy to miss.
- Assessment of the thyroid gland.
- Cardiac ultrasound analysis, specifically left ventricle segmentation in echocardiography.
- Surgical planning and intraoperative/postoperative observation, where segmentation feeds into the computer-assisted surgery workflow.
Industry relevance. Automatic segmentation reduces the time and error associated with manual labeling of diseased areas, which the authors note is time-consuming and susceptible to mistakes. The work is supported by the National Natural Science Foundation of China (Grants 62173014 and U22A2051), the National Key Research and Development Program of China (Grant 2022YFC2405401), and the Natural Science Foundation of Beijing Municipality (Grants L232037 and L242166), indicating institutional investment in this clinical direction.
Future Directions
- Reported evaluation is incomplete in the provided content: the specific MWNet numbers for DSC, IoU, MAE, parameter count, and FPS across all four datasets are not shown, so the magnitude of the claimed improvements cannot be verified from what is presented here.
- Memory bank scaling. The paper identifies that simplistic accumulation of long-term memory features rapidly increases memory consumption with longer sequences; how well the restricted memory banks and similarity-driven updates scale to substantially longer videos than the 31-to-196-frame range reported for VTUS is an open question.
- Extension beyond the four evaluated modalities. The method was tested on thyroid nodule, thyroid gland, and cardiac data; generalization to other ultrasound targets with different frequency characteristics is untested.
- Annotation dependence. The paper positions its memory design as avoiding annotation supervision, in contrast to methods such as confidence-based memory updates that depend on keyframe annotations; whether this holds under low-annotation or unannotated clinical deployment settings remains to be shown.
Target Audience
Researchers and graduate students working on medical image and video segmentation, particularly those interested in frequency-domain methods, wavelet-based network design, and memory-bank temporal modeling. It is also relevant to clinical engineers and medical imaging practitioners who need automated organ and lesion delineation from ultrasound video, and to readers specifically interested in small-object tracking in long sequences. The paper is not a beginner-level read: comfort with wavelet decompositions, lifting schemes, ConvLSTM, and attention mechanisms is needed to follow the methodology.
Authors’ abstract
Medical ultrasound videos are widely used for medical inspections, disease diagnosis and surgical planning. High-fidelity lesion area and target organ segmentation constitutes a key component of the computer-assisted surgery workflow. The low contrast levels and noisy backgrounds of ultrasound videos cause missegmentation of organ boundary, which may lead to small object losses and increase boundary segmentation errors. Object tracking in long videos also remains a significant research challenge. To overcome these challenges, we propose a memory bank-based wavelet filtering and fusion network, which adopts an encoder-decoder structure to effectively extract fine-grained detailed spatial features and integrate high-frequency (HF) information. Specifically, memory-based wavelet convolution is presented to simultaneously capture category, detailed information and utilize adjacent information in the encoder. Cascaded wavelet compression is used to fuse multiscale frequency-domain features and expand the receptive field within each convolutional layer. A long short-term memory bank using cross-attention and memory compression mechanisms is designed to track objects in long video. To fully utilize the boundary-sensitive HF details of feature maps, an HF-aware feature fusion module is designed via adaptive wavelet filters in the decoder. In extensive benchmark tests conducted on four ultrasound video datasets (two thyroid nodule, the thyroid gland, the heart datasets) compared with the state-of-the-art methods, our method demonstrates marked improvements in segmentation metrics. In particular, our method can more accurately segment small thyroid nodules, demonstrating its effectiveness for cases involving small ultrasound objects in long video. The code is available at https://github.com/XiAooZ/MWNet.