Research
Towards Effective Waste Segmentation for Automated Waste Recycling in Cluttered Background
Overview Research area: Computer vision for automated waste recycling — specifically semantic segmentation of recyclable waste objects in cluttered scenes. Technical level: Intermediate (readers shoul

- arXiv
- 2606.13587
- Published
- 2026-06-11
- Authors
- Mamoona Javaid, Mubashir Noman, Abdul Hannan, Shah Nawaz, Mustansar Fiaz, Sajid Ghuffar
AI summary
Overview
- Research area: Computer vision for automated waste recycling — specifically semantic segmentation of recyclable waste objects in cluttered scenes.
- Technical level: Intermediate (readers should be comfortable with convolutional segmentation encoders/decoders, attention-style feature weighting, and frequency-domain filtering).
- Scope: The paper introduces EWSegNet, a waste segmentation network that combines spatial and spectral (frequency-domain) feature extraction with a boundary/blob enhancement module, and evaluates it on three waste segmentation datasets.
What This Paper Is About
Automated waste recycling needs segmentation models that can pick out translucent, deformable and thin waste objects from messy, cluttered backgrounds, but existing methods rely on large backbone networks that are computationally expensive and still lose accuracy in cluttered scenes. The paper's goal is to design a single network, EWSegNet, that improves segmentation in clutter while using fewer encoder parameters and lower latency than the current best-performing method.
Key Contributions
- An end-to-end waste segmentation network (EWSegNet) that aims to improve computational efficiency without compromising segmentation performance.
- A frequency context module (FCM) that captures global context using data-dependent kernels in the frequency domain.
- A spatial context module (SCM) that performs feature excitation and weighting in the spatial domain to complement the FCM.
- An auxiliary feature enhancement module (AFEM) that uses difference of Gaussian filtering and pooled attention to emphasize waste object boundaries and blob regions in cluttered scenes.
Main Findings
- Efficiency versus accuracy trade-off on ZeroWaste-f: EWSegNet reaches 56.44% mIoU and 91.75% pixel accuracy with 23.3M encoder parameters, 20.5 GFLOPs and 64.8 msec latency. COSNet, the stated prior state of the art, reaches 56.67% mIoU and 91.91% pixel accuracy with 27.3M encoder parameters, 24.4 GFLOPs and 73.6 msec latency. The paper describes this as similar mIoU with fairly reduced parameters and latency.
- Gains on a specific class: On ZeroWaste-f class-wise IoU, EWSegNet scores 35.05% on Metal versus 29.61% for COSNet, which the paper reports as an IoU gain of 5.44%. EWSegNet also scores higher on Background (91.45 vs 91.44) and Cardboard (59.24 vs 59.13), but lower than COSNet on Soft Plastic (63.17 vs 65.92) and Rigid Plastic (33.28 vs 37.24).
- Large improvement on ZeroWaste-aug: EWSegNet achieves 74.10% mIoU and 84.31% mF1, an absolute gain of 10.8% mIoU over LWCHNet (63.16% mIoU, 76.03% mF1). The authors attribute this to the augmented objects in the dataset reducing class imbalance.
- Best overall mIoU on SpectralWaste: EWSegNet reaches 71.03% mIoU versus 69.96% for COSNet. It scores higher on four object types, with the largest absolute gain being 4.63% on Cardboard (79.77 vs 75.14). COSNet remains better on Video Tape (42.95 vs 42.05) and Trash Bag (71.38 vs 68.96).
- Ablation results isolate each component: Starting from a baseline of 47.32% mIoU / 90.77% pixel accuracy on ZeroWaste-f, adding SCM alone gives 51.63% mIoU (a 4.31% gain), adding FCM alone gives 53.05% mIoU, combining FCM + SCM gives 54.11% mIoU (a further 1.06%), and adding AFEM brings the full model to 56.44% mIoU and 91.75% pixel accuracy.
- AFEM behaves as intended: Visualizations show that the boundary emphasis part highlights object boundaries and the blob amplification part emphasizes blob regions, and their combined output feeds forward as enhanced features.
- Hyperparameter sensitivity: Raising the initial learning rate from 5e-5 to 1e-4 improved ZeroWaste-f mIoU from 56.44% to 57.14%, with Metal IoU rising from 35.05% to 38.78% and Soft Plastic from 63.17% to 64.07%, while Cardboard and Rigid Plastic declined slightly.
- Documented failure cases: Qualitative results on ZeroWaste-f show all compared methods sometimes falsely segment paper as cardboard, and additional visualizations show EWSegNet and COSNet both produce false detections and miss some cardboard objects.
Methodology in Plain English
EWSegNet is built as an encoder–decoder segmentation network. The encoder has four stages, each containing a number of "efficient waste feature extraction" (EWFE) layers; the stage depths are (2, 2, 8, 2) and the channel dimensions are (80, 160, 320, 640). Convolution layers before each stage downsample the feature maps and increase channels, with the first acting as a stem layer.
Each EWFE layer contains three blocks. The spatial context module takes an input, projects it with a 5x5 group-wise convolution, splits it into three channel groups, and uses two of them to compute weights (via channel mean plus sigmoid, and spatial mean plus softmax) that re-weight the third group; the two weighted results are concatenated and passed through a 1x1 convolution. This captures relationships among nearby pixels. The frequency context module applies a 1x1 convolution to produce two feature maps, transforms both with the Fourier transform, multiplies them in the frequency domain (equivalent to convolution in the spatial domain), and transforms the result back with the inverse Fourier transform — an efficient way to capture global relationships. A multilayer perceptron follows.
The auxiliary feature enhancement module takes the third encoder stage's features and does two things. For boundary emphasis, it transforms the features to the frequency domain, multiplies them with two Gaussian functions with different sigma values, transforms both back, and subtracts one from the other (a difference of Gaussian operation) to isolate high-frequency boundary information; spatial-mean-derived channel weights then boost the relevant channels. For blob amplification, it generates query, key and value maps, applies average pooling to the queries and max pooling to the keys within an n×n neighborhood, and computes self-attention against the values. The two enhanced outputs are concatenated and passed through a 1x1 convolution. The enhanced features are added back to the third stage's feature maps before the fourth stage, and all multiscale features go to the decoder.
Training used MMSegmentation, a single Quadro RTX 6000 GPU, batch size 8, an ImageNet-1k pretrained encoder with a randomly initialized decoder, and the UPerNet decoder. Augmentations were random resize, random crop to 512x512, and random horizontal flip, with the AdamW optimizer, initial learning rate 5e-5, and 40k training iterations on all three datasets. At evaluation, the shorter image side was resized to 512 while preserving aspect ratio. The encoder itself was pretrained on ImageNet-1k for 600 epochs, reaching 81.7 Top-1 accuracy.
Why This Matters
Impact on research. The paper argues that prior waste segmentation methods inspired by spatial enhancement convolutions only capture a small neighborhood context and become expensive as filter sizes grow, and that lightweight transformer-based methods struggle in clutter. EWSegNet offers an alternative route — combining spatial excitation with frequency-domain global context — and reports efficiency gains alongside competitive accuracy, which is directly relevant to the reported scale of the waste problem (the paper cites an estimate of three billion tons of annual waste generation by 2050).
Real-world applications.
- Automated waste recycling (AWR) pipelines that separate recyclable items from mixed solid waste.
- Recycling plants that need to process material quickly without exposing workers to pointed and unhygienic waste objects.
- Robotic or automated sorting stations that must handle translucent, deformable and thin elongated items such as film, filaments and video tape.
- Household or municipal garbage detection and sorting systems of the kind targeted by earlier YOLO-based detection work cited in the paper.
Industry relevance. The paper reports encoder parameter counts, GFLOPs and latency on a 512x512 RGB input, framing efficiency as a practical requirement for AWR deployment rather than an afterthought. EWSegNet's 23.3M encoder parameters, 20.5 GFLOPs and 64.8 msec latency are lower than COSNet's 27.3M, 24.4 GFLOPs and 73.6 msec, which the paper presents as making the method more appropriate for deployment.
Future Directions
- Use hyperspectral data. SpectralWaste contains RGB and hyperspectral images, but this work used only the RGB version "for fair comparison"; the hyperspectral channel remains unexplored.
- Close the remaining accuracy gap. On ZeroWaste-f, EWSegNet's mIoU (56.44%) is slightly below COSNet's (56.67%), and COSNet still leads on the Video Tape and Trash Bag classes of SpectralWaste — suggesting headroom in those categories.
- Address documented failure modes. The additional visualizations show both EWSegNet and COSNet producing false detections and missing cardboard objects on ZeroWaste-f, which the method does not yet resolve.
- Tune and extend training settings. The learning-rate tuning in the appendix improved mIoU from 56.44% to 57.14%, indicating that the reported configuration is not necessarily optimal.
Target Audience
Researchers and engineers working on semantic segmentation, waste detection and sorting, and automated recycling systems will benefit most, particularly those interested in resource-efficient architectures for cluttered, non-rigid or translucent objects. The paper is also relevant to practitioners evaluating models for real-time or embedded deployment, since it reports parameter counts, GFLOPs and latency alongside accuracy, and to readers interested in frequency-domain feature design in deep networks.
Authors’ abstract
Rapid expansion of urban areas and population growth is causing an immense increase in waste production, which demands the need for efficient and automated waste management. In this scenario, automated waste recycling (AWR) using deep learning methods can assist humans in optimal waste management. Recent deep learning approaches for AWR provide promising waste segmentation performance, however, these methods rely on large backbone networks that are inefficient for AWR systems and suffer from performance deterioration in cluttered scenes. To this end, an optimal waste segmentation network is introduced which effectively utilizes the spatial domain to capture localized structural dependencies and the spectral domain to efficiently extract global contextual relationships. This cascaded design allows the network to progressively leverage both local and global representations across complementary domains to highlight the semantic information necessary for effective segmentation of various waste objects. Furthermore, auxiliary feature enhancement module (AFEM) is introduced to enhance the target objects' boundaries and blob amplification for better segmentation in cluttered scenarios. Extensive experimentation on ZeroWaste-aug, ZeroWaste-f and SpectralWaste datasets reveals the merits of the proposed method.