Research
DepthCropSeg++: Scaling a Crop Segmentation Foundation Model With Depth-Labeled Data
Overview Research area: Computer vision for agriculture — semantic segmentation of crops using foundation-model training techniques, depth-derived pseudo-labels, and transformer architectures. Technic

- arXiv
- 2601.12366
- Published
- 2026-01-18
- Authors
- Jiafei Zhang, Songliang Cao, Binghui Xu, Yanan Li, Weiwei Jia, Tingting Wu, Hao Lu, Weijuan Hu, Zhiguo Han
AI summary
Overview
Research area: Computer vision for agriculture — semantic segmentation of crops using foundation-model training techniques, depth-derived pseudo-labels, and transformer architectures.
Technical level: Advanced. The paper assumes familiarity with Vision Transformers, semantic segmentation metrics (mIoU, bIoU), self-training, and monocular depth estimation.
Scope: This paper describes DepthCropSeg++, a crop segmentation foundation model built on a ViT-Adapter backbone with FADE dynamic upsampling, trained on a 28,601-image cross-species, cross-scene dataset whose labels are generated almost entirely from depth maps rather than manual pixel annotation.
What This Paper Is About
Crop segmentation — separating plants from soil and background in field images — is normally solved with models trained on small, single-crop, manually annotated datasets, so performance collapses when the crop variety, lighting, or field condition changes. The authors extend their earlier near-unsupervised method DepthCropSeg by scaling the training data roughly an order of magnitude (from 3,577 to 28,406 images, as stated in the abstract) and upgrading the architecture, aiming for one model that segments arbitrary crop species in arbitrary field conditions.
Key Contributions
- DepthCropSeg++ itself: a crop segmentation foundation model with claimed cross-species and cross-scene capability, built on a ViT-Adapter backbone.
- A scaled-up depth-labeled dataset: the abstract reports 28,406 images spanning 30+ species and 15 environmental conditions, produced by depth-informed pseudo-labeling plus fast manual screening.
- Architectural upgrade: the base segmentation architecture of DepthCropSeg is replaced with the state-of-the-art ViT-Adapter, and its bilinear upsampling is replaced by the task-agnostic dynamic upsampling operator FADE for better boundary and fine-structure handling.
- A two-stage self-training pipeline that uses the first-stage model's predictions to filter pseudo-labels, excluding inconsistent pixels from the loss in the second stage.
Main Findings
- Headline accuracy: DepthCropSeg++ reaches 93.11% mIoU on a comprehensive test set of 6,760 images, exceeding the fully supervised baseline by +0.36% and SAM by +48.57% (abstract figures).
- Ablation on components (Table IV): semi-supervised baseline 92.36% mIoU; adding depth-informed pseudo-masks 92.86%; adding two-stage self-training 92.88%; adding FADE 93.11%. The authors note the self-training gain is small because performance is near saturation on this dataset.
- Comparison to other models (Table IV): SAM 32.80%, HQ-SAM 44.54%, the GWFSS competition top solution 69.50%, fully supervised model 92.75%.
- Upsampling comparison (Table V): bilinear 92.88% mIoU (570.59 M params, 2473 GFLOPs, 1.44 FPS, 7.38 G memory); CARAFE 92.93% (+0.35 M params, +2 GFLOPs); DySample 92.93% (+0.10 M, +0.3 GFLOPs); FADE 93.11% (+0.44 M, +3 GFLOPs). FADE surpasses bilinear by 0.23% in Table V, while the discussion describes an approximate 0.25% increase.
- Statistical comparison (Table VI): DepthCropSeg++ runs give 93.47, 93.43, 93.47 (mean ± std 93.457 ± 0.023, variance 0.0005); fully supervised runs give 93.02, 93.01, 93.07 (mean ± std 93.033 ± 0.032, variance 0.0010); 95% CI (0.362, 0.485), t-statistic 19.259, P-value 0.0001 (text states p-value less than 0.001).
- Challenging subsets (Table VII):
- Unseen soybean: DepthCropSeg++ 90.09% mIoU / 17.38% bIoU, versus SAM 29.98/1.16, HQ-SAM 35.15/4.17, GWFSS 42.35/12.28, DepthCropSeg 76.46/29.04, fully supervised 89.59/15.48. The paper reports bIoU improving by approximately 1.9% over the fully supervised model (17.38 vs 15.48).
- Nighttime rice: DepthCropSeg++ 86.90% mIoU, versus HQ-SAM 69.12, GWFSS 84.30, DepthCropSeg 62.21, fully supervised 86.67, SAM 49.04.
- Full-coverage canopy: DepthCropSeg++ 99.86% mIoU / 47.42% bIoU, versus fully supervised 96.54/20.36, SAM 72.32/12.89, HQ-SAM 78.97/15.83, DepthCropSeg 68.85/26.77, GWFSS 36.01/9.00.
- Labeling efficiency: DepthCropSeg screened 12,927 high-quality pseudo-labeled samples from 260,260 images in about 12 hours; the paper states comparable manual annotation would need a minimum of six months. Selecting full-coverage samples (3,980 training, 643 testing, all-ones masks) took 2 hours.
- Qualitative robustness: the model segmented crops correctly across cloudy soil fields, nighttime rice paddies with strong reflections, sunny high-contrast fields with shadows, densely planted crops, and pest/disease-damaged plants; it also handled low-light, unseen soybean, unseen ornamental plants, and different viewing angles (0°, 45°, 90°).
- Failure cases: poles mis-segmented as vegetation in high-coverage scenes; dried weeds partly classified as healthy vegetation; faint linear artifacts from repetitive row-wise soil patterns; fragmented output on certain complex morphologies; soil clumps segmented as vegetation under extremely low light; discontinuous segmentation in large-scale aerial views.
- Dataset totals: the final training set is reported as 28,601 images (11,406 supervised plus 17,195 pseudo-labeled) in Section II-E and Table III, while the abstract states 28,406 images; the test set is 6,760 images from 10 datasets, plus 72 nighttime rice and 47 manually labeled nighttime soybean images used for generalization assessment.
Methodology in Plain English
The authors avoid paying for pixel-level annotation by using depth. Their earlier DepthCropSeg system runs Depth Anything V2 on a single RGB crop photo to get a monocular depth map, then finds the crop boundary with a gradient-guided histogram thresholding procedure (depth normalization, edge detection, gradient-weighted enhancement, sigmoid curve fitting, then picking the threshold at the point of maximum gradient change in the fitted curve). Because only the relative depth gap between plant and ground matters, this works across lighting conditions, and Depth Anything V2 handles hard cases such as water reflections and shadows. Humans then quickly filter out bad masks.
They apply this to five unlabeled public datasets plus one manually annotated nighttime rice dataset, and add fully covered (all-plant, no background) images with trivial all-ones masks, since these defeat the depth pipeline. Preprocessing includes random scaling to 3584×896, cropping, 50% horizontal flipping, color perturbations for brightness/contrast/saturation/hue, dataset-specific normalization, and zero padding.
The segmentation network is BEiT-Adapter-Large (a ViT-Adapter variant) initialized from COCO-Stuff 164k weights trained for 80k iterations, fine-tuned with AdamW at a learning rate of 2e-5, weight decay 0.05, and layer-wise learning rate decay 0.9. All bilinear upsampling layers in the decoder are swapped for FADE, which generates upsampling kernels by adaptively gating encoder (shallow CNN) and decoder (transformer) features — favoring encoder features near object boundaries for sharp edges and decoder features inside objects for semantic consistency.
Training is two-stage: stage one trains on the initial pseudo-masks; the resulting model re-predicts on the whole training set; a trimap is formed by comparing predictions with the original masks, keeping consistent pixels and marking inconsistent ones as 255 so they are ignored in the loss; stage two retrains with this refined supervision. Experiments ran on four 48GB RTX A6000 GPUs with two 10-core Intel Xeon Silver 4210R CPUs and 256GB RAM, under Ubuntu 20.04, Python 3.8, PyTorch 1.9.0, and CUDA 11.8.
Why This Matters
Research impact: The work argues that data scale and label-generation method, not just architecture, drive generalization in agricultural vision. It also reports that a depth-label-trained model can beat the upper bound of fully supervised training on this test set (93.11% vs 92.75% mIoU), and that general-purpose foundation models such as SAM and HQ-SAM transfer poorly to agricultural imagery (32.80% and 44.54% mIoU).
Real-world applications:
- Plant phenotyping pipelines, where segmentation is the prerequisite for measuring traits.
- Density estimation and canopy/coverage measurement, including the Green fraction (GF), Green area index (GAI), and leaf angle distribution (LAD) mentioned for remote sensing of vegetation.
- Weed control and cover crop identification, where distinguishing crop from weed or background is required.
- Nighttime and high-throughput field monitoring, where the model reports 86.90% mIoU on nighttime rice and 99.86% on full-coverage canopies.
Industry relevance: The authors emphasize deployability on resource-constrained platforms such as UAVs and field robots, citing FADE's computational efficiency and independence from high-resolution guidance features; ViT-Adapter's lightweight adapters are likewise framed as suited to limited compute. The stated payoff is avoiding repeated data collection and re-annotation for each new crop, given growth cycles of three to ten months.
Future Directions
- Better structural cue extraction and adaptive scale modeling, which the authors propose directly in response to observed failures with poles, dried weeds, row-wise soil patterns, and aerial views.
- Reducing remaining long-tail failures under extreme low light, extreme scale compression, and complex plant morphologies, where segmentation fragments or misclassifies soil clumps.
- Closing the training-label gap, since stage two still depends on pseudo-labels and only excludes inconsistent pixels rather than correcting them; the paper notes two-stage self-training gave limited gains on this dataset.
- Evaluating generalization beyond agriculture-specific benchmarks, as the authors report that most publicly available labeled data were already absorbed into their training and test sets, leaving only qualitative visualizations for maize and wheat at multiple growth stages.
Target Audience
Researchers and engineers working on agricultural computer vision, plant phenotyping, and remote sensing who need segmentation that transfers across crop species and field conditions; practitioners deploying vision models on UAVs or field robots with limited compute; and machine learning researchers interested in weak supervision, pseudo-labeling from monocular depth, and self-training pipelines for domain-specific foundation models. Readers seeking a beginner-level introduction to segmentation or transformer architecture will find this paper demanding, since it presupposes familiarity with ViT-Adapter, dynamic upsampling operators such as CARAFE and IndexNet, and boundary-sensitive evaluation metrics.
Authors’ abstract
DepthCropSeg++: a foundation model for crop segmentation, capable of segmenting different crop species under open in-field environment. Crop segmentation is a fundamental task for modern agriculture, which closely relates to many downstream tasks such as plant phenotyping, density estimation, and weed control. In the era of foundation models, a number of generic large language and vision models have been developed. These models have demonstrated remarkable real world generalization due to significant model capacity and largescale datasets. However, current crop segmentation models mostly learn from limited data due to expensive pixel-level labelling cost, often performing well only under specific crop types or controlled environment. In this work, we follow the vein of our previous work DepthCropSeg, an almost unsupervised approach to crop segmentation, to scale up a cross-species and crossscene crop segmentation dataset, with 28,406 images across 30+ species and 15 environmental conditions. We also build upon a state-of-the-art semantic segmentation architecture ViT-Adapter architecture, enhance it with dynamic upsampling for improved detail awareness, and train the model with a two-stage selftraining pipeline. To systematically validate model performance, we conduct comprehensive experiments to justify the effectiveness and generalization capabilities across multiple crop datasets. Results demonstrate that DepthCropSeg++ achieves 93.11% mIoU on a comprehensive testing set, outperforming both supervised baselines and general-purpose vision foundation models like Segmentation Anything Model (SAM) by significant margins (+0.36% and +48.57% respectively). The model particularly excels in challenging scenarios including night-time environment (86.90% mIoU), high-density canopies (90.09% mIoU), and unseen crop varieties (90.09% mIoU), indicating a new state of the art for crop segmentation.