Research
SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
Overview Research area: Computer vision for forestry automation — specifically instance segmentation of tree trunks and fine-grained species classification from ground-level, under-canopy images in na

- arXiv
- 2510.09458
- Published
- 2025-10-10
- Authors
- David-Alexandre Duclos, William Guimont-Martin, Gabriel Jeanson, Arthur Larochelle-Tremblay, Martine Lapointe, Théo Defosse, Frédéric Moore, Philippe Nolet, François Pomerleau, Philippe Giguère
AI summary
Overview
Research area: Computer vision for forestry automation — specifically instance segmentation of tree trunks and fine-grained species classification from ground-level, under-canopy images in natural forests.
Technical level: Intermediate. The paper is readable without deep forestry or robotics background, but assumes familiarity with object detection, instance segmentation, and standard COCO-style metrics (mAP, mAR, AP50, AR50, IoU).
Scope in one sentence: The paper introduces SilvaScenes, a benchmark dataset of 164 under-canopy images containing 1421 annotated trees from 28 species across five bioclimatic domains in Quebec, Canada, and benchmarks modern deep learning models on it to establish how feasible joint tree detection and species classification actually is.
What This Paper Is About
Tree detection and taxonomic classification are considered core perception tasks for automating forestry operations such as field surveys and harvesting, and these operations often have to happen at ground level beneath a closed canopy. Existing image datasets for this setting either lack species labels entirely, focus on urban trees, or cover only a handful of visually distinct species — so it was unclear whether state-of-the-art models can simultaneously detect tree trunks and identify their species in real, high-diversity natural forests. The authors build a dataset that unifies pixel-precise trunk segmentation with expert species labels, and benchmark current architectures to measure how hard that combined task really is.
Key Contributions
-
SilvaScenes benchmark dataset: 164 colour images collected in June and July 2025 across five bioclimatic domains in Quebec, Canada, containing 1421 unique trees from 28 species, with instance segmentation masks for trunk detection and fine-grained species annotations produced with input from forestry experts.
-
A systematic evaluation of current deep learning approaches: CNN-based models (YOLOv11 and YOLOv12, in Small and X-Large variants) and a ViT-based model (Mask2Former with Swin-Small and Swin-Large backbones) are benchmarked on both species-aware and species-agnostic tree segmentation using stratified five-fold cross-validation.
-
A characterization of why the task is difficult: The authors isolate the effect of adding a classification component, analyze model confusion between visually similar species, and examine tree occlusion as a key failure factor.
-
Public release of dataset, source code, and models at https://github.com/norlab-ulaval/SilvaScenes.
Main Findings
-
Trunk segmentation is feasible; species-aware segmentation is not yet solved. The best model, Mask2Former with Swin-Large, reaches an mAP of 69.9% and an mAR of 76.4% for tree segmentation, but only an mAP of 39.2% and an mAR of 68.6% for species segmentation.
-
Adding species classification causes a large performance drop. The highest degradation appears in mAP and AP50, which suffer a 30.7-point and 39.6-point loss respectively when the classification component is included.
-
AP and AR diverge sharply on the species task. Because of high visual similarity between certain species, models appear to emit multiple predictions under ambiguity rather than withholding a prediction, producing high recall but low precision.
-
Bigger models do better, but no architecture dominates across all metrics. Mask2Former with Swin-Large is the clear winner when compute is unconstrained; Mask2Former with Swin-Small (68.8 M parameters) and YOLOv11 X-Large (56.9 M parameters) are similar in size yet each is stronger on a different task.
-
YOLOv11 X-Large outperforms YOLOv12 X-Large on most species metrics (mAP 35.2% versus 31.4%, AP50 48.0% versus 42.7%, mAR 64.1% versus 62.7%, AR50 85.9% versus 83.8%), suggesting attention mechanisms may not help on these tasks.
-
Tree detection is easy even for small models. Most architectures achieve comparable AP50 and AR50 for species-agnostic segmentation; Mask2Former with Swin-Large reaches AP50 of 90.8% and AR50 of 98.1%, while YOLOv11 Small still reaches 87.2% and 96.5%.
-
Speed and accuracy trade off substantially. On an NVIDIA RTX 4090 with BF16-mixed precision, Mask2Former with Swin-Large runs at 4.7 FPS (216.0 M parameters, 868.0 B FLOPs), while YOLOv11 Small runs at 57.7 FPS (9.4 M parameters, 35.5 B FLOPs).
-
Occlusion is pervasive. Nearly half of the trees are occluded by at least 25%, and almost one out of seven trees is occluded by 75% or more.
-
The dataset is highly imbalanced, which is typical of natural forests: sugar maple (243 trees) and balsam fir (250 trees) dominate, while species such as butternut, black ash, red ash, swamp white oak, eastern white pine, tamarack, pin cherry, and white elm each appear only once to three times.
-
Higher image resolutions produce significant performance gains and are identified in the abstract as likely fundamental to these tasks going forward; the detailed resolution results are in a section that falls outside the provided content.
-
Results compare favourably to prior work on detection quality. Against CanaTree100 (mAP 60.0%, AP50 87.2%, mAR 65.2%, AR50 91.5%), the authors attribute their gains to differences in network architecture, higher image quality, and annotation methodology.
Methodology in Plain English
Data collection. The team walked off-trail through forests rather than using a robot or drone, which gave them control over motion blur, camera angle, and camera settings. They used a Fujifilm GFX100S with a 43.8 × 32.9 mm, 102 MP sensor (11 648 × 8736 px) and a Fujifilm GF23mmF4 R LM WR lens with a 99.9° diagonal field of view, shooting around f/6.4 and 1/50 s to maximize depth of field and minimize noise. Because deep learning scales poorly with resolution, images were downsampled to 1.6 MP (1456 × 1092 px), matching prior work.
Site selection. Images were gathered across five Quebec bioclimatic domains, from the species-rich Sugar maple–Bitternut hickory domain (which alone contains 48 tree species in Quebec) to the conifer-dominated Balsam fir–White birch southern boreal domain, with multiple sites per domain to capture both inter- and intra-domain diversity.
Annotation. Species ground truth was largely established in situ by forestry experts who could use bark, leaves, shoots, cones, shapes, and environmental context. Masks covered trunks only — branches and foliage were excluded, because they are hard to annotate and not needed for operations such as harvesting. Occluded trunk sections were labelled when their shape could be inferred, except where another segmented trunk overlapped them. Trunks forking below breast height (1.3 m) were counted as separate trees, and very small trees (median width under 16 px in the downsampled images) were not annotated. Trees that could not be reliably identified were grouped as Unknown. The authors tried using the Segment Anything model family for automatic annotation but found the masks noisy, imprecise, and misaligned with their guidelines, so all masks were drawn and revised by humans using the full 102 MP images.
Benchmarking. All models were implemented in PyTorch, pre-trained on COCO, and trained with their native augmentation pipelines. Mask2Former's cross-entropy classification loss was replaced with focal loss to address class imbalance. Hyperparameters were tuned per experiment via Bayesian search with Weights & Biases. Experiments used stratified five-fold cross-validation, with folds split so that each held roughly 20% of every species' trees. To keep classes viable, species needed at least 16 specimens to be evaluated individually, leaving the 20 most common species; the eight species below that threshold were merged with Unknown into an "Other" class, giving 21 classes in total. Performance was measured with COCO metrics (mAP, AP50, mAR, AR50) macro-averaged across classes, plus parameter count, FLOPs, and FPS.
Why This Matters
Impact on research. The paper fills a concrete gap: previous under-canopy datasets either provide no class labels (ForTrunkDet, CanaTree100), cover only three genera (FinnWoodlands), or sidestep detection entirely by using close-up bark images of isolated trees (BarkNet 1.0, CentralBark). SilvaScenes is the first to combine dense instance masks with fine-grained species labels in high-diversity natural forests, giving the field a common yardstick and a quantified baseline showing how much room remains for improvement.
Real-world applications:
- Plot-level forest inventories: automating the species tally that surveyors currently perform on foot.
- Tree harvesting: giving forestry machinery the perception needed to identify and grasp the correct stems, echoing prior work on log grasping and segmentation under occlusion.
- Autonomous navigation in forests: trunk segmentation supports traversability assessment for under-canopy robots and UAVs.
- Ecosystem monitoring and mapping: tracking species distribution across bioclimatic domains, with the dataset's explicit bioclimatic design supporting generalization testing.
Industry relevance. The paper directly addresses forestry automation, which the authors frame around anticipated cost reductions, improved worker safety, and more sustainable practices. By reporting parameter counts, FLOPs, and FPS alongside accuracy, the work speaks to the practical trade-offs of deploying perception on low-compute mobile systems versus accepting the 4.7 FPS of the most accurate model.
Future Directions
-
Exploiting higher image resolution. The abstract reports that higher resolutions yield significant performance gains; determining how much of the species-classification gap can be closed by resolution alone — and how to afford the associated computational cost — is an open question.
-
Mitigating class imbalance and occlusion. The authors identify species imbalance and tree occlusion as among the most pressing issues for precise segmentation and identification, implying that sampling strategies, occlusion-aware architectures, or synthetic occlusion data are natural next steps.
-
Handling species ambiguity and confusion. The high-AR, low-AP pattern suggests models need better mechanisms for abstaining or disambiguating between visually similar species such as maples, red oak, ironwood, and aspens.
-
Adding depth and separate confidence estimation. The qualitative analysis suggests that depth images or depth estimation models could resolve cases where multiple trees are predicted as one, and that introducing a detection confidence independent of classification uncertainty could recover predictions currently discarded for low confidence.
-
Scaling annotation with foundation models. The authors note that with a larger amount of data, models such as SAM could offer a useful quality-versus-quantity trade-off despite their current imprecision on this task.
Target Audience
This paper is most useful to computer vision researchers working on instance segmentation in cluttered outdoor environments, robotics researchers developing perception for field or forestry robots, and forestry scientists interested in automating plot-level surveys and species inventories. It is also relevant to practitioners building deployable perception systems who need to weigh accuracy against FPS and parameter budgets. Readers should have some grounding in segmentation metrics and model architectures to get the most from the benchmark tables, though the problem framing and dataset description are accessible to a broader audience.
Authors’ abstract
Interest in forestry automation is growing alongside rapid advances in deep learning. In particular, tree detection and taxonomic classification are seen as core tasks required for automating field surveys and forestry equipment. These operations must often be performed in under-canopy settings, which pose challenging conditions for perception systems, including heavy occlusion, variable lighting, and dense vegetation. Despite this necessity, current work has yet to properly establish the feasibility of simultaneously executing tree detection and taxonomic classification in natural forests, as available datasets primarily focus on urban settings or on a limited number of species. To address this gap, we present SilvaScenes, a benchmark dataset for instance segmentation of tree species from under-canopy images in natural forests. Collected across five bioclimatic domains in Quebec, Canada, our dataset features 1421 trees from 28 species, with segmentation masks for pixel-precise tree trunk detection and fine-grained species annotations from forestry experts. We demonstrate the relevance and difficult nature of SilvaScenes by evaluating modern deep learning approaches, showing that while trunk segmentation is feasible, with a top mean average precision (mAP) of 69.9% and mean average recall (mAR) of 76.4%, species-aware segmentation remains a significant challenge with an mAP and an mAR of only 39.2% and 68.6%, respectively. Alongside additional experiments, we highlight key challenges, namely that species imbalance and tree occlusion figure among the most pressing issues for precise segmentation and identification. Meanwhile, higher image resolutions contribute to significant performance gains and will likely prove fundamental to these tasks moving forward. Our dataset, source code, and models will be made available at https://github.com/norlab-ulaval/SilvaScenes.