Phenotype-aware imitation learning for small, oriented object detection in remote sensing

Publication Type

Journal Article

Publication Date

3-2026

Abstract

Imitation learning has emerged as a promising paradigm for detecting small, arbitrarily oriented objects in remote sensing imagery. However, existing methods largely focus on global feature alignment between object instances and high-quality exemplars. This often leads to overhomogenized features across instances and distortion of intrinsic semantic structures, which in turn introduces geometric ambiguities in orientation, shape, and scale. Such issues are especially detrimental in remote sensing, where small objects exhibit significant variation and complex spatial distributions. To address these limitations, we propose a phenotype-aware imitation learning (PIL) framework that explicitly preserves geometric diversity and semantic structure for accurate object localization. PIL consists of two components: a geometric disentangling branch and a semantic structure registration branch. The geometric disentangling branch learns instance-specific deformation fields that adaptively normalize orientation, scale, and shape variations between exemplars and instances. By disentangling geometric differences, this branch allows the imitation process to focus on meaningful semantic discrepancies. The semantic structure registration branch further improves alignment by decoupling imitation into registration and imitation spaces. In the registration space, instances align their semantic structures by modeling spatial relations within the exemplar’s global context. In the imitation space, features that have been geometrically normalized and semantically aligned perform pixelwise imitation, supervised by a mixed-grained contrastive loss (MC-loss) that balances both local and global semantic cues. To mitigate geometric detail loss in cross-scale imitation, we also introduce a contextual micro-to-macro feature pyramid network (CM2-FPN). This module leverages low-layer spatial detail to complement the contextual information of large objects in high-layer features. Extensive experiments on DOTA, DIOR-R, and SODA-A datasets show that PIL-Net consistently outperforms existing methods. The performance gains are especially significant in challenging settings dominated by small, densely distributed, and arbitrarily oriented targets.

Keywords

Arbitrary-oriented small-object detection, feature pyramid network (FPN), imitation learning, remote sensing images

Discipline

Numerical Analysis and Scientific Computing | Remote Sensing

Publication

IEEE Transactions on Geoscience and Remote Sensing

Volume

64

First Page

1

Last Page

21

ISSN

0196-2892

Identifier

10.1109/TGRS.2026.3675321

Additional URL

https://doi.org/10.1109/TGRS.2026.3675321

This document is currently not available here.

Share

COinS