Deep Reinforcement Learning–Based Missile Guidance for Highly Maneuvering Targets Using Proximal Policy Optimization


Sazak M. D., Ahrazoglu M. A., Canpolat S¸Ahin M., Cansız B., Taşkıran M.

2026 11th International Conference on Recent Advances in Air and Space Technologies (RAST), İstanbul, Türkiye, 13 - 15 Mayıs 2026, ss.1-6, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/rast69551.2026.11672620
  • Basıldığı Şehir: İstanbul
  • Basıldığı Ülke: Türkiye
  • Sayfa Sayıları: ss.1-6
  • Yıldız Teknik Üniversitesi Adresli: Evet

Özet

Proportional Navigation Guidance (PNG) has long been widely used due to its practicality. Despite its simplicity, it is struggling to maintain consistent performance against unpredictable and aggressive target maneuvers. Its fixed-gain architecture creates challenges in adapting to varying maneuverability. This restriction requires a guidance framework that dynamically responds to engagement conditions. To address this shortcoming, the guidance problem is reframed as a continuous control reinforcement learning task and solved using Proximal Policy Optimization (PPO). The training environment is built around two-dimensional pursuit–evasion kinematics that incorporate realistic actuator dynamics, randomized initial conditions, and piecewise-varying target maneuvers. Comparative simulations demonstrate a clear performance advantage of the learned policy over fixed-gain PNG, achieving an interception rate of 96.77% against 91.03% for PNG across 836 randomized Monte Carlo scenarios, with an average miss distance of 17.83 m compared to 145.78 m. In particular, the policy also exhibits an emerging lofting behavior that was not explicitly encoded during training. Processor-in-the-loop experiments confirm successful interception across all tested scenarios within real-time bounds, establishing the proposed approach as a genuinely deployable alternative to conventional guidance methods.