Skip to the content.

Deep Reinforcement Learning (DRL)

DRL combines neural networks (function approximation) with reinforcement learning (trial-and-error learning via rewards) for sequential decision-making.

Agent-Environment Loop

Agent (policy π) ──action aₜ──▶ Environment
       ▲                              │
       └──── state sₜ₊₁, reward rₜ ──┘
Component Description Telecom Example
State (s) Current observation Network KPIs, topology, alarms
Action (a) Agent’s decision Reroute traffic, restart VNF, adjust power
Reward (r) Feedback signal +1 SLA met, -1 degradation
Policy (π) State → action mapping The neural network

Key Algorithms

Algorithm Type Best For
DQN Value-based Discrete action spaces
PPO Actor-Critic General purpose, most popular
SAC Actor-Critic Continuous control, robust exploration
DDPG / TD3 Actor-Critic Continuous action output
A3C / A2C Actor-Critic Distributed/parallel training

Why GNN + DRL

Standard DRL treats state as a flat vector, losing topology structure. GNN + DRL preserves the graph:

This combination is the de facto standard for telecom optimization problems (O-RAN xApps, resource allocation, power control).