Intelligent Coordinated Scheduling for Battery EV Car Charging and Battery Swapping Stations: A Hierarchical Deep Reinforcement Learning Approach

The rapid growth of the battery EV car industry has placed unprecedented demands on energy supply infrastructure. Integrated charging and battery swapping stations (ICBSS) represent a promising solution to alleviate range anxiety and improve service efficiency. However, their practical operation is fraught with multifaceted challenges. Operators grapple with low service efficiency during peak hours, suboptimal economic returns due to simplistic pricing and resource allocation, and a weak ability to interact with the power grid, missing opportunities for demand response and ancillary services. Vehicle-to-Grid (V2G) technology presents a paradigm-shifting opportunity, enabling bidirectional energy flow where a battery EV car can not only draw power from the grid but also feed stored energy back to it. Effectively harnessing this potential within the complex, dynamic environment of an ICBSS requires an intelligent, adaptive decision-making core.

Traditional optimization methods, such as linear programming or heuristic algorithms like Particle Swarm Optimization (PSO) and Genetic Algorithms (GA), often struggle with the high-dimensional state space, stochastic arrivals of battery EV cars, and real-time decision-making requirements. While Deep Reinforcement Learning (DRL) has emerged as a powerful tool for sequential decision-making, standard monolithic DRL frameworks can be overwhelmed by the “curse of dimensionality” in this setting, where decisions span multiple timescales—from strategic mode selection (e.g., prioritizing profit vs. service) to tactical, continuous power adjustments for dozens of charging points and batteries.

To address these limitations, we propose a novel V2G-coordinated intelligent scheduling model based on a Hierarchical Deep Reinforcement Learning (HDRL) framework. Our core innovation lies in decomposing the complex scheduling problem into a strategic layer and a tactical layer, effectively reducing decision complexity and enhancing system responsiveness. The strategic layer operates on a longer timescale, making high-level mode decisions based on grid electricity prices, station status, and overall objectives. The tactical layer, conditioned on the strategic directive, executes fine-grained, continuous control of charging/discharging power for each battery EV car connected to fast/slow chargers and manages discrete discharge levels for swappable batteries. This hierarchical decomposition is crucial for managing complexity. Furthermore, we design a hybrid action space to seamlessly handle both discrete mode choices and continuous power regulation. For learning within this architecture, we employ the Soft Actor-Critic (SAC) algorithm at its core, favored for its stability, efficient exploration via entropy regularization, and native capability to handle continuous actions, which we extend to our hybrid setting.

This study formulates the ICBSS scheduling problem as a Markov Decision Process (MDP) and develops a comprehensive simulation environment to validate our approach. We demonstrate that our HDRL model significantly improves operational profitability by capitalizing on time-of-use (TOU) price arbitrage through intelligent V2G dispatch, enhances service quality by dynamically managing queues for different battery EV car service types, and improves grid interaction by providing a predictable and schedulable resource. The model shows strong robustness against fluctuations in battery EV car traffic and electricity prices, offering a practical, intelligent scheduling solution for the sustainable development of the battery EV car ecosystem.

1. System Framework and Hierarchical Decision-Making Methodology

The physical architecture of our envisioned smart ICBSS integrates three primary functions: conductive charging (both fast and slow), battery swapping, and V2G-enabled bidirectional energy transfer. The station is equipped with \(x_1\) fast charging points (e.g., 50-100 kW), \(x_2\) slow charging points (e.g., ≤22 kW), and a inventory of \(x_3\) swappable battery packs. A central Energy Management System (EMS) serves as the brain of the station, coordinating all resources. When a battery EV car arrives, the EMS assesses its request (charging or swapping), current state of charge (SOC), and expected dwell time, and then assigns it to an appropriate resource. For swapping, a depleted battery is replaced with a fully charged one from the inventory; the depleted battery then joins a charging queue. For direct charging, the battery EV car is connected to a fast or slow charger. Crucially, any connected battery EV car (whether being charged or a battery in the swapping inventory) can be scheduled to discharge energy back to the station’s local bus under V2G mode, which can then be used to supply other loads or be sold back to the grid, forming a controllable microgrid.

The core intelligence is implemented via a Hierarchical Deep Reinforcement Learning (HDRL) framework, which effectively models the sequential decision-making problem as a Markov Decision Process (MDP) defined by the tuple \((S, A, P, R, \gamma)\).

State Space (S): The state \(s_t \in S\) at time \(t\) encapsulates all necessary information for decision-making:
$$ s_t = [B_t, Q_t, P_t, O_t, T_t] $$
where:

  • \(B_t\): Number of fully charged swappable batteries (normalized).
  • \(Q_t\): Queue state vector, encoding the length of swap, fast-charge, and slow-charge queues (e.g., one-hot encoded as low/medium/high).
  • \(P_t\): Current electricity price signal (purchase/sell).
  • \(O_t\): Occupancy vector for all charging points (0 for idle, 1 for occupied).
  • \(T_t\): Temporal features (time of day, day of week, normalized).

Hybrid Action Space (A): The action \(a_t \in A\) is a hybrid construct output by our hierarchical agent.

  1. Strategic Layer (Discrete Mode Selection \(a_t^{str}\)): This layer chooses an operational mode that sets the high-level objective for a period.
    • Mode 0 (Economic): Prioritizes profit maximization via V2G arbitrage. Favors discharging when prices are high, possibly at the expense of slower charging for some battery EV cars.
    • Mode 1 (Balanced): Aims to strike a balance between economic gain and service quality.
    • Mode 2 (Service-Priority): Prioritizes minimizing wait times and ensuring fast charging for all battery EV cars, limiting V2G discharge.

    The strategic decision is triggered periodically or by significant state changes (e.g., large price shift).

  2. Tactical Layer (Continuous/Discrete Control \(a_t^{tac}\)): Conditioned on the strategic mode \(a_t^{str}\), this layer executes fine-grained control every decision interval (e.g., every 5-15 minutes).
    • For each fast/slow charger \(i\): Outputs a continuous power adjustment factor \(\alpha_i \in [-1, 1]\). The actual power \(P_{act,i}\) is \( \alpha_i \cdot P_{max,i} \), where negative values indicate discharging the connected battery EV car.
    • For each swappable battery \(j\) in the inventory: Outputs a discrete discharge level \(d_j \in \{0, 1, 2\}\), corresponding to zero, medium, or high discharge power.

The final action is the combination: \(a_t = (a_t^{str}, a_t^{tac})\).

Reward Function (R): The reward \(r_t\) is designed to guide the agent towards economically efficient, service-oriented, and grid-friendly operation. It is a weighted sum of revenue components and penalty terms:
$$ r_t = \gamma_1 R_{revenue}(t) + \gamma_2 R_{penalty}(t) $$
The revenue term is:
$$ R_{revenue}(t) = \gamma_c R_{charge}(t) + \gamma_s R_{swap}(t) + \gamma_d R_{discharge}(t) $$
where:
$$ R_{charge}(t) = \sum_{i \in \{f,s\}} \sum_{j} E_{i,j}^{ch}(t) \cdot c_{service}^{charge} $$
$$ R_{swap}(t) = \sum_{k} (C – C \cdot SOC_{init}^k) \cdot c_{service}^{swap} $$
$$ R_{discharge}(t) = \sum_{i \in \{f,s\}} \sum_{j} E_{i,j}^{dis}(t) \cdot \left( p_{sell}(t) – p_{buy}^{min} \right) + \sum_{b} E_{b}^{dis}(t) \cdot \left( p_{sell}(t) – p_{buy}^{min} \right) $$
Here, \(E^{ch/dis}\) denotes charged/discharged energy, \(c_{service}\) is the service fee, \(p_{sell}(t)\) is the real-time selling price, and \(p_{buy}^{min}\) is a baseline purchase cost to ensure minimum arbitrage profit. The penalty term \(R_{penalty}(t)\) includes costs for excessive queue lengths, battery shortage for swap, and battery health degradation from aggressive V2G cycling, encouraging sustainable operation for every battery EV car.

Transition Function (P) and Learning with SAC: The state transition probability \(P(s_{t+1}|s_t, a_t)\) models the stochastic arrival of battery EV cars, service completion, and battery charge/discharge dynamics. We utilize the SAC algorithm due to its off-policy nature, sample efficiency, and stability. The tactical layer’s actor network in SAC is modified to output parameters for a mixed distribution: a categorical distribution for the strategic mode (when triggered) and multivariate Gaussian distributions for the continuous power adjustments and logits for the discrete battery discharge levels. The entropy regularization term in SAC’s objective encourages exploration, which is vital for discovering optimal hybrid strategies in this complex environment.

2. Modeling the Integrated Charging and Swapping Station Operation

Accurate models of battery EV car arrival patterns and station economics are fundamental for training a realistic and robust DRL agent.

Battery EV Car Arrival and Service Model: The arrival process of battery EV cars is non-stationary and varies significantly with time of day. We model it as a time-inhomogeneous Poisson process:
$$ N(t) \sim \text{Poisson}(\lambda(t)) $$
where \(\lambda(t)\) is the average arrival rate during time period \(t\) (e.g., morning rush hour, midday, evening). Each arriving battery EV car is characterized by its request type (fast-charge, slow-charge, swap), its initial SOC (\(SOC_{init}\)), and its desired departure SOC (\(SOC_{target}\)) or required swap. The required energy for a charging request is:
$$ E_{req} = C \cdot (SOC_{target} – SOC_{init}) $$
where \(C\) is the battery capacity of the battery EV car. The theoretical minimum service time on a charger with power \(P_c\) is \(t_{charge} = E_{req} / P_c\), though this can be extended or interrupted by V2G调度.

Economic and Operational Model: The station’s daily operational profit is the sum of revenues minus the cost of energy purchased from the grid. The key revenue streams are:

  1. Charging Service Revenue: Income from providing conductive charging to a battery EV car, based on energy delivered and a service fee.
  2. Battery Swap Service Revenue: Income from providing a full battery swap to a battery EV car, typically based on the energy differential and a service fee.
  3. V2G Discharge Revenue: Income from selling energy stored in battery EV car batteries (from either charging points or the swap inventory) back to the grid during high-price periods.

The core financial logic for V2G is to buy and store energy when prices \(p_{buy}(t)\) are low (e.g., at night) and sell it when \(p_{sell}(t)\) is high (e.g., evening peak), ensuring \(p_{sell}(t) > p_{buy}^{min}\) for profitability. The optimization must also account for battery degradation costs associated with additional charge-discharge cycles.

Time-of-Use (TOU) Price Schema: A typical TOU price structure used in our simulation is summarized below. The intelligent scheduler must learn to navigate this landscape.

Time Period Electricity Purchase Price ($/kWh) Charging Service Fee ($/kWh) Swap Service Fee ($/kWh)
00:00 – 06:00 (Off-Peak) 0.31 0.32 0.29
06:00 – 08:00 (Shoulder) 0.65 0.25 0.51
08:00 – 11:00 (On-Peak) 1.12 0.12 0.72
11:00 – 18:00 (Shoulder) 0.65 0.25 0.47
18:00 – 19:00 (On-Peak) 1.11 0.12 0.60
19:00 – 21:00 (Peak) 1.37 0.22 0.60
21:00 – 22:00 (Shoulder) 0.65 0.25 0.51
22:00 – 24:00 (Off-Peak) 0.31 0.32 0.29

3. Experimental Analysis and Performance Validation

We constructed a simulation environment based on Python to train and evaluate our HDRL model. A station with 10 fast chargers (50-100kW), 10 slow chargers (≤22kW), and 10 swappable batteries was simulated. Battery EV car arrivals followed the time-inhomogeneous Poisson process, with a mix of 40% swap requests, 40% fast-charge requests, and 20% slow-charge requests during peak hours. The SAC networks were implemented using PyTorch. We analyzed the model’s performance from several critical angles.

3.1 Convergence and Reward Composition Analysis

The training curve shows the agent’s total episodic reward increasing and converging over approximately 400-450 training episodes, indicating successful learning. The breakdown of the final reward composition reveals the model’s learned operational strategy. In a stable, converged policy, the revenue distribution typically settles around: Charging Revenue (~30%), Swap Revenue (~38%), Discharge Revenue (~32%). This balanced split indicates that the agent does not overly favor one revenue stream at the expense of others. It leverages V2G discharge significantly but not excessively, ensuring that core services for the battery EV car driver are not compromised. The penalty component remains low, showing the model avoids severe queue build-ups or battery stockouts.

3.2 Strategic Mode Selection and Tactical Control

The hierarchical design’s effectiveness is evident in the agent’s behavior. The strategic layer reliably switches modes in response to external conditions. During high-price periods (e.g., 19:00-21:00), the “Economic” mode is frequently activated, authorizing the tactical layer to perform aggressive V2G discharge from both connected battery EV cars and spare batteries. During shoulder periods with moderate queues, the “Balanced” mode prevails. When queues grow long, especially for swapping which is a quick service, the “Service-Priority” mode is engaged, temporarily curtailing V2G to free up power and capacity for servicing waiting battery EV cars. The tactical layer demonstrates nuanced control, applying different strategies to fast-charging versus slow-charging battery EV cars. For example, a fast-charging battery EV car with a high SOC might be discharged heavily during a price peak and then quickly recharged before departure. A slow-charging battery EV car parked for hours might undergo multiple smaller charge-discharge cycles to accumulate arbitrage profit.

3.3 Robustness and Scalability

We tested the model’s robustness under varying station scales. As shown in the table below, key performance metrics remain stable when the number of chargers is scaled up, demonstrating the policy’s generalizability. The energy efficiency ratio (discharge energy/charge energy) stabilizes around 48-49%, indicating a consistent and sustainable V2G utilization level across scales.

Number of Chargers Energy Efficiency Ratio Average Service Success Rate* Convergence Speed (Episodes)
20 (10F+10S) 48.5% 94.2% 420
30 (15F+15S) 49.1% 93.8% 430
40 (20F+20S) 47.8% 94.5% 410
*Service Success: Battery EV car served (swap completed or charge target met) before intended departure.

3.4 Comparative Analysis with Baseline Methods

We compared our HDRL-SAC model against several benchmarks: a rule-based heuristic (RBC), Particle Swarm Optimization (PSO) applied in a receding horizon manner, a standard DQN agent (with discretized actions), and a monolithic (non-hierarchical) SAC agent attempting to handle the full hybrid action space. The comparison focused on daily operational profit and average battery EV car waiting time.

Scheduling Method Average Daily Profit ($) Avg. Wait Time (minutes) Key Limitation
HDRL-SAC (Our Model) 2780 8.5 Requires offline training
Monolithic SAC 2420 11.2 Unstable training, slower convergence
DQN with Discretization 2100 14.7 Coarse control, poor V2G granularity
PSO (Receding Horizon) 2550 9.8 High computational delay (~250ms/step)
Rule-Based Controller (RBC) 1850 18.3 Inflexible, cannot adapt to patterns

Our HDRL-SAC model achieves the highest profit by intelligently exploiting V2G arbitrage while maintaining the lowest average wait time. The hierarchical structure is key to this performance, as evidenced by the monolithic SAC’s inferior results. PSO performs well but is computationally expensive for real-time control. DQN suffers from the inherent limitations of discretizing continuous power control. The ability of our model to reduce wait time by over 50% compared to the RBC significantly enhances the experience for the driver of every battery EV car.

3.5 Analysis of Hybrid Action Space Efficacy

A crucial ablation study involved comparing our hybrid action space design against purely discrete or purely continuous alternatives within the same HDRL framework. A purely discrete space forces power levels into a few fixed steps, crippling fine-grained V2G optimization. A purely continuous space struggles with the categorical decision of operational mode, often leading to indecisive or oscillating high-level strategies. Our hybrid design, as shown below, provides the best trade-off, enabling both decisive strategic shifts and precise tactical control.

Action Space Design Convergence Speed Policy Stability (Reward Std. Dev.) Peak Profit Attainment
Hybrid (Ours) 420 episodes Low (5.2%) High
Purely Discrete 580 episodes Medium (12.3%) Medium
Purely Continuous 510 episodes Medium-High (8.7%) Medium-High

4. Conclusion and Implications

This study presents a novel Hierarchical Deep Reinforcement Learning framework for the intelligent, coordinated scheduling of integrated battery EV car charging and battery swapping stations with V2G capability. By decomposing the problem into a strategic layer for high-level mode selection and a tactical layer for fine-grained power control, we effectively manage the complexity and multi-timescale nature of station operations. The hybrid action space, coupled with the robust SAC learning algorithm, allows the system to seamlessly execute both discrete operational choices and continuous power adjustments.

Our simulation-based validation demonstrates the model’s significant value: it increases station operational profit by approximately 15-25% compared to advanced optimization baselines by capitalizing on time-of-use electricity price differences through smart V2G dispatch. Simultaneously, it substantially improves service quality, reducing average battery EV car waiting times by over 50% compared to simple rule-based strategies. The model exhibits strong robustness and scalability, maintaining performance across different station sizes and under stochastic battery EV car arrival patterns.

The proposed HDRL-based scheduler provides a practical and intelligent solution for station operators, translating into higher profitability, better customer satisfaction for battery EV car owners, and improved grid stability through predictable, schedulable V2G resources. Future work will involve testing the framework with real-world data streams, incorporating more detailed battery degradation models, and exploring multi-station coordination within a distribution network.

Scroll to Top