0049_22.Safe Reinforcement Learning for Grid-forming Inverter Based Frequency Regulation with Stability Guarantee
摘要:JOURNAL OF MODERN POWER SYSTEMS AND CLEAN ENERGY, VOL. 13, NO. 1, January 2025 ### Safe Reinforcement Learning for Grid-forming Inverter Based Frequency Regulation with ### Stability Guarantee Hang Shuai, Member, IEEE, Buxin She, Student Member, IEEE, Jinning Wang, Student Member,, and Fan
自动提取内容
JOURNAL OF MODERN POWER SYSTEMS AND CLEAN ENERGY, VOL. 13, NO. 1, January 2025
Safe Reinforcement Learning for Grid-forming Inverter Based Frequency Regulation with
Stability Guarantee
Hang Shuai, Member, IEEE, Buxin She, Student Member, IEEE, Jinning Wang, Student Member,, and Fangxing Li, IEEE Fellow, IEEE
Abstract——This study investigates a safe reinforcement learn‐ ing algorithm for grid-forming (GFM) inverter based frequency regulation. To guarantee the stability of the inverter-based re‐
source (IBR) system under the learned control policy, a model- based reinforcement learning (MBRL) algorithm is combined
with Lyapunov approach, which determines the safe region of states and actions. To obtain near optimal control policy, the
control performance is safely improved by approximate dynam‐ ic programming (ADP) using data sampled from the region of
attraction (ROA). Moreover, to enhance the control robustness against parameter uncertainty in the inverter, a Gaussian pro‐
cess (GP) model is adopted by the proposed algorithm to effec‐ tively learn system dynamics from measurements. Numerical
simulations validate the effectiveness of the proposed algorithm.
Index Terms——Inverter-based resource (IBR), virtual synchro‐ nous generator (VSG), safe reinforcement learning, Lyapunov function, frequency regulation, grid-forming inverter.
I. INTRODUCTION OWER system frequency control is critical for maintain‐
Ping grid stability when imbalance between generation
and load occurs. As the penetration of inverter-based resourc‐ es (IBRs) such as renewable energy and battery storage con‐ tinues to increase, modern power systems are facing signifi‐ cant challenges due to reduced mechanical inertia and in‐ creased disturbances. Therefore, power system stability con‐ trol has recently spurred much interest from both academia and industry [1], [2]. Various control methods have been proposed for IBRs to provide frequency regulation services [1], [3], [4]. For in‐ stance, both conventional synchronous generators (SGs) and IBR employ the frequency droop control strategy, which ad‐ justs the active power output in response to frequency devia‐
Manuscript received: November 13, 2023; revised: February 11, 2024; accept‐ ed: March 28, 2024. Date of CrossCheck: March 28, 2024. Date of online publi‐ cation: April 9, 2024. This work was funded in part by the CURENT Research Center and in part by the National Science Foundation (NSF) (No. ECCS-2033910).
tion 4.0 International License (http://creativecommons.org/licenses/by/4.0/). This article is distributed under the terms of the Creative Commons Attribu‐
H. Shuai, B. She, J. Wang, and F. Li (corresponding author) are with the De‐ partment of Electrical Engineering and Computer Science, University of Tennes‐ see, Knoxville, TN, 37996, USA (e-mail: hshuai1@utk.edu; bshe@vols.utk.edu; jwang175@vols.utk.edu; fli6@utk.edu). DOI: 10.35833/MPCE.2023.000882
tions. Droop-control-based inverters barely provide inertia support to the grid. Consequently, a droop-control-based net‐ work is typically characterized by a lack of inertia and being sensitive to faults [5]. In the event of a disturbance, the sys‐ tem frequency may undergo abrupt changes, potentially lead‐ ing to the tripping of generators or the unnecessary shedding of loads. To alleviate the negative impact of low inertia, the virtual synchronous generator (VSG) [6], [7] control was de‐ veloped. This control strategy emulates the frequency re‐ sponse characteristics of SGs, augmenting the system with virtual inertia and damping properties. Additionally, the val‐ ues of inertia and damping in VSGs are more flexible than those in SGs, which are not limited by physical conditions such as rotating mass. Therefore, IBRs can adjust the inertia adaptively to obtain faster and more stable power output [8]- [10]. However, traditional frequency regulation strategies for IBRs were usually designed based on linearized small-signal models [8], [9], [11], which makes the control performance deteriorate quickly when frequency deviations are large. Due to the challenges posed by the low inertia and nonlinearity of IBRs, advanced controls are needed to ensure grid stabili‐ ty. To deal with the challenges, various advanced frequency controllers are developed recently [12] - [14]. Among these methods, reinforcement learning (RL) technique is one of the most promising approaches. In [13], a model-free deep reinforcement learning (DRL) based load frequency control method was designed. The challenge of designing DRL- based power system stability controller lies in guaranteeing the control strategy won ’ t lead to unstable condition after disturbances. However, the conventional model-free RL- based controllers mentioned above do not yield any stability guarantees. Therefore, [15] proposed a Lyapunov-based mod‐ el-free RL strategy for primary frequency control of the pow‐ er system, which can guarantee that the system frequency reaches stable equilibrium after disturbances. In [15] and [16], Lyapunov stability theory was utilized to design the ar‐ chitecture of recurrent neural network (RNN) controllers for power networks. However, the system parameters (e.g., iner‐ tia of SGs) need to be known in prior in order to train the neural Lyapunov function [16], and whether the learned func‐ tion satisfies the Lyapunov conditions for all points in a re‐
gion needs further investigation. Given the frequent adjust‐ ì dθ ïï ïï d = ω ments of virtual inertia and damping parameters in IBRs, the t íïï (1) development of a robust DRL-based frequency regulation dω M = Pset-Pi-Dω-u(θω) controller for IBRs could enhance their integration into the îïï dt power system. where u(×) is the control action function of the battery energy The primary contribution of this work is the development storage systems (BESSs), which denotes the active charging of a safe model-based reinforcement learning (MBRL) algo‐ power (i.e., P in Fig. 1) of the BESS; M and D are the vir‐ b rithm for grid-forming (GFM) inverter based frequency regu‐ tual inertia and damping constant of the inverter, respective‐ lation. Inspired by [17], this algorithm addresses the chal‐ ly; P and P are the set point and real-time measurement set i lenges of ensuring controller stability and effectively dealing of the active power output of the inverter, respectively; and with system parameter uncertainty. In the proposed algo‐ θ and ω are the voltage phase and angular frequency devia‐ rithm, a Gaussian process (GP) model is adopted to learn tion of the inverter, respectively. More specifically, ω = ωi- the unknown nonlinear dynamics of the inverter system, and ω, where ω is the generated angular frequency of the in‐ n i approximate dynamic programming (ADP) [18], [19] is used verter output voltage, and ωn is the nominal angular frequen‐ to improve the control performance of the algorithm. More‐ cy of the inverter. In Fig. 1, Pi can be calculated as [21]: over, to guarantee the system stability under the learned con‐ trol policy, Lyapunov function is used to obtain the region of Pi=∑ViVj(Bijsin(θi-θj) +Gijcos(θi-θj))
(2) jÎ{ig} attraction (ROA). Different from pre-training a neural Lyapu‐ nov function according to system dynamics in [16], we de‐ where Bij and Gij are the susceptance and conductance com‐
sign the Lyapunov function as the value function of the Bell‐ ponents of the (ij) element of the admittance matrix Y, re‐
man ’ s equation in ADP. This allows both the Lyapunov func‐ spectively; Vi and θi are the voltage magnitude and phase of
tion and the control policy to update during training, leading node i, respectively; and θg is the voltage phase of the main
to an enlarged ROA and improved control performance si‐ grid. Note that lossy power flow model is adopted in (2).
multaneously. In addition, the controller based on the pro‐ We aim to propose a control policy to improve the dynam‐
posed algorithm is adaptive to the adjustment of inverter pa‐ ic performance of VSG after disturbances with the minimal
rameters (i. e., virtual inertia and damping coefficients), which cost. The optimal control problem can be formulated as:
means the controller will be more robust to parameter uncer‐ ìmin (u T Ru+ x T Qx) ïï u tainty. ïï ïïs.t. (1) This paper is organized as follows. Section II formulates í (3) the GFM inverter based frequency regulation problem. In ïïï - u £ u £ uˉ ï u is stabilizing Section III, the GFM inverter based frequency regulation via îïï the safe MBRL controller is designed. The numerical simula‐ where x= (θω) is the state of the VSG; Q and R are the pos‐ tions are presented in Section IV. Section V concludes the itive definite matrices; u and uˉ are the lower and upper limi‐ paper. tations of the control actions - u, respectively, which are deter‐ mined by the maximum charging and discharging capacities II. FORMULATION OF GFM INVERTER BASED FREQUENCY of BESSs; and u is the vector of u(×). As shown in Fig. 1, REGULATION PROBLEM the control action is optimized using the proposed algorithm. The diagram of GFM inverter based primary frequency control is depicted in Fig. 1. III. GFM INVERTER BASED FREQUENCY REGULATION VIA
IBR RfLf i f i RcLc SAFE MBRL CONTROLLER o GridThe primary objective of the controller is to safely learn LC filter Cf v
| LC filter C | v | |||
|---|---|---|---|---|
| v | i ,i ,v | Measurement P ,Q | BESS | |
| Voltage and | (ω,V) | VSG-based | (ω,θ) frequency regulation Safe MBRL-based | |
| current control Fig. 1. Diagram of GFM inverter based primary frequency control. | power control ω V P | P | control |
oabout the frequency dynamics of VSG from measurements and adapt the control policy π for optimal performance, with‐
*** out encountering unstable system states. This implies that f o o i i the adjustment of the control policy throughout the learning process must be performed in such a way that the system state remains within the ROA. The parameter uncertainty band nonlinearity of the AC power flow, as described in (2),
n n setQset make the design of controllers for (1) challenging. The pro‐ posed controller for GFM inverter based frequency regula‐ tion is depicted in Fig. 2. In the proposed controller, the fre‐ quency dynamics of VSG are learned by the GP model with We assume the bus voltage magnitudes to be 1 p. u., and system measurements. The ROA for a fixed policy is deter‐ neglect the reactive power flows. The frequency dynamics of mined using Lyapunov functions. And the control policy is VSG-based power control loop of the GFM inverter can be updated by ADP-based RL approach to expand the ROA. given by the swing equations [5], [16], [20]: The details of the proposed policy are presented below.
SHUAI et al.: SAFE REINFORCEMENT LEARNING FOR GRID-FORMING INVERTER BASED FREQUENCY REGULATION WITH...
Updated GP model dynamics outlined in (1) will remain stable. Conversely, if ADP: of GFM inverter Update GP model with
| ADP: | of GFM inverter | Update GP model with | |||
|---|---|---|---|---|---|
| compute policy π using (7) | new measurements | ||||
| Control | Lyapunov | Measurements | θ | ||
| policy π | function | Select most uncertain | of system states | ω | |
| ROA: {(x,u) | u | (x,π(x))≤c | }, state-action pairs | GFM inverter dynamics | |
| where c | =arg max c | from ROA | |||
| s.t. "xΘ(c)∩X | , u (x,π(x)) | θ u | ω θ | =f ω θ, u | |
| v(x)<L Fig. 2. Proposed algorithm for GFM inverter based frequency regulation. | τ | ω, |
the state ventures outside this region, the system is prone to instability. We can use Lyapunov function v to determine ROA for a fixed control policy π. Lyapunov function v is a kk+1 +1 continuously differentiable function with v(0) = 0 and v(x) > 0 for all x ¹ 0 [25]. Therefore, Lyapunov function is Lv-Lip‐ n n schitz continuous. Based on the Lyapunov stability theory, nwe have the following theorem [23], [25]. τ n k+1 k k Δv
(k)k+1 k Theorem 1 If v( f (xπ(x))) <v(x) for all x within the lev‐
kk()
el set Θ(c) ={x Î χ{0}|v(x) £c} (χ is the state space, c> 0), then Θ(c) is an ROA, so that x₀Î Θ(c) implies xkÎ Θ(c) for all k> 0 and lim xk= 0. k® ¥ By discretizing the dynamic model shown in (1) and (2), The theorem indicates that when a fixed policy π is em‐ the dynamics can be reformulated as the following nonlinear ployed, applying the dynamics f(×) to the state consistently discrete-time system: results in decreasing values in the Lyapunov function. Conse‐ ìθk+ 1= θk+hωk quently, the system state is assured to converge inevitably to‐ ïï í h (4a) wards the equilibrium point. Further details can be found in ïïωk+ 1= ωk+ M (Psetk-Pik-Dωk-uk) [23]. According to the theorem, the determination of the îïï where h is the step size for the discrete simulation; and the ROA Θ(c) is achieved by examining a level set of the Lyapu‐
subscript k denotes the discrete time index. nov function. To compute ROA, the crucial steps involve
Equation (4a) can be expressed as: identifying an appropriate Lyapunov function and determin‐ ing Θ(c) that ensures the condition v( f (xπ(x))) <v(x) holds xk+ 1= f (xkuk) =h(xkuk) +g(xkuk) (4b) for all x Î Θ(c). where f(×) denotes the true dynamics of the VSG, compris‐ The dynamics of VSG f(×) are uncertain, leading to uncer‐ ing two components: a known model represented by h(×), and tainty in v( f(×)). This introduces an additional challenge in a priori unknown model errors denoted by g(×). In inverters, determining Θ(c) using the above theorem. According to the the parameter (e. g., M and D in (1)) can undergo dynamic GP model, v( f (xu)) is contained in ϒ (xu): =[v(μ (xu)) ± n n- 1 changes, which introduces uncertainties. To ensure the stabil‐ L β σ (xu)] with probability higher than 1 -δ. L is the v n n- 1 v ity and predictability of the system, we assume the dynamic Lipschitz constant of the Lyapunov function v(×). To ensure of the VSG is Lf-Lipschitz continuous, which means that the safe state-actions are always safe, we define the upper bound dynamic does not change too rapidly between any two of v( f (xu)) as un(xu): = maxCn(xu), where Cn(xu) = points in its domain. This assumption holds true for the Cn- 1(xu) ϒn(xu). Therefore, in accordance with the afore‐ VSG system as described in (1), with a supporting proof pro‐ mentioned theorem and considering v( f (xu)) £un(xu), the vided in Appendix A. To enable safe learning, we adopt GP model to learn a re‐ system stability in (1) is assured if un(xu) <v(x) is satisfied
liable statistical system model described by (1) and (2). GP for all x Î Θ(c). Nevertheless, determining Θ(c) becomes im‐
model is a powerful method in machine learning and statisti‐ practical when attempting to identify all states x on the con‐
cal modeling. GP consists of random variables, and any fi‐ tinuous domain that satisfy un(xu) <v(x). To address this
nite group of them follows a joint Gaussian distribution. In challenge, we can discretize the state space into cells denot‐ system modeling, GP is often used to capture complex rela‐ ed as χτ, such that ||x-[x]τ||1£ τ. In this context, [x]τ repre‐ tionships in data [22]. According to the GP theory [23], sents the cell with the minimal distance to x. Considering there exists a parameter βn> 0 such that with probability at the system dynamic is Lf-Lipschitz continuous and the con‐ least 1 -δ it holds for all n³ 0 that ||f (xu) -μn(xu)||1£ trol policy is Lπ-Lipschitz continuous, we can get the follow‐ ing theorem [17]. The proof is discussed in Appendix A. βnσn(xu). μn(×) and σn(×) = trace∑(×) are the posterior Theorem 2 If un(xu) <v(x) -LDvτ holds for all x Î
(n) Θ(c) χτ and for some n³ 0, then v( f (xπ(x))) <v(x) holds mean and covariance matrix functions of the GP model of for all x Î Θ(c) with probability at least 1 -δ, where LDv= the VSG dynamics in (4b) conditioned on n measurements, L L (L + 1) +L. And Θ(c) is an ROA for the dynamics f un‐ respectively. In this way, we can use a GP model to build v f π v der policy π. confidence intervals on the inverter dynamics, which can In this way, under a fixed policy π, the ROA can be iden‐ cover the true dynamics with probability 1 -δ. tified within the discretized state space as follows: After learning about the inverter dynamics from measure‐ ments, the goal is to safely adapt the optimal control policy Dn={(xu)|un(xπ(x)) -v(x)<-LDvτ} (5)
without leading to unstable system conditions. The safety of It should be noted that the ROA is dependent on the poli‐ the controller is characterized by the safe region of states cy. To get the largest possible ROA, we can optimize the and actions, commonly referred to as the ROA [24]. When policy using (6). The corresponding optimal policy for cn the system state falls within the boundaries of the ROA, the is πn.
c n = max c "x Î Θ(c) χτ(xπ(x))ÎDn
(6) IV. SIMULATION RESULTS π ÎΠPc Î R> 0
where ΠP is the set of safe policies. A case study was conducted on a GFM inverter system, as shown in Fig. 1, to demonstrate the effectiveness of the The ROA optimized by (6) is contained in true ROA with proposed algorithm for system frequency regulation. The probability at least 1 -δ for all n> 0. Precisely solving (6) is step size of the discrete simulation of the system was set to intractable, thus we adopt the ADP [18] technique to im‐ be 0.01 s and the total simulation time horizon was 15 s. We prove the performance of the policy from data, as shown be‐ used GP model to learn the frequency dynamics of the VSG. low: The mean dynamics of the VSG were characterized by a lin‐ πn= arg min∑r(xπW(x)) + γJπ W (μn- 1(xπW(x))) + earized model of the true dynamics, as shown in (B1) in Ap‐ π W ÎΠP x Î χτ pendix B, accounting for inaccuracies in the values of M and λ(un(xπW(x)) -v(x) +LDvτ) (7) D. Consequently, the optimal policy designed for the mean
where πW is the policy with parameters W; γ is the discount dynamics exhibited suboptimal performance with a limited factor; λ is a Lagrange multiplier for the safety constraint; ROA, primarily due to underactuation of the system. We ad‐ T Topted a hybrid approach employing both linear and Matérn r(xπW(x)) =u Ru+ x Qx ³ 0 is the cost function; and Jπ W
(×) kernels (refer to Appendix C) [22], [26]. This combination is the value function of the Bellman ’ s equation, which is ap‐ enabled us to effectively capture model errors stemming proximated using piecewise linear approximations [18] in from inaccuracies in parameters. As for the policy network, this work, and Jπ W
(x) =r(xπ(x)) + γJπ W ( f (xπ(x))). Consider‐ a neural network featuring two hidden layers was implement‐
ing the cost function is strictly positive, we use Jπ W
(×) as the ed, each comprising 32 neurons with rectified linear unit Lyapunov function. In (7), the objective of the optimization (ReLU) activation functions. The states θ and ω were dis‐ cretized into 2000 and 1500 intervals, respectively. The ac‐ is to minimize the cost and make sure the safety constraint tion space was discretized into 55 intervals. R and Q in (3) holds, and stochastic gradient descent (SGD) based optimiza‐ were set to be 0.1 and êêêê é0.1 0ùúúúú, respectively. tion method can be utilized. ë 0 2û For the proposed algorithm, a safe initial point is essential The case study was conducted on an Intel Core i7-8650U for initiating the learning process. Consequently, an initial @ 1.90 GHz Windows based computer with 16 GB RAM. policy is required, ensuring the asymptotic stability of the The convergence process of the training for the proposed al‐ system origin in (1) within a confined set of states. In this gorithm is illustrated in Fig. 3. work, we utilize a linear-quadratic regulator (LQR) control‐
0.91 ler as our initial policy. In addition, to expand the ROA throughout the learning process, the agent strategically ex‐0.90 plores the state-action pairs for which the system dynamics 0.89 are most uncertain. To achieve this, we meticulously choose0.88 measurement data points based on:
0.87 (xnun) = arg max (un(xu) -ln(xu))
(8) 0.86 (xu)ÎDn Normalized objective
0.85 where ln(xu) is the lower bound of v( f (xu)). The proposed algorithm is summarized in Algorithm 1.0.840 5 10 15 20 25 30 Number of iterations Algorithm 1: safe MBRL algorithm for GFM inverter based frequency regulation Fig. 3. Convergence process of training for proposed algorithm. Load the power system simulation environment; initialize the LQR-based initial policy; initialize the parameters of the policy πW; initialize the GP model for VSG dynamics and ADP value functions; set the total The proposed algorithm exhibited remarkable conver‐ number of episodes Ne; and set the training step n= 1 gence, typically requiring only a few tens of iterations. Un‐ Get the initial safe set based on the initial LQR controller and the corre‐der the obtained control policy, the ROA is shown in Fig. 4 sponding initial Lyapunov function by the dark green area, where the light green area denotes for n £ Nedo the state space and the blue cross marks denote the data for i= 12N do points the agent selected to explore the safe region. From Based on (8), select a new safe sample of the state-action pair (xu)the result, the ROA was determined based on the informa‐ Update the GP model for VSG dynamics based on the actively select‐tion from multiple measurements. ed new data point We investigated the frequency control performance of the Optimize policy πn by solving (7) using the SGD-based optimization proposed algorithm, as depicted in Fig. 5, where the frequen‐ methodcy deviation f is derived from the angular frequency devia‐ Update the Lyapunov function (i.e., value function Jπ W
(×))tion ω, with the relationship expressed as f = ω/(2π), so in Using the updated policy, calculate cn in (6) to ensure that "x Î Θ(c) χτ, Fig. 5(a) - (c), the values of ω t® 0+ are -1, -2, and -3 rad/s, u n (xπ(x)) -v(x)<-LDvτ holds respectively. From Fig. 5, it is evident that when the inverter Compute and update the safe set (i.e., ROA) experiences frequency deviation, the proposed algorithm effi‐ Return the well-trained policy πW ciently restores the system to a stable state using BESSs. In
contrast, without any control, the system became unstable af‐ To demonstrate the superiority of the proposed algorithm ter the disturbance. Additionally, while traditional linear over traditional model-free DRL algorithms (e. g., the deep droop control can stabilize the system under certain levels of deterministic policy gradient (DDPG) algorithm outlined in disturbance, it fails to maintain stability when the distur‐ [14]), we conducted a comparative analysis of the proposed bance is significant, as shown in Fig. 5(c). The results also algorithm against DDPG and soft actor critic (SAC) algo‐ indicate that the linearized control policy could lead to a rap‐ rithms for frequency regulation. The results are depicted in id deterioration in control performance in the presence of Fig. 6. The analysis reveals that while the DDPG and SAC large frequency deviations. algorithms achieve satisfactory control performance under
1.00 relatively mild disturbances, as shown in Fig. 6(a) and (b), 0.75 managing to stabilize the inverter frequency within several seconds after the disturbance, their effectiveness diminishes
0.50 ω
0.25 with increasing disturbance magnitude. In contrast, the pro‐ posed algorithm not only restores inverter frequency more 0swiftly than the model-free algorithms in scenarios with rela‐ -0.25 tively minor disturbances but also maintains robust frequen‐ Normalized -0.50cy control under more significant disturbances, as shown in -0.75
Fig. 6(c).
0.15 -1.0-0.5 0 0.5 1.0 Normalized θ 0.10 (Hz) f
Fig. 4. ROA under safe MBRL-based control policy. 0.05
0.6 0.2 0 (Hz) f
0.1-0.05 0.4 0 -0.10 -0.1
0.2 Frequency deviation-0.15 -0.2
0 -0.3-0.200 5 10 15 -0.4 Time (s) Frequency deviation -0.2-0.5 BESS charging action (p.u.) (a) 0 5 10 15
0.2 Time (s)
(a) 0.6 0.2 (Hz)
0.1 f
0.1 0 f (Hz)0.4 0 -0.1
0.2-0.1 -0.2 -0.2 0 -0.3 Frequency deviation -0.3 -0.2 -0.4 -0.4 Frequency deviation BESS charging action (p.u.) 0 5 10 15 -0.4-0.5 0 5 10 15 Time (s) Time (s) (b)
(b) 0.8 0.6 0.50 0.6 (Hz) f
0.4 0.25 (Hz) f
0.4 0.2 0.2 0 0 0 -0.2 -0.25-0.2 -0.4 Frequency deviation BESS charging action (p.u.)Frequency deviation-0.4 -0.6-0.50 0 5 10 15-0.6 Time (s) 0 5 10 15
(c) Time (s) f without control; f of linear droop control; f of proposed algorithm (c) BESS charging action of proposed algorithm DDPG algorithm; SAC algorithm; Proposed algorithm
Fig. 5. BESS charging action of proposed algorithm and frequency control Fig. 6. Frequency control performance of proposed algorithm and model-
performance of proposed algorithm and comparing algorithms. (a) ft® 0+ = free DRL algorithms. (a) f t® 0+ = -0.159 Hz. (b) ft® 0+ = -0.318 Hz. (c) ft® 0+ = -0.159 Hz. (b) ft® 0+ = -0.318 Hz. (c) ft® 0+ = -0.478 Hz.-0.478 Hz.
The stable control performance of the proposed algorithm V. CONCLUSION can be largely attributed to the integration of Lyapunov sta‐ In this paper, we presented a novel safe MBRL algorithm bility theory into the learning process, which provides a safe‐ for GFM inverter based frequency regulation with stability ty guarantee characteristic. More specifically, the proposed guarantee. The proposed algorithm ensures stability by learn‐ algorithm selects optimal control actions within the ROA, en‐ ing a Lyapunov function and utilizes ADP-based RL to en‐ suring a level of safety that model-free DRL algorithms can‐ hance control performance. Additionally, the GP modeling not guarantee for the learned policy. was employed to capture VSG dynamics and enhance robust‐ Furthermore, to test the robustness of the proposed algo‐ ness to parameter uncertainty. The proposed algorithm offers rithm against inverter parameter variations, such as M and D a safe and robust controller for GFM inverter based frequen‐ in (1), we evaluated the performance of the well-trained safe cy regulation. Simulation results demonstrated that the per‐ MBRL controller under different parameter settings. Fig‐ formance of the proposed algorithm surpasses that of tradi‐ ure 7(a) illustrates the frequency response of the inverter tional droop control and model-free DRL algorithms. More‐ with varying D values (70% ×Dbase£ D£ 130% ×Dbase) while over, the proposed algorithm only requires the measurements the virtual inertia setting was held constant at Mbase. In of the voltage phase and angular frequency of the inverter,
Fig. 7(b), the frequency response to varying M values, devi‐
which are easily accessible in modern power systems. The ating by ±30% from the base value, was examined, while ease of implementation of the proposed algorithm enhances maintaining the damping coefficient steady at Dbase. Observa‐ its potential for practical applications. tions from Fig. 7(a) and (b) indicate that the safe MBRL- based control policy was able to effectively and safely con‐ APPENDIX A trol the BESS to provide frequency regulation, regardless of the M and D adjustments. This adaptability underscores the Lemma 1 The control policy πW is Lipschitz continuous capability of the controller to handle dynamic changes and with Lipschitz constant Lπ. uncertainties within the system, affirming its robustness Proof 1 In this work, πW= ϕ(x). ϕ(x) is the output of a K- against a wide range of operational conditions. layer network, which is given by:
0.1ϕ(x) =ϕK(ϕK- 1(ϕ₁(x;W₁); W₂);WK) (A1) In the hidden layers, ReLU activation functions are used. 0For the kth layer, there exits a constant L k > 0 such that ||ϕk(x; Wk) -ϕk(x + r; Wk)|| £Lk||r|| holds for all x and r. Here, (Hz) f-0.1 0.065 Increasing D r is a vector satisfying ||r|| £ε and ε is a small enough positive
0.060number. The output layer utilizes tanh activation function, thus -0.2
0.055 K the network satisfies ||ϕ(x) -ϕ(x + r)|| £Lπ||r||, with Lπ=∏Lk.
0.050 k= 1 -0.3 This means the control policy π is Lipschitz continuous with
0.045 W Frequency deviation
1.8 2.0 2.2 2.4 2.6 2.8 3.0Lipschitz constant L π. -0.4 Lemma 2 The closed-loop dynamics of the VSG given in (4b) are Lipschitz continuous with Lipschitz constant Lf. -0.5 0 Proof 2 From the dynamics given in (1) and Lemma 1, the 5 10 15 Time (s)dynamic function of VSG is a continuously differentiable func‐
(a) tion. Any continuously differentiable function is locally Lip‐
0.1 schitz. Therefore, the closed-loop dynamics of VSG given in (4b) are Lipschitz continuous with Lipschitz constant Lf. 0 Lemma 3 The Lyapunov function v is Lipschitz continuous
(Hz) f-0.1 0.08 with Lipschitz constant Lv. Increasing M Proof 3 In this work, the Lyapunov function is set as the
0.07 value function Jπ W of the ADP method. The value function is -0.2 0.06
0.05 approximated using a piecewise linear function that is continu‐ ous. Given that the slopes of this piecewise linear function are -0.3 0.04 bounded, the Lyapunov function exhibits Lipschitz continuity
0.03 Frequency deviation 1.0 1.5 2.0 2.5 3.0 3.5 4.0 with a Lipschitz constant denoted by Lv. -0.4 Theorem 2 can be proofed as follows. According to Lemma 1 of [17], v( f (xπ(x))) -v(x) < 0 for all continuous states -0.50 x Î Θ(c) with probability higher than 1 -δ. So, it can be con‐ 5 10 15 Time (s)cluded based on Theorem 1 that Θ(c) is an ROA for the system.
(b) Fig. 7. Robustness of safe MBRL-based control policy against D and MAPPENDIX B
uncertainties with ft® 0+ = -0.478 Hz. (a) Robustness of safe MBRL-based control policy against D uncertainty. (b) Robustness of safe MBRL-basedThe LQR-based initial policy is designed based on the lin‐ control policy against M uncertainty. earized VSG dynamics. According to formulas (1) and (2), the
linearized small-signal model of VSG around an given operat‐ ing point is obtained as:
̇ ùúú é êêê 0 1 ù úúú éêêêê Dθ éêêê 0 ù úúú éêê Dθ ùúúúú êê úú = ê 1 D ú + ê 1 úu ëDω̇ û êê- M (Bcos θ-Gsin θ)- M úú ëDωû êê- M úú ë û ë û (B1)
The eigenvalues of the system are: 2 -D ± D- 4M (Bcos θ-Gsin θ) (B2) λ 12= 2M where G+ jB = Yig is the mutual admittance between the IBR node and the main grid. As shown in Fig. 1, the mutual admit‐ tance can be calculated using the line parameters as [21]: 1-RcXc Yig= -= 2 2 + j 2 2(B3) Rc+ jXcR c + XcRc+ Xc
It can be found that the eigenvalues depend on the operating point, virtual inertia, and damping coefficients, and line param‐ eters Rc and Xc. In this work, Yig= -0.495 + j4.95. The per unit values of M and D are set to be 5 and 1, respectively.
APPENDIX C
The linear kernel is given by: k L (xx′) =x T x′ (C1)
where x and x′ are two distinct states. The Matérn kernel is given by: ν̂ 1 2ν̂ 2ν̂ k M(xx′) =ν- 1d(xx′) Kνd(xx′) (C2) Γ(ν̂)2(l) (l)
where l is a length-scale parameter; d(×) is the Euclidean dis‐ tance; Kν(×) is a modified Bessel function; Γ(×) is the Gamma function; and ν̂ is the parameter that regulates the smoothness of the function.
REFERENCES [1] A. Bidram, A. Davoudi, and F. L. Lewis, “A multiobjective distributed control framework for islanded AC microgrids,” IEEE Transactions on Industrial Informatics, vol. 10, no. 3, pp. 1785-1798, Aug. 2014. [2] D. Chen, K. Chen, Z. Li et al., “PowerNet: multi-agent deep reinforce‐ ment learning for scalable power grid control,” IEEE Transactions on Power Systems, vol. 37, no. 2, pp. 1007-1017, Mar. 2022. [3] Z. A. Obaid, L. M. Cipcigan, L. Abrahim et al., “Frequency control of future power systems: reviewing and evaluating challenges and new control methods,” Journal of Modern Power Systems and Clean Ener‐ gy, vol. 7, no. 1, pp. 9-25, Jan. 2019. [4] P. Verma, K. Seethalekshmi, and B. Dwivedi, “A cooperative approach of frequency regulation through virtual inertia control and enhance‐ ment of low voltage ride-through in DFIG-based wind farm,” Journal of Modern Power Systems and Clean Energy, vol. 10, no. 6, pp. 1519- 1530, Nov. 2022. [5] X. Meng, J. Liu, and Z. Liu, “A generalized droop control for grid- supporting inverter based on comparison between traditional droop control and virtual synchronous generator control,” IEEE Transactions on Power Electronics, vol. 34, no. 6, pp. 5416-5438, Jun. 2019. [6] J. Liu, Y. Miura, H. Bevrani et al., “Enhanced virtual synchronous generator control for parallel inverters in microgrids,” IEEE Transac‐ tions on Smart Grid, vol. 8, no. 5, pp. 2268-2277, Sept. 2017. [7] K. Sakimoto, Y. Miura, and T. Ise, “Stabilization of a power system with a distributed generator by a virtual synchronous generator func‐ tion,” in Proceedings of 8th International Conference on Power Elec‐ tronics, Jeju, South Korea, Jun. 2011, pp. 1498-1505. [8] P. He, Z. Li, H. Jin et al., “An adaptive VSG control strategy of bat‐ tery energy storage system for power system frequency stability en‐
hancement,” International Journal of Electrical Power & Energy Sys‐
[9] M tems. Li, vol, W.. Huang 149, p., N109039. Tai ,et al Jul..,2023 “A dual-adaptivity inertia control strat. ‐ egy for virtual synchronous generator,” IEEE Transactions on Power Systems, vol. 35, no. 1, pp. 594-604, Jan. 2020. [10] J. Alipoor, Y. Miura, and T. Ise, “Power system stabilization using vir‐ tual synchronous generator with alternating moment of inertia,” IEEE Journal of Emerging and Selected Topics in Power Electronics, vol. 3, no. 2, pp. 451-458, Jun. 2015. [11] F. Wang, L. Zhang, X. Feng et al., “An adaptive control strategy for virtual synchronous generator,” IEEE Transactions on Industry Appli‐ cations, vol. 54, no. 5, pp. 5124-5133, Sept. 2018. [12] A. Ademola-Idowu and B. Zhang, “Frequency stability using MPC- based inverter power control in low-inertia power systems,” IEEE Transactions on Power Systems, vol. 36, no. 2, pp. 1628-1637, Mar.
[13] Z power systems. Yan and Y. Xu : a deep reinforcement learning method with continuous, “Data-driven load frequency control for stochastic
action search,” IEEE Transactions on Power Systems, vol. 34, no. 2, pp. 1653-1656, Mar. 2019. [14] Y. Li, W. Gao, W. Yan et al., “Data-driven optimal control strategy for virtual synchronous generator via deep reinforcement learning ap‐ proach,” Journal of Modern Power Systems and Clean Energy, vol. 9, no. 4, pp. 919-929, Aug. 2021. [15] W. Cui, Y. Jiang, and B. Zhang, “Reinforcement learning for optimal primary frequency control: a Lyapunov approach,” IEEE Transactions on Power Systems, vol. 38, no. 2, pp. 1676-1688, Mar. 2023. [16] W. Cui and B. Zhang, “Lyapunov-regularized reinforcement learning for power system transient stability,” IEEE Control Systems Letters, vol. 6, pp. 974-979, Jun. 2022. [17] F. Berkenkamp, M. Turchetta, A. Schoellig et al., “Safe model-based reinforcement learning with stability guarantees,” in Proceedings of the 31st International Conference on Neural Information Processing Systems, California, USA, Dec. 2017, pp. 908-919. [18] H. Shuai, J. Fang, X. Ai et al., “Stochastic optimization of economic dispatch for microgrid based on approximate dynamic programming,” IEEE Transactions on Smart Grid, vol. 10, no. 3, pp. 2440-2452, May
[19] H. Shuai, J. Fang, X. Ai et al., “Optimal real-time operation strategy for microgrid: an ADP-based stochastic nonlinear optimization ap‐ proach,” IEEE Transactions on Sustainable Energy, vol. 10, no. 2, pp. 931-942, Apr. 2019. [20] D. Raisz, D. Deepak, F. Ponci et al., “Linear and uniform swing dy‐ namics in multimachine converter-based power systems,” International Journal of Electrical Power & Energy Systems, vol. 125, p. 106475, Feb. 2021. [21] V. Vittal, J. D. McCalley, P. M. Anderson et al., Power System Control and Stability. New York: John Wiley & Sons, 2019. [22] C. E. Rasmussen and C. K. I. Williams, Gaussian Processes for Ma‐ chine Learning. Cambridge: The MIT Press, 2005. [23] F. Berkenkamp, R. Moriconi, A. P. Schoellig et al., “Safe learning of regions of attraction for uncertain, nonlinear systems with Gaussian processes,” in Proceedings of 2016 IEEE 55th Conference on Decision and Control, Las Vegas, USA, Dec. 2016, pp. 4661-4666. [24] B. She, J. Liu, F. Qiu et al., “Systematic controller design for inverter- based microgrids with certified large-signal stability and domain of at‐ traction,” IEEE Transactions on Smart Grid, doi: 10.1109/TSG.2023. 3330705 [25] H. K. Khalil, Nonlinear Systems. London: Prentice Hall, 1996. [26] D. Duvenaud. (2014, May). The kernel cookbook: advice on covari‐ ance functions. [Online]. Available: https://www. cs. toronto. edu/duven‐ aud/cookbook
Hang Shuai received the B.Eng. degree from Wuhan Institute of Technolo‐ gy (WIT), Wuhan, China, in 2013, and the Ph. D. degree in electrical engi‐ neering from Huazhong University of Science and Technology (HUST), Wu‐ han, China, in 2019. He was also a Visiting Student Researcher with the University of Rhode Island (URI), Kingston, USA, from 2018 to 2019. He was a Postdoctoral Researcher with the URI and University of Tennessee (UTK), Knoxville, USA, from 2019 to 2022. Currently, he is a Research As‐ sistant Professor with the UTK. His research interests include reinforcement learning for power system, microgrid operation and control, and bulk power system resilience.
Buxin She received the B.S.E.E and M.S.E.E degrees from Tianjin Universi‐
ty, Tianjin, China, in 2017 and 2019, respectively, and the Ph. D. degree from the University of Tennessee, Knoxville, USA, in 2023, all in electrical engineering. He is currently a Research Engineer in Pacific Northwest Na‐ tional Laboratory (PNNL), Richland, USA. He served as a Student Guest Editor of IET-RPG. He was an outstanding reviewer of IEEE Access Jour‐ nal of Power and Energy (OAJPE) (2020) and Journal of Modern Power Systems and Clean Energy (MPCE) (2022 and 2023). His research interests include microgrid operation and control, machine learning in power sys‐ tems, distribution system operation and plan, and power grid resilience.
Jinning Wang received the B.S. and M.S. degrees in electrical engineering from Taiyuan University of Technology, Taiyuan, China, in 2017 and 2020, respectively. He is currently pursuing a Ph.D. degree in electrical engineer‐ ing at the University of Tennessee, Knoxville, USA. He is the author of AMS, a power system dispatch simulator, which is a key component of the CURENT Large-scale Testbed. He also curates the list Popular Open Source Libraries for Power System Analysis. His research interests include data
mining, scientific computation, and power system simulation.
Fangxing Li (also known as Fran Li) received the B. S. E. E. and M. S. E. E. degrees from Southeast University, Nanjing, China, in 1994 and 1997, re‐ spectively, and the Ph. D. degree from Virginia Tech, Blacksburg, USA, in
- He is currently the John W. Fisher Professor of electrical engineering and the Director of CURENT with the University of Tennessee, Knoxville, USA. From 2020 to 2021, he was the Chair of IEEE PES Power System Operation, Planning and Economics (PSOPE) Committee. He has been the Chair of IEEE WG on Machine Learning for Power Systems since 2019 and the Editor-in-Chief of IEEE Open Access Journal of Power and Energy (OAJPE) since 2020. He was the recipient of numerous awards and honors, including R&D 100 Award in 2020, IEEE PES Technical Committee Prize Paper awards in 2019 and 2024, five best or prize paper awards at interna‐ tional journals, and seven best papers/posters at international conferences. His research interests include resilience, artificial intelligence in power, de‐ mand response, distributed generation and microgrid, and electricity market.