14. Neural Networks in the Loop: QCs, IQCs & Dissipativity

Stability of NN-controlled systems via quadratic constraints, Zames–Falb multipliers, offset-free tracking, dissipative RNNs, synthesis with guarantees

Before you start

This module assumes:

Book contents · Apply this chapter to a real decision · Glossary

Contents
1. The Feedback Setup and Loop Transformations 2. Local Sector QCs and the Stability LMI (Yin, Seiler, Arcak) 3. Dynamic Multipliers: Acausal Zames–Falb (Pauli et al.) 4. Offset-Free Setpoint Tracking With NN Controllers (Pauli et al.) 5. Dissipativity Analysis and Training of RNNs (Pauli et al.) 6. Synthesis With Guarantees: Dissipativity-Constrained RL and Youla-REN 7. Reachability of NN Loops and Verified Lyapunov Controllers 8. Towards Scale: ReLU IQCs and Incremental Analysis (2025–2026) Walkthrough: Stability LMI for a Linear Plant With a One-Layer NN Interactive: A Neural Controller in Closed Loop Graded practice: Easy / Medium / Hard Original research exercises Application lab & chapter review Exercises and readiness check Key Papers Flashcards
Study route — Keep the equilibrium, domain and time quantifier visible

On a first pass, follow the links below and solve the Easy set before opening the research exercises. Try each question on paper; open its hint only when stuck, then compare your reasoning with the worked solution. Use Medium questions to connect the algebra and Hard questions to check the assumptions.

§1: equilibrium and shifted signals · §2: local certificate and invariance · §3: pointwise versus summed constraints · §5: storage and gain

Practice: 4 Easy · 4 Medium · 4 Hard · original research exercises. Check readiness or use the course study guide.

Read one closed-loop energy calculation first

For the scalar loop $x^+=0.8x$, choose $V=x^2$. Substitution gives $V(x^+)-V(x)=(0.8^2-1)x^2=-0.36x^2$. This proves a decrease at every nonzero state. With a neural controller, the expression also contains activation variables $w$; a QC lets us prove the decrease without solving for every possible $w$ separately.

Track three questions alongside every theorem: around which equilibrium are the differences taken, on which state/pre-activation set is the QC valid, and at which times is it nonnegative? A local sector needs an invariant set inside its domain. A hard IQC may be nonnegative only after summing across time. Review Lyapunov sublevel sets, quadratic-constraint signs, and induction as needed.

A reinforcement-learning agent (see Primer E) or an imitation of an expensive MPC (see Primer D) hands you a neural network $\pi$ that seems to control the plant well in simulation. Before it runs on hardware you want a certificate: a proof that the closed loop is stable, together with a region of initial states from which this is guaranteed. The network is a composition of affine maps and scalar activations (see Primer E), and every common activation obeys simple quadratic inequalities (sector and slope bounds). Pulling all activations into one diagonal block turns the closed loop into a Lur'e system, a linear system in feedback with a static nonlinearity, and sixty years of robust control apply: the S-procedure, integral quadratic constraints (IQCs), dissipativity and multipliers. This page follows the line that Patricia Pauli and co-authors built on the template of Yin, Seiler and Arcak (local sector constraints, then Zames–Falb multipliers, offset-free tracking and recurrent networks, for which see Primer E), then turns to synthesis, to reachability and to verified neural Lyapunov functions, and ends with the 2025–2026 work on scale. The tools are those of Module 2; the network-level constraints are those of Module 12.

Notation on this page
Time index $t$ (the papers write $k$). Plant $x_{t+1}=Ax_t+Bu_t$ (see Primer D) with $x\in\mathbb R^{n_x}$, $u\in\mathbb R^{n_u}$. A network with $\ell$ hidden layers: $w^0=x$, $v^i=W_{i-1}w^{i-1}+b_{i-1}$, $w^i=\phi(v^i)$ for $i=1,\dots,\ell$, and $u=W_\ell w^\ell+b_\ell$. We index weights from $0$ as Pauli et al. do; Yin et al. write $W^1,\dots,W^{\ell+1}$ for our $W_0,\dots,W_\ell$. The vectors $v,w\in\mathbb R^{n_\phi}$ stack all pre- and post-activations ($n_\phi$ = number of neurons; stacked vectors as in Primer A), $\varphi:\mathbb R\to\mathbb R$ is the scalar activation and $\phi$ applies it elementwise. Sector bounds $[\alpha,\beta]$ and slope bounds $[\mu,\nu]$ (Module 12 writes slope bounds as $[\alpha,\beta]$; we need both at once in Section 3), per neuron collected in $D_\alpha=\mathrm{diag}(\alpha_i)$, $D_\beta=\mathrm{diag}(\beta_i)$; this $\beta$ is unrelated to the GP scaling $\beta_t$ of Modules 3–6. Multiplier $T=\mathrm{diag}(\lambda_1,\dots,\lambda_{n_\phi})\succeq0$ (see Primer A) as in Module 2: the $\lambda_i$ are S-procedure multipliers, not the CMDP multiplier of Module 8. Lyapunov or storage matrix $P\succ0$, $V(x)=x^\top Px$, ellipsoid $\mathcal E(P,x_*)=\{x:(x-x_*)^\top P(x-x_*)\le1\}$ (see Primer A). The stability LMI matrix is $\mathcal M$; an IQC multiplier is factorised as $\Pi=\Psi^{*}M\Psi$ (see Primer D) with middle matrix $M$ (Pauli et al. call the middle matrix $P$ and the Lyapunov matrix $X$; we translate). $\gamma$ is an $\ell_2$ gain (see Primer D), never a discount. Equilibrium quantities carry a star, deviations a tilde: $\tilde x=x-x_*$.

1. The Feedback Setup and Loop Transformations

Start with the plainest object: a linear time-invariant plant (see Primer D) and a feedforward network used as a state-feedback controller. (Output feedback $u=\pi(Cx)$ only changes the first weight to $W_0C$; Yin et al., Remark 2.)

Definition — LTI plant with a feedforward NN controller
$$\begin{aligned}&x_{t+1}=Ax_t+Bu_t,\qquad u_t=\pi(x_t):\\ &w^0_t=x_t,\qquad v^i_t=W_{i-1}w^{i-1}_t+b_{i-1},\qquad w^i_t=\phi(v^i_t)\quad(i=1,\dots,\ell),\\ &u_t=W_\ell w^\ell_t+b_\ell .\end{aligned}$$ An equilibrium is a tuple $(x_*,u_*,v_*,w_*)$ with $x_*=Ax_*+Bu_*$ and $u_*=\pi(x_*)$, where $v_*,w_*$ are obtained by propagating $x_*$ through the layers. Its region of attraction (ROA) is $\{x_0:\ x_t\to x_*\}$. The goal is an ellipsoid $\mathcal E(P,x_*)$ that is provably contained in the ROA, as large as possible (Yin et al., Def. 1 and eq. 1).

The first move, used by Fazlyab, Morari & Pappas (DeepSDP) for a network alone and by Yin, Seiler & Arcak, TAC 2022 for the loop, is to separate what is linear from what is not.

Key equation — Isolating the nonlinearities (LFT = linear fractional transformation; Lur'e form)
$$\begin{bmatrix}u_t\\ v_t\end{bmatrix}=\underbrace{\begin{bmatrix}N_{ux} & N_{uw} & N_{ub}\\ N_{vx} & N_{vw} & N_{vb}\end{bmatrix}}_{N}\begin{bmatrix}x_t\\ w_t\\ 1\end{bmatrix},\qquad w_t=\phi(v_t),$$ with $N_{ux}=0$, $N_{uw}=\begin{bmatrix}0 & \cdots & 0 & W_\ell\end{bmatrix}$, $N_{ub}=b_\ell$ and $$N_{vx}=\begin{bmatrix}W_0\\ 0\\ \vdots\\ 0\end{bmatrix},\qquad N_{vw}=\begin{bmatrix}0 & & & \\ W_1 & 0 & & \\ & \ddots & \ddots & \\ & & W_{\ell-1} & 0\end{bmatrix},\qquad N_{vb}=\begin{bmatrix}b_0\\ \vdots\\ b_{\ell-1}\end{bmatrix}.$$ The closed loop is the feedback interconnection of a linear system and the diagonal static map $\phi$: $\ x_{t+1}=Ax_t+BN_{uw}w_t+BN_{ub}$, $\ v_t=N_{vx}x_t+N_{vw}w_t+N_{vb}$, $\ w_t=\phi(v_t)$.
linear part: plant (A, B) and all affine layers weights W₀…Wℓ, biases, matrix N φ = diag(φ, …, φ) nφ scalar activations, static v w

All pre-activations $v$ leave the linear block, all activations $w$ return to it; the state and the control stay inside the linear block.

Why this form. Every certificate on this page has the same shape: a quadratic Lyapunov or storage function for the linear block, plus quadratic constraints that everything in the $\phi$ block is known to satisfy, glued together by the S-procedure. Because $N_{vw}$ is strictly block lower triangular, $w$ can be computed layer by layer, so the loop is well posed. Implicit (equilibrium) networks and RENs allow a full $N_{vw}$; then well-posedness needs a condition such as $2T-TN_{vw}-N_{vw}^\top T\succ0$ for a diagonal $T\succ0$ and activations slope-restricted in $[0,1]$ (Revay, Wang & Manchester, TAC 2024, eq. 12).

Derivation — Removing the biases by an equilibrium shift

Step 1 (equilibrium equations). An equilibrium satisfies $x_*=Ax_*+Bu_*$, $\begin{bmatrix}u_*\\ v_*\end{bmatrix}=N\begin{bmatrix}x_*\\ w_*\\ 1\end{bmatrix}$, $w_*=\phi(v_*)$ (Yin et al., eq. 12). Finding one is a fixed-point problem $x_*=Ax_*+B\pi(x_*)$; for a bias-free network whose activation satisfies $\varphi(0)=0$ (tanh, ReLU), $x_*=0$ always works, with $v_*=w_*=0$. Bias-free is not enough on its own: a sigmoid has $\sigma(0)=1/2$, so $\pi(0)\ne0$ in general.

Step 2 (subtract). Subtracting the equilibrium equations from the loop equations, the constant column $N_{\cdot b}$ multiplies $1-1=0$:

$$\tilde x_{t+1}=A\tilde x_t+B\tilde u_t,\qquad \begin{bmatrix}\tilde u_t\\ \tilde v_t\end{bmatrix}=\begin{bmatrix}N_{ux} & N_{uw}\\ N_{vx} & N_{vw}\end{bmatrix}\begin{bmatrix}\tilde x_t\\ \tilde w_t\end{bmatrix},\qquad \tilde w_t=\tilde\phi(\tilde v_t),$$

with the shifted activations $\tilde\varphi_i(s)=\varphi(s+v_{*,i})-\varphi(v_{*,i})$.

Step 3 (what survives). (i) $\tilde\varphi_i(0)=0$, so the origin is the equilibrium of the shifted loop. (ii) Every chord of $\tilde\varphi_i$ is a chord of $\varphi$, so slope bounds are inherited unchanged. (iii) What is lost: if $v_{*,i}\ne v_{*,j}$ then $\tilde\varphi_i\ne\tilde\varphi_j$. The stacked map is then no longer a repeated nonlinearity (the same scalar function in every channel), and multipliers that exploit a repeated nonlinearity (coupling different neurons through monotonicity, such as ZF-R) become invalid between such neurons (Section 3, Exercise 14.6); multipliers that use only each neuron's own sector or slope bounds, including the full-block circle multipliers of Section 3, stay valid. Pauli et al. (CDC 2021) make exactly this point: the shift "eliminates the bias terms", but "in general" the components "are not identical", so they use repeated-nonlinearity IQCs only for bias-free networks at $x_*=0$. That case (with $\varphi(0)=0$, so that every $v_{*,i}=0$) is sufficient, not necessary: coupling stays valid within any group of neurons that carry the same activation with equal $v_{*,i}$ and share their slope bounds (Section 3). A general equilibrium shift destroys this.

Definition — Sector and slope restrictions, global and local
Let $\mathcal I\subseteq\mathbb R$ be an interval containing $v_*$, and write $\Delta v=v-v_*$, $\Delta\varphi=\varphi(v)-\varphi(v_*)$.
  • $\varphi$ satisfies the (offset) sector $[\alpha,\beta]$ around $(v_*,\varphi(v_*))$ on $\mathcal I$ if $(\Delta\varphi-\alpha\,\Delta v)(\beta\,\Delta v-\Delta\varphi)\ge0$ for all $v\in\mathcal I$ (Yin et al., Defs. 2–4). Global: $\mathcal I=\mathbb R$; local: a bounded $\mathcal I$; centred: $v_*=0$, $\varphi(0)=0$.
  • $\varphi$ is slope-restricted in $[\mu,\nu]$ on $\mathcal I$ if $\mu\le\frac{\varphi(v)-\varphi(v')}{v-v'}\le\nu$ for all $v\ne v'$ in $\mathcal I$.
Both are quadratic constraints: the sector condition is $$\begin{bmatrix}\Delta v\\ \Delta\varphi\end{bmatrix}^{\top}\begin{bmatrix}-2\alpha\beta & \alpha+\beta\\ \alpha+\beta & -2\end{bmatrix}\begin{bmatrix}\Delta v\\ \Delta\varphi\end{bmatrix}\ge0,$$ and slope restriction is the same inequality for the pair $(v-v',\ \varphi(v)-\varphi(v'))$ with $(\mu,\nu)$ in place of $(\alpha,\beta)$. Slope restriction on $\mathcal I$ implies the sector $[\mu,\nu]$ around every point of the graph over $\mathcal I$; the converse fails. tanh and ReLU are globally slope-restricted, hence sector-bounded, in $[0,1]$. On $|v|\le\bar v$ the centred tanh has the local sector $[\tanh(\bar v)/\bar v,\ 1]$ and the local slope bounds $[1-\tanh^2\bar v,\ 1]$ (Exercise 14.2).
Why local bounds are unavoidable
The global sector $[0,1]$ contains the function $w\equiv0$, the controller switched off. A certificate built only from the global sector therefore proves stability for every function in the sector, including that one, and cannot exist when the open-loop plant is unstable (the walkthrough makes this exact). Near the equilibrium tanh is almost the identity, and a local sector $[\alpha,1]$ with $\alpha\gt0$ excludes the switched-off controller. The price is that the constraint is valid only while the pre-activations stay in a box, so the certificate must also prove that they do. For the lower sector bound use the chord slope $\tanh(\bar v)/\bar v$, not the smallest slope $1-\tanh^2\bar v$: the latter is the right lower bound for slope restriction and a valid but loose sector bound (for $\bar v=2$: $0.482$ against $0.071$; Module 2 compares the two).
Key equation — Loop transformation: any sector to $[0,1]$
For an offset sector everything below is applied to the deviations $\tilde x=x-x_*$, $\tilde v=v-v_*$, $\tilde w=w-w_*$ (the biases then cancel); we drop the tildes. Given a sector $[D_\alpha,D_\beta]$ with $D_\beta-D_\alpha\succ0$, define $\hat w=(D_\beta-D_\alpha)^{-1}(w-D_\alpha v)$, i.e. $w=D_\alpha v+(D_\beta-D_\alpha)\hat w$. Per neuron, $w_i-\alpha_iv_i=(\beta_i-\alpha_i)\hat w_i$ and $\beta_iv_i-w_i=(\beta_i-\alpha_i)(v_i-\hat w_i)$, so $$(w_i-\alpha_iv_i)(\beta_iv_i-w_i)=(\beta_i-\alpha_i)^2\,\hat w_i\,(v_i-\hat w_i):$$ $w$ is in the sector $[\alpha,\beta]$ iff $\hat w$ is in the sector $[0,1]$. Substituting $w$ into the linear block gives $(I-N_{vw}D_\alpha)v=N_{vx}x+N_{vw}(D_\beta-D_\alpha)\hat w$, and $I-N_{vw}D_\alpha$ is invertible (block lower triangular with identity blocks on the diagonal), so the transformed loop is again a linear system in feedback with a $[0,1]$-sector block. Pauli et al. use the multiplier-level version of this move: Zames–Falb multipliers for slope $[0,\infty]$ are mapped to slope $[\mu,\nu]$ by a congruence with $L_{\mu\nu}=\begin{bmatrix}I\otimes\mathrm{diag}(\mu) & -I\\ -I\otimes\mathrm{diag}(\nu) & I\end{bmatrix}$ (CDC 2021, Sec. II-D; the paper calls it $T$, renamed here because $T$ is our diagonal multiplier). The Kronecker product $I\otimes\mathrm{diag}(\mu)$ (Primer A) repeats $\mathrm{diag}(\mu)$ once for every time sample in the multiplier's window, so $L_{\mu\nu}$ maps the stacked window of $(v,w)$ to $(\mu v-w,\ w-\nu v)$, sample by sample. Per neuron this turns the graph of $\varphi$ into the monotone relation $\nu_iv_i-w_i\mapsto w_i-\mu_iv_i$ (see the background below; $L_{\mu\nu}$ produces the negatives of these two quantities, which leaves their product unchanged). Neurons with different $(\mu_i,\nu_i)$ get different relations even when $\varphi$ is the same, so multipliers that couple neurons need common slope bounds within each coupled group (Section 3).
Background — Monotone relations

A relation is a set of pairs $(a,b)$; unlike a function it may assign several $b$ values to one $a$. It is monotone when $(a-a')(b-b')\ge0$ for every two of its pairs. Here $a=\nu v-w$ and $b=w-\mu v$: a chord of $\varphi$ with slope $s\in[\mu,\nu]$ gives $\Delta a=(\nu-s)\Delta v$, $\Delta b=(s-\mu)\Delta v$, so $(\Delta a)(\Delta b)=(\nu-s)(s-\mu)(\Delta v)^2\ge0$. If $s=\nu$ on some interval, then $\Delta a=0$ while $\Delta b\ne0$: one value of $a$ carries several values of $b$, which is why the transformed object is a relation rather than a function. The monotonicity arguments of Section 3 use only such product inequalities (and $ab\ge0$ through the origin), so multivaluedness does no harm.

Going deeper — Absolute stability: Lur'e systems, circle criterion, Zames–Falb

The Lur'e problem asks for stability of a linear system $G$ in feedback with any nonlinearity from a class (a sector, or a slope restriction). Two classical warnings shape everything below. First, checking the linear gains in the sector is not enough in general (Aizerman's conjecture is false; see Primer D); certificates must use the constraint itself. Second, the answer depends on the class: sector-bounded, slope-restricted and slope-restricted and odd nonlinearities admit progressively richer multipliers.

Static multipliers (a constant $\lambda\ge0$ times the sector QC, the S-procedure) give circle-criterion type tests; by the KYP lemma (Module 2, Rantzer 1996) the frequency-domain and LMI forms are equivalent (the strict versions without any controllability assumption). Dynamic multipliers exploit that the nonlinearity has no memory and bounded slopes. The continuous-time Zames–Falb theorem, in the form collected by Carrasco, Turner & Heath, EJC 2016 (Thm. 1): let $G$ be a stable, proper, real-rational scalar transfer function in negative feedback $v=f-Gw$, $w=\varphi(v)+e$ with zero initial state, let $\varphi(0)=0$ with every chord slope in $[0,k]$, $0\lt k\lt\infty$, and assume that the loop is well posed (unique causal solutions for every input of finite energy on each finite interval). If $M(j\omega)=m_0-H(j\omega)$, where $H$ has a possibly two-sided, integrable impulse response $h$ (see Primer D) with $\|h\|_1\lt m_0$ and $h\ge0$ unless $\varphi$ is odd, and there is an $\varepsilon\gt0$ with $$\mathrm{Re}\big\{M(j\omega)\,(1+kG(j\omega))\big\}\ge\varepsilon\quad\forall\omega\in\mathbb R,$$ then the loop is $L_2$-stable: there is a $c\lt\infty$ with $\|(v,w)\|_{L_2[0,T]}\le c\,\|(f,e)\|_{L_2[0,T]}$ for every horizon $T$, a bound from input energy to signal energy, not a region-of-attraction statement. With $H=0$ this is the circle-type test. Both qualifications matter. For $G(s)=-s/(s+1)=-1+1/(s+1)$, $M=1$ and $k=1$ one has $\mathrm{Re}\{1+G(j\omega)\}=1/(1+\omega^2)\gt0$ at every frequency, but not uniformly. With $\varphi(v)=v$ that loop is not well posed: the direct feedthrough $G(\infty)=-1$ makes the coefficient of the instantaneous (algebraic) loop $1+G(\infty)=0$. And $(1+G)^{-1}=s+1$ has a frequency response that grows without bound, so it has no finite $L_2$ gain. Sections 2–3 use discrete-time, time-domain versions of both tests.

2. Local Sector QCs and the Stability LMI (Yin, Seiler, Arcak)

Yin, Seiler & Arcak (TAC 67(4):1980–1987, 2022; arXiv 2020) gave the template that the rest of this page extends: a quadratic Lyapunov function, local sector constraints on the activations obtained by interval bound propagation, and one SDP whose solution is an ellipsoidal inner approximation of the ROA. Their Theorem 2 adds plant perturbations described by IQCs.

Lemma — Stacked offset local sector QC (Yin et al., Lemma 1)
Let $\alpha,\beta,\underline v,\bar v,v_*\in\mathbb R^{n_\phi}$ with $\alpha\le\beta$ and $\underline v\le v_*\le\bar v$ elementwise, $w_*=\phi(v_*)$, and suppose the $i$-th activation satisfies the offset local sector $[\alpha_i,\beta_i]$ around $(v_{*,i},w_{*,i})$ for $v_i\in[\underline v_i,\bar v_i]$. Then for every $T=\mathrm{diag}(\lambda)$, $\lambda\ge0$, $$\begin{gathered}\begin{bmatrix}v-v_*\\ w-w_*\end{bmatrix}^{\top}\underbrace{\begin{bmatrix}-2D_\alpha D_\beta T & (D_\alpha+D_\beta)T\\ (D_\alpha+D_\beta)T & -2T\end{bmatrix}}_{Q_{\alpha\beta}(T)}\begin{bmatrix}v-v_*\\ w-w_*\end{bmatrix}\ \ge\ 0\\ \text{for all }v\in[\underline v,\bar v],\ w=\phi(v).\end{gathered}$$ Yin et al. write the same matrix as $\Psi_\phi^{\top}M_\phi(\lambda)\Psi_\phi$ with $\Psi_\phi=\begin{bmatrix}D_\beta & -I\\ -D_\alpha & I\end{bmatrix}$, $M_\phi(\lambda)=\begin{bmatrix}0 & T\\ T & 0\end{bmatrix}$ (multiply out). Proof. The form equals $2\sum_i\lambda_i(\Delta w_i-\alpha_i\Delta v_i)(\beta_i\Delta v_i-\Delta w_i)$, a nonnegative combination of nonnegative numbers. $\blacksquare$

The bounds $[\underline v,\bar v]$ must be computed, and they must contain every pre-activation the closed loop can produce. Yin et al. choose a box for the first layer and push it through the network with interval bound propagation (Gowal et al., IBP):

Algorithm 1: Local sector bounds by interval bound propagation (Yin et al., Sec. II-D)
  1. Choose a first-layer box $[\underline v^1,\bar v^1]$ symmetric about $v^1_*$, e.g. $\bar v^1-v^1_*=v^1_*-\underline v^1=\delta\,\mathbf 1$.
  2. For $i=1,\dots,\ell-1$: // monotone activation
  3. output box $[\underline w^i,\bar w^i]=[\phi(\underline v^i),\phi(\bar v^i)]$; centre $c=\tfrac12(\bar w^i+\underline w^i)$, radius $r=\tfrac12(\bar w^i-\underline w^i)$;
  4. $\bar v^{i+1}_j=(W_i)_jc+(b_i)_j+\sum_k|(W_i)_{jk}|r_k$, $\ \underline v^{i+1}_j=(W_i)_jc+(b_i)_j-\sum_k|(W_i)_{jk}|r_k$ // exact max/min of an affine function over a box
  5. For each neuron compute a valid offset sector on its interval. For tanh: $\beta_i=1$ and $\alpha_i=\min\Big\{\frac{\tanh\bar v_i-\tanh v_{*,i}}{\bar v_i-v_{*,i}},\ \frac{\tanh v_{*,i}-\tanh\underline v_i}{v_{*,i}-\underline v_i}\Big\}$ (Yin et al., Sec. II-C; $\beta_i$ can be tightened further).

Why the interval step is exact, and where it loses. For a row $a=(W_i)_j$ write each output as $w_k=c_k+e_k$ with $|e_k|\le r_k$. Then $aw=ac+\sum_ka_ke_k$ is largest when every $e_k=r_k\,\mathrm{sign}(a_k)$, which gives $ac+\sum_k|a_k|r_k$ (smallest with the opposite signs). The $e_k$ can be chosen independently only because the set is a box; the true outputs of a layer are correlated, so the box contains combinations the network never produces, and the overestimate compounds layer by layer. In the last step $\beta_i=1$ is a valid, not always the tightest, upper bound; if an endpoint coincides with $v_{*,i}$, its chord ratio is read as the limit $\tanh'(v_{*,i})$, and a degenerate interval $\{v_{*,i}\}$ satisfies the sector QC trivially.

Theorem — Local stability and an ellipsoidal ROA estimate (Yin, Seiler & Arcak, TAC 2022, Thm. 1)
Assumptions. (i) LTI plant and $\ell$-layer feedforward network as above, with an equilibrium $(x_*,u_*,v_*,w_*)$. (ii) A first-layer box $\bar v^1$, $\underline v^1=2v^1_*-\bar v^1$, and vectors $\alpha,\beta$ such that $\phi$ satisfies the offset local sector $[\alpha,\beta]$ around $(v_*,w_*)$ for all $v$ in the box obtained by Algorithm 1 (their Property 1). (iii) With $R_V=\begin{bmatrix}I & 0\\ N_{ux} & N_{uw}\end{bmatrix}$ and $R_\phi=\begin{bmatrix}N_{vx} & N_{vw}\\ 0 & I\end{bmatrix}$ there exist $P\succ0$ and $T=\mathrm{diag}(\lambda)$, $\lambda\ge0$, such that $$\mathcal M(P,T):=R_V^{\top}\begin{bmatrix}A^{\top}PA-P & A^{\top}PB\\ B^{\top}PA & B^{\top}PB\end{bmatrix}R_V+R_\phi^{\top}Q_{\alpha\beta}(T)R_\phi\ \prec\ 0,$$ $$\begin{bmatrix}(\bar v^1_i-v^1_{*,i})^2 & (W_0)_i\\ (W_0)_i^{\top} & P\end{bmatrix}\succeq0,\qquad i=1,\dots,n_1,$$ where $(W_0)_i$ is the $i$-th row of the first weight.
Statement. (a) $x_*$ is locally asymptotically stable (the paper says "locally stable"; the proof gives asymptotic stability) and (b) $\mathcal E(P,x_*)$ is an inner approximation of the ROA.
In words. On the ellipsoid, every first-layer pre-activation stays in its box, so every activation obeys its local sector; the first LMI says that $V$ strictly decreases for every nonlinearity obeying those sectors.
Why each assumption matters. The equilibrium: the QCs are centred at it, and without one there is nothing to be stable. The box and Property 1: the sector bounds are only true inside the box; a box that misses reachable pre-activations makes the certificate unsound. The second LMI: it is what keeps the state inside the box (below), without it the local sectors may be used where they are false (Exercise 14.4). Strictness of the first LMI: it yields a margin $\varepsilon\|x-x_*\|^2$ and hence convergence, not just boundedness.
Proof sketch — Theorem 1 of Yin et al., with the reason for every step

Step A (ellipsoid inside the box). The Schur complement of the second LMI with respect to $P\succ0$ gives $(W_0)_iP^{-1}(W_0)_i^{\top}\le(\bar v^1_i-v^1_{*,i})^2$. For $x\in\mathcal E(P,x_*)$ put $y=x-x_*$, $y^{\top}Py\le1$. By Cauchy–Schwarz in the $P$ inner product,

$$|(W_0)_iy|=\big|\big(P^{-1/2}(W_0)_i^{\top}\big)^{\top}\big(P^{1/2}y\big)\big|\le\sqrt{(W_0)_iP^{-1}(W_0)_i^{\top}}\ \sqrt{y^{\top}Py}\le\bar v^1_i-v^1_{*,i}.$$

Since $v^1-v^1_*=W_0(x-x_*)$, every $x\in\mathcal E$ produces $v^1\in[\underline v^1,\bar v^1]$, and by Algorithm 1 every later pre-activation is in its box. So on $\mathcal E$ the QC of the Lemma holds. (Yin et al. cite a lemma of Hindi and Boyd for this containment; the Cauchy–Schwarz line above is its proof, and the bound is attained, so the condition is also necessary.)

Step B (the quadratic form is the Lyapunov difference plus the QC). Strictness gives $\varepsilon\gt0$ with $\mathcal M\preceq-\varepsilon I$. Multiply by $\zeta_t=(\tilde x_t,\tilde w_t)$ on both sides. Because $R_V\zeta_t=(\tilde x_t,\tilde u_t)$, $R_\phi\zeta_t=(\tilde v_t,\tilde w_t)$ and $\tilde x_{t+1}=A\tilde x_t+B\tilde u_t$:

$$\begin{gathered}\underbrace{V(x_{t+1})-V(x_t)}_{\text{first term}}+\underbrace{\begin{bmatrix}\tilde v_t\\ \tilde w_t\end{bmatrix}^{\top}Q_{\alpha\beta}(T)\begin{bmatrix}\tilde v_t\\ \tilde w_t\end{bmatrix}}_{\ge0\ \text{if}\ x_t\in\mathcal E}\ \le\ -\varepsilon\|\zeta_t\|^2\le-\varepsilon\|x_t-x_*\|^2,\\ V(x)=(x-x_*)^{\top}P(x-x_*).\end{gathered}$$

Step C (induction removes the circularity; proof by induction: Primer 0). If $x_t\in\mathcal E$, Step A makes the QC valid at time $t$, so $V(x_{t+1})\le V(x_t)-\varepsilon\|x_t-x_*\|^2\le1$, i.e. $x_{t+1}\in\mathcal E$. Only the QC at the current time is used, so there is no circular argument: invariance is proved one step at a time.

Step D (convergence). Summing, $\varepsilon\sum_t\|x_t-x_*\|^2\le V(x_0)\le1$. The terms of a convergent series tend to zero (Primer 0): otherwise $\|x_t-x_*\|\ge a$ for some $a\gt0$ and infinitely many $t$, and each such $t$ would add at least $\varepsilon a^2$ to the sum. So $x_t\to x_*$. Stability in the sense of Lyapunov (starting close means staying close) comes from the invariant sublevel sets $\{V\le c\}$, $c\le1$, and $\lambda_{\min}(P)\|x-x_*\|^2\le V(x)\le\lambda_{\max}(P)\|x-x_*\|^2$; with $V$ a strict Lyapunov function on the invariant set $\mathcal E$ this is local asymptotic stability (Khalil, Thm. 4.1, as cited by Yin et al.). $\blacksquare$

Read the two LMIs as two different jobs

The containment LMI checks where the local activation model is allowed: every state in the proposed ellipsoid must place the first pre-activation inside its chosen box. The decrease LMI then checks what happens next from any such state. Starting in the ellipsoid, one step lowers $V$, keeps the state inside the same ellipsoid, and permits the argument again.

This order matters. First establish containment at the current state, then decrease, then invariance by induction. A simulation that stays inside the box is useful evidence about that trajectory, but it cannot replace containment for every state in the certified set. The Easy ellipsoid question and Hard invariance question let you practice these jobs separately.

Computing a large certified ellipsoid. For a fixed box both LMIs are linear in $(P,\lambda)$, and Yin et al. solve

$$\min_{P\succ0,\ \lambda\ge0}\ \mathrm{trace}(P)\quad\text{s.t. the two LMIs of Theorem 1}.$$

A small $\mathrm{trace}(P)=\sum_j1/r_j^2$ (with $r_j=1/\sqrt{\lambda_j(P)}$ the semi-axes) means a large ellipsoid. It is only a surrogate: it favours long axes but neither maximises the volume nor guarantees that the result contains a competing ellipsoid. The volume $\propto\det(P)^{-1/2}$ would be more natural, but minimising $\log\det P$ is minimising a concave function, whereas the trace is linear. The box size is not a convex variable: shrinking $\delta$ tightens the sectors ($\alpha$ grows) but confines the ellipsoid to a smaller box, enlarging it loosens the sectors until the first LMI becomes infeasible (their Remark 1). Yin et al. grid $\delta$; Pauli et al. (CDC 2021) bisect for the largest feasible $\delta$ and then run a golden-section search on $\mathrm{trace}(P)$. The explorer below does the same: it brackets $\delta\in[0.02,40]$, takes 18 bisection steps at the geometric midpoint $\sqrt{\delta_{lo}\delta_{hi}}$ (bisection in $\log\delta$), then 14 golden-section steps for $\mathrm{trace}(P)$ on $[0.05\,\delta_{\max},\,0.998\,\delta_{\max}]$. These are stopping rules: they locate the feasibility boundary and the best box only approximately.

Background — Bisection and golden-section search

Bisection keeps an interval $[\delta_{lo},\delta_{hi}]$ with $\delta_{lo}$ feasible and $\delta_{hi}$ infeasible, tests the midpoint and keeps the half where the switch happens, so the interval halves at every step. It finds the boundary only if feasibility is nested (every box smaller than a feasible one is feasible) and every feasibility test is reliable; here that is an assumption, not a proved property. Golden-section search minimises a function of one variable on an interval by comparing its values at two interior points and discarding the part beyond the worse one; the interval shrinks by the factor $0.618$ per step, and the method is guaranteed to approach the minimiser only if the function is unimodal (decreasing, then increasing). Unimodality of $\delta\mapsto\min\mathrm{trace}(P)$ has not been shown, so the outer search is a heuristic. Whatever box it returns, the certificate for that box is valid.

Theorem — Robust version with IQCs (Yin et al., Thm. 2, structure)
The plant becomes $x_{t+1}=Ax_t+B_1q_t+B_2u_t$, $p_t=Cx_t+D_1q_t+D_2u_t$, $q=\Delta(p)$, where the perturbation $\Delta$ (unmodelled dynamics, saturation, delay, slope-restricted nonlinearities) satisfies a hard time-domain IQC: a stable filter $\Psi_\Delta$ with zero initial state maps $(p,q)$ to $r$ and $\sum_{t=0}^{N}r_t^{\top}M_\Delta r_t\ge0$ for all $N$. Assumption 1: the equilibrium is at the origin and $\Delta(0)=0$ (and, implicitly, the uncertain interconnection is well posed: it has a unique solution). With the extended state $\zeta=(x,\psi)$ (plant and filter), the first LMI gains the filter term $[C_e\ D_e]^{\top}M_\Delta[C_e\ D_e]$ (assembled in the derivation below), the containment LMI uses the zero-padded first weight, and the conclusion is local stability for every $\Delta$ satisfying the IQC with $\mathcal E(P_x,x_*)$, $P_x$ the upper-left block of $P$, inside the robust ROA (the filter starts at $\psi_0=0$, so only that slice of the ellipsoid is relevant). The same machinery admits an off-by-one IQC on the activations, a causal Zames–Falb multiplier of order one (Lessard, Recht & Packard, 2016) that relates the pairs $(v_t,w_t)$ and $(v_{t-1},w_{t-1})$ at consecutive times and encodes the local slope bounds; its filter stores $w_{t-1}-\mathrm{diag}(\nu)v_{t-1}$ ($\nu$ the upper local slope bounds) and adds $n_\phi$ states.
Derivation — Assembling the robust LMI, and the off-by-one IQC written out

The extended system. Work in deviations (tildes dropped) and keep $s=(w,q)$ as the free signals. Substituting $u=N_{ux}x+N_{uw}w$ gives $x^+=A_cx+B_ss$ and $p=C_cx+D_ss$ with $A_c=A+B_2N_{ux}$, $B_s=[B_2N_{uw}\ \ B_1]$, $C_c=C+D_2N_{ux}$, $D_s=[D_2N_{uw}\ \ D_1]$. Write the filter as $\psi^+=A_f\psi+B_{fp}p+B_{fq}q$, $r=C_f\psi+D_{fp}p+D_{fq}q$, and let $J_w=[I\ 0]$, $J_q=[0\ I]$ pick $w$ and $q$ out of $s$. Substituting $p$ into the filter gives $\zeta^+=A_e\zeta+B_es$, $r=C_e\zeta+D_es$ and $(v,w)=F_\phi(\zeta,s)$ with

$$\begin{gathered} A_e=\begin{bmatrix}A_c&0\\B_{fp}C_c&A_f\end{bmatrix},\quad B_e=\begin{bmatrix}B_s\\B_{fp}D_s+B_{fq}J_q\end{bmatrix},\\ C_e=[D_{fp}C_c\ \ C_f],\quad D_e=D_{fp}D_s+D_{fq}J_q,\quad F_\phi=\begin{bmatrix}N_{vx}&0&N_{vw}J_w\\0&0&J_w\end{bmatrix},\\ \begin{bmatrix}A_e&B_e\end{bmatrix}^{\top}P\begin{bmatrix}A_e&B_e\end{bmatrix}-\mathrm{diag}(P,0)+\begin{bmatrix}C_e&D_e\end{bmatrix}^{\top}M_\Delta\begin{bmatrix}C_e&D_e\end{bmatrix}+F_\phi^{\top}Q_{\alpha\beta}(T)F_\phi\prec0 .\end{gathered}$$

The last inequality has size $n_x+n_\psi+n_\phi+n_q$ ($n_\psi$ filter states, $n_q$ uncertainty channels); the containment LMIs are those of Theorem 1 with the row $[(W_0)_j\ \ 0]$. Multiplying by $(\zeta_t,s_t)$ gives $V(\zeta_{t+1})-V(\zeta_t)+r_t^{\top}M_\Delta r_t+(\text{sector QC})\le-\varepsilon\|(\zeta_t,s_t)\|^2$. The IQC term is nonnegative only as a sum, so one sums over $t$ exactly as in the hard-IQC proof of Section 3 (Steps 2–4 there, including the extension argument for the local sectors); with $\psi_0=0$ this keeps $\zeta_t$ in the ellipsoid and drives it to zero.

The off-by-one IQC for one neuron. With local slope bounds $[\mu,\nu]$ put $a_t=\nu v_t-w_t$ and $b_t=w_t-\mu v_t$, a monotone relation through the origin (Section 1). Yin et al. (Sec. III-D) take the filter state $\psi_{t+1}=w_t-\nu v_t=-a_t$ with $\psi_0=0$, the output $r_t=(a_t-a_{t-1},\ b_t)$ and the middle matrix $\begin{bmatrix}0&\eta\\ \eta&0\end{bmatrix}$, $\eta\ge0$, so $r_t^{\top}Mr_t=2\eta\,b_t(a_t-a_{t-1})$ (with $a_{-1}=0$). Over a horizon $\sum_tb_t(a_t-a_{t-1})=\mathbf b^{\top}H\mathbf a$, where $H$ has $1$ on the diagonal and $-1$ on the subdiagonal: off-diagonal entries $\le0$ and all row and column sums $\ge0$, so the rearrangement lemma of Section 3 gives $\mathbf b^{\top}H\mathbf a\ge0$ for every $N$. The terms themselves can be negative: for $\varphi(v)=v/2$, $[\mu,\nu]=[0,1]$ and $(v_0,v_1)=(2,1)$ they are $1$ and $-1/4$, the partial sums $1$ and $3/4$. That is exactly the difference between a hard IQC and a pointwise QC. Summing such terms over the neurons with weights $\eta_i\ge0$ gives the $n_\phi$-state filter of the theorem.

Examples in the paper. (1) An inverted pendulum ($m=0.15$ kg, $l=0.5$ m, friction $0.5$ Nms/rad, input saturation at $0.7$ Nm, i.e. clipping to $[-0.7,0.7]$, a projection onto a box as in Primer B) with a policy-gradient controller: two tanh layers of 32 neurons, biases fixed to zero during training so that $x_*=0$, sampling time $0.02$ s. The term $\theta-\sin\theta$ is treated as a perturbation that is slope-restricted in $[0,0.2548]$ and sector-bounded in $[0,0.087]$ for $|\theta|\le0.73$, the saturation by a local sector, and the first-layer box uses $\delta=0.1$; both assumptions are checked on the resulting ellipsoid. (2) Vehicle lateral dynamics with a 32–32 tanh policy, saturation at $\pi/6$ and a norm-bounded LTI actuator uncertainty ($\|\Delta\|_\infty\le0.1$): adding the off-by-one IQC reduces $\mathrm{trace}(P_x)$ from $4.4$ to $2.9$, increases $\det(P_x^{-1})$ (proportional to the squared volume) from $3.2\times10^5$ to $1.1\times10^6$, and raises the largest feasible box from $\delta=0.67$ to $\delta=1.4$.

Complexity. The first LMI has size $n_x+n_\phi$ and, with a diagonal $T$, about $n_x(n_x+1)/2+n_\phi$ variables; the $n_1$ containment LMIs have size $1+n_x$. By the usual interior-point estimate (each Newton step costs on the order of $mn^3+m^2n^2$ floating-point operations, or flops, for $m$ variables and matrix size $n$, and about $\sqrt n$ steps are needed, times a factor that grows with the logarithm of the requested accuracy) a wide network costs roughly $n_\phi^4$ per step and $n_\phi^{4.5}$ in total, up to constants ($O(\cdot)$ notation: Primer 0; doubling $n_\phi$ multiplies $n_\phi^4$ by $16$): comfortable for hundreds of neurons, hopeless for millions, which motivates Section 8.

Caveat — What the certificate covers
It covers the model $(A,B)$ exactly (or the IQC-described family of Theorem 2), the network exactly as trained (a changed weight invalidates it), and the ellipsoid only; states outside it may or may not converge. The interval propagation is itself conservative in deep networks (boxes inflate layer by layer), which loosens the sectors of later layers. Floating-point: an SDP solver returns a numerically feasible point; one re-checks the unrounded LMIs with a margin (the explorer uses approximate eigenvalues); a formal certificate additionally needs validated rounding-error bounds.

3. Dynamic Multipliers: Acausal Zames–Falb (Pauli et al.)

A static QC uses one time instant. But an activation is memoryless and slope-restricted, so values at different times are related too: $w_t-w_{t-1}=\tilde\varphi(v_t)-\tilde\varphi(v_{t-1})$ has the slope of a chord. Pauli, Gramlich, Berberich & Allgöwer, CDC 2021 brought the full class of dynamic multipliers for slope-restricted nonlinearities to NN loops: full-block circle and Yakubovich-type multipliers (following Fetzer and Scherer) combined with acausal Zames–Falb multipliers realised by FIR filters (see Primer D).

Definition — Integral quadratic constraints (Megretski & Rantzer 1997; discrete-time form of Pauli et al., Defs. 4–5)
Signals $v,w\in\ell_2^n$ satisfy the frequency-domain IQC defined by $\Pi:\{|z|=1\}\to\mathbb C^{2n\times2n}$ if $\int_0^{2\pi}\begin{bmatrix}\hat v(e^{j\omega})\\ \hat w(e^{j\omega})\end{bmatrix}^{*}\Pi(e^{j\omega})\begin{bmatrix}\hat v(e^{j\omega})\\ \hat w(e^{j\omega})\end{bmatrix}d\omega\ge0$. With a factorisation $\Pi=\Psi^{*}M\Psi$, $\Psi$ stable, let $r=\Psi\begin{bmatrix}v\\ w\end{bmatrix}$ (zero initial filter state). The pair satisfies the soft IQC if $\sum_{t=0}^{\infty}r_t^{\top}Mr_t\ge0$ and the hard IQC if $\sum_{t=0}^{N}r_t^{\top}Mr_t\ge0$ for every $N\ge0$. An operator satisfies an IQC if all its input–output pairs do. A static QC is the special case $\Psi=I$ that holds at every $t$.
The symbols. $\ell_2^n$ is the set of sequences $(v_t)_{t\ge0}$ in $\mathbb R^n$ with finite energy $\sum_t\|v_t\|^2\lt\infty$; $\hat v(e^{j\omega})=\sum_{t\ge0}v_te^{-j\omega t}$ is the discrete-time Fourier transform (Primer D); ${}^{*}$ is the conjugate transpose, and $\Pi(e^{j\omega})$ is Hermitian, so the integrand is real. With $\Pi=\Psi^{*}M\Psi$, Parseval's theorem turns the integral into $2\pi\sum_{t\ge0}r_t^{\top}Mr_t$, which is why the soft IQC is the same condition written in time. The hard IQC is applied to finite prefixes of a trajectory, so the stability proofs below never assume in advance that the closed-loop signals have finite energy.
Theorem — IQC stability theorem (Megretski & Rantzer, TAC 1997)
Let $G$ be a stable LTI system and $\Delta$ a bounded causal operator. Assume (i) the loop $v=Gw+f$, $w=\Delta(v)+e$ is well posed for $\tau\Delta$, every $\tau\in[0,1]$; (ii) $\tau\Delta$ satisfies the IQC defined by $\Pi$ for every $\tau\in[0,1]$; (iii) there is $\varepsilon\gt0$ with $$\begin{bmatrix}G(j\omega)\\ I\end{bmatrix}^{*}\Pi(j\omega)\begin{bmatrix}G(j\omega)\\ I\end{bmatrix}\preceq-\varepsilon I\qquad\forall\omega.$$ Then the loop is stable in the input–output sense: for zero initial conditions there is a $c\lt\infty$ with $\|(v,w)\|_{2,[0,T]}\le c\,\|(f,e)\|_{2,[0,T]}$ for every horizon $T$ (Megretski & Rantzer, TAC 42(6), 1997; stated in continuous time there, with $G(j\omega)$ and time integrals; the discrete-time version uses $G(e^{j\omega})$ and sums). This is a gain bound, not a region-of-attraction statement. In words: if the linear part maps every signal pair into the region where the IQC is violated, the only pair consistent with both is the zero pair, robustly. Why the assumptions: the homotopy in $\tau$ (from the trivially stable loop $\tau=0$ to the real one) is how the proof works with soft IQCs; well-posedness makes the loop define signals at all.
Caveat — Two proof routes, two sets of hypotheses
The homotopy proof above works with soft IQCs and a frequency-domain inequality. Yin et al. and Pauli et al. follow the other route, surveyed by Scherer, IEEE CSM 2022: a hard factorised IQC, a storage function for the plant plus filter states, and a dissipation inequality summed over time. That route needs the hard property (the inequality on every finite horizon), and is what makes local, region-of-attraction statements possible. Do not mix the hypotheses of one route with the conclusion of the other.
Definition — Discrete-time acausal Zames–Falb multipliers with FIR factorisation (Pauli et al., Sec. II-D)
For monotone nonlinearities with $\varphi(0)=0$, i.e. slope restriction $[0,\infty]$ (Pauli et al. write "the sector $[0,\infty]$, which corresponds to monotonicity"; a sector bound alone would not do, because the cross terms compare values at different times or of different neurons), take FIR coefficients $M_{-\ell},\dots,M_{\ell}\in\mathbb R^{n\times n}$ and the block matrix $$\tilde P=\begin{bmatrix}M_0 & M_{-1} & \cdots & M_{-\ell}\\ M_1 & 0 & \cdots & 0\\ \vdots & \vdots & \ddots & \vdots\\ M_\ell & 0 & \cdots & 0\end{bmatrix},\qquad \Pi^{ZF}_{[0,\infty]}=\Big\{\begin{bmatrix}0 & \tilde P^{\top}\\ \tilde P & 0\end{bmatrix}\Big\},$$ acting on the window $(v_t,\dots,v_{t-\ell},w_t,\dots,w_{t-\ell})$, i.e. with the FIR filter $$\Psi^{ZF}(z)=\mathrm{blkdiag}\big([I,z^{-1}I,\dots,z^{-\ell}I]^{\top},\ [I,z^{-1}I,\dots,z^{-\ell}I]^{\top}\big).$$ The coefficients must satisfy the row and column sum conditions $$\Big(\sum_{j=-\ell}^{\ell}M_j\Big)\mathbf 1\ge0,\qquad \mathbf 1^{\top}\Big(\sum_{j=-\ell}^{\ell}M_j\Big)\ge0,$$ together with sign conditions: all entries except the diagonal of $M_0$ are nonpositive; for odd nonlinearities these are replaced by absolute-value dominance conditions of the diagonal of $M_0$ over all other entries. Slope bounds $[\mu,\nu]$ follow by the loop transformation of Section 1: $\Pi^{ZF}_{[\mu,\nu]}=L_{\mu\nu}^{\top}\Pi^{ZF}_{[0,\infty]}L_{\mu\nu}$. Per-neuron bounds $\mu,\nu\in\mathbb R^n$ are fine for diagonal $M_j$; an entry of $M_j$ that couples neurons $i$ and $k$ also needs $\mu_i=\mu_k$ and $\nu_i=\nu_k$ (theorem below).
The odd case. $\varphi$ is odd if $\varphi(-s)=-\varphi(s)$. Then the sign conditions are replaced by: for every channel $i$ (writing $m_{j,ik}$ for the entries of $M_j$) $$\begin{gathered}m_{0,ii}\ge\sum_{k\ne i}|m_{0,ik}|+\sum_{j\ne0}\sum_{k}|m_{j,ik}|,\\ m_{0,ii}\ge\sum_{k\ne i}|m_{0,ki}|+\sum_{j\ne0}\sum_{k}|m_{j,ki}|,\end{gathered}$$ i.e. each diagonal entry of $M_0$ dominates the absolute sum of all other entries in its row, and in its column, over all channels and delays; off-diagonal entries of either sign are then allowed. (Pauli et al. print the second sum over all $j$; with $j=0$ included the entries of $M_0$ would be counted twice, so we read it as $j\ne0$.) The loop-transformed channels of an odd $\varphi$ are odd again, but a shifted activation $\tilde\varphi(s)=\tanh(v_*+s)-\tanh(v_*)$ with $v_*\ne0$ is not odd, so this larger class is available only for equilibria at $v_*=0$ (e.g. bias-free networks with $x_*=0$).

Written out, the quadratic form is $2w_t^{\top}(M_0v_t+M_{-1}v_{t-1}+\dots+M_{-\ell}v_{t-\ell})+2\sum_{j=1}^{\ell}w_{t-j}^{\top}M_jv_t$. The multiplier $\sum_jM_jz^{-j}$ contains positive powers of $z$ (future values), which is what acausal means; the factorisation nevertheless uses only delays, because the future terms are moved to the other signal: "$v_t$ against a past $w$" plays the role of "$w$ against a future $v$". The IQC is a property of the signals, not a filter one implements, so non-causality costs nothing.

Theorem — Hard IQC for acausal Zames–Falb multipliers (Pauli et al., CDC 2021, Thm. 1)
Let $\tilde\phi:\mathbb R^n\to\mathbb R^n$ be a diagonally repeated nonlinearity (the same scalar function in every channel) that vanishes at zero and is slope-restricted in $[\mu,\nu]$ with scalar bounds $\mu,\nu$ common to all channels, and let $\Pi\in\Pi^{ZF}_{[\mu,\nu]}$ with middle matrix $M$. Then all $v,w\in\ell_2^n$ with $w_t=\tilde\phi(v_t)$ (and $v_t=w_t=0$ for $t\lt0$) satisfy $$\sum_{t=0}^{N}\begin{bmatrix}v_t\\ \vdots\\ v_{t-\ell}\\ w_t\\ \vdots\\ w_{t-\ell}\end{bmatrix}^{\top}M\begin{bmatrix}v_t\\ \vdots\\ v_{t-\ell}\\ w_t\\ \vdots\\ w_{t-\ell}\end{bmatrix}\ge0\qquad\forall N\in\mathbb N_0,$$ a hard IQC with the FIR filter $\Psi^{ZF}$. If $\tilde\phi$ is not diagonally repeated, or its slope bounds $\mu_i,\nu_i$ differ between neurons, the statement holds with diagonal $M_j=\mathrm{diag}(m_j)$; block-diagonal $M_j$ (ZF-RL) need the repeated function and the common bounds only within each block.
In words: over every finite horizon, the multiplier-weighted products of current and delayed inputs and outputs of the nonlinearity add up to something nonnegative. Why the assumptions: slope restriction (not only a sector) is what makes products at different times valid; the repeated structure is what allows full $M_j$ to couple different neurons (derivation below, Exercise 14.6), and after the loop transformation the channels are repeated only if they share $(\mu,\nu)$. The paper writes the transformation with per-neuron vectors $\mu,\nu$; with coupled $M_j$ that is not enough: for $\varphi(s)=s/2$ in two channels, $\mu=(1/4,0)$, $\nu=(1,3/4)$ and $M_0=\begin{bmatrix}1 & -1\\ -1 & 1\end{bmatrix}$ (all other $M_j=0$), the pair $v=(1,1)$, $w=(1/2,1/2)$ gives the value $-1/8$ at a single time step. The zero values before $t=0$ (zero initial filter state) give the finite-horizon, hard version that the Lyapunov argument needs.
Derivation — Why the conditions work: doubly hyperdominant matrices and rearrangement

A matrix is doubly hyperdominant if its off-diagonal entries are nonpositive and its row and column sums are nonnegative. The key lemma (Willems and Brockett, as used by Pauli et al., Lemma 2): if $\varphi$ is monotone ($\mathrm{slope}[0,\infty]$) with $\varphi(0)=0$ and $\mathsf M$ is doubly hyperdominant, the repeated map $\phi$ satisfies $v^{\top}\mathsf M\phi(v)\ge0$ for all $v$.

The $2\times2$ case shows both ingredients. Take $\mathsf M=\begin{bmatrix}m & -c\\ -c & m\end{bmatrix}$ with $m\ge c\ge0$. Then

$$v^{\top}\mathsf M\phi(v)=(m-c)\big[v_1\varphi(v_1)+v_2\varphi(v_2)\big]+c\,\big(v_1-v_2\big)\big(\varphi(v_1)-\varphi(v_2)\big).$$

The first bracket is $\ge0$ by the sector property ($v\varphi(v)\ge0$); the second product is $\ge0$ by monotonicity of one and the same function evaluated at $v_1$ and $v_2$ (a rearrangement inequality). If the two channels carried different functions $\varphi_1\ne\varphi_2$, the second term would read $(v_1-v_2)(\varphi_1(v_1)-\varphi_2(v_2))$, which can be negative: that is why this monotonicity-based coupling of different neurons needs a repeated nonlinearity, while coupling the same neuron at different times (the off-diagonal time structure of $\tilde P$) is always fine for a time-invariant activation.

Hard IQC. Stack the finite sequence $(v_0,\dots,v_N)$ and $(w_0,\dots,w_N)$. The sum over $t$ of the FIR quadratic form equals $\mathbf v^{\top}\mathsf M_N\mathbf w$ for a banded block-Toeplitz matrix $\mathsf M_N$ (constant along its block diagonals, Primer A) built from the $M_j$; conditions (row/column sums and signs) make $\mathsf M_N$ doubly hyperdominant for every $N$, and the lemma applied to the long vector gives nonnegativity for every horizon, i.e. a hard IQC (Pauli et al., Appendix B).

Example: one channel, one delay. Summing the written-out form over $t=0,1,2$ (with $v_{-1}=w_{-1}=0$) gives $2\mathbf w^{\top}H\mathbf v$ with $H=\begin{bmatrix}m_0&m_1&0\\ m_{-1}&m_0&m_1\\ 0&m_{-1}&m_0\end{bmatrix}$, so $\mathsf M_N=2H^{\top}$. Every interior row and column contains $m_{-1}+m_0+m_1\ge0$; the first and last ones miss an off-diagonal entry, which is $\le0$, so their sums are even larger. Truncating to a finite horizon therefore keeps $\mathsf M_N$ doubly hyperdominant (in the odd case it only removes terms from the absolute sums). The zero padding before $t=0$ is legitimate because the shifted nonlinearity passes through zero.

Slope bounds $[\mu,\nu]$. The lemma is applied to the loop-transformed channels $a_i=\nu_iv_i-w_i$, $b_i=w_i-\mu_iv_i$. For a chord of slope $s\in[\mu_i,\nu_i]$, $\Delta a_i\,\Delta b_i=(\nu_i-s)(s-\mu_i)\Delta v_i^2\ge0$, so each $a_i\mapsto b_i$ is monotone (possibly multi-valued), and it passes through $0$ when $\varphi(0)=0$. The rearrangement step for coupled neurons $i,k$ needs these two relations to be the same. For a repeated $\varphi$ that is guaranteed when $\mu_i=\mu_k$ and $\nu_i=\nu_k$, and fails in general otherwise. Per-neuron local bounds from interval propagation usually differ, so a coupled multiplier needs one common pair of bounds, valid on the common interval of the group, as in Step 4 of the proof sketch below. The counterexample in the theorem's "why" line fails already at a single time step.

Definition — Full-block circle, circle–Yakubovich and combined multipliers (Pauli et al., Sec. II-C and II-E)
Let $\tilde\phi$ vanish at zero and lie in $\mathrm{sec}[\alpha,\beta]\cap\mathrm{slope}[\mu,\nu]$ channel by channel (different channels may carry different functions). The sector says that $w_i=\delta_iv_i$ for some $\delta_i\in[\alpha_i,\beta_i]$, i.e. $w=\Delta v$ with a diagonal $\Delta\in[\alpha,\beta]$ that may change with time. The full-block circle class is $$\begin{gathered}\Pi_{[\alpha,\beta]}=\Big\{\Pi\in\mathbb S^{2n}:\ \begin{bmatrix}I\\ \Delta\end{bmatrix}^{\top}\Pi\begin{bmatrix}I\\ \Delta\end{bmatrix}\succeq0\\ \text{for every diagonal }\Delta\in[\alpha,\beta]\Big\},\end{gathered}$$ and every $\Pi$ in it gives $\begin{bmatrix}v_t\\ w_t\end{bmatrix}^{\top}\Pi\begin{bmatrix}v_t\\ w_t\end{bmatrix}=v_t^{\top}\begin{bmatrix}I\\ \Delta_t\end{bmatrix}^{\top}\Pi\begin{bmatrix}I\\ \Delta_t\end{bmatrix}v_t\ge0$ at every time: a static QC ($\Psi=I$), hence a hard IQC. It contains the diagonal multipliers $Q_{\alpha\beta}(T)$ of Section 2, and its off-diagonal entries are justified by the "for every $\Delta$" condition, not by a repeated nonlinearity. Slope restriction adds $w_t-w_{t-1}=\Delta'_t(v_t-v_{t-1})$ with a diagonal $\Delta'_t\in[\mu,\nu]$. With the filter output $r^{\rm cy}_t=(v_t,\ v_t-v_{t-1},\ w_t,\ w_t-w_{t-1})$ (filter $1-z^{-1}$, zero initial state) the circle–Yakubovich class $\Pi^{\rm cy}\subset\mathbb S^{4n}$ is defined by the same condition with $\mathrm{diag}(\Delta,\Delta')$ ranging over the box $[(\alpha,\mu),(\beta,\nu)]$; again $(r^{\rm cy}_t)^{\top}\Pi\,r^{\rm cy}_t\ge0$ at every $t$. The combined class stacks the filters, $\Psi^{\rm comb}=\begin{bmatrix}\Psi^{ZF}\\ \Psi^{\rm cy}\end{bmatrix}$, with $\Pi^{\rm comb}=\mathrm{blkdiag}(\Pi^{ZF},\Pi^{\rm cy})$, each block satisfying its own conditions (for coupled $M_j$ the repeated-channel conditions above): a sum of two hard IQCs is a hard IQC. Local bounds again need the containment LMI and the extension argument of Step 4 below.
Making "for every $\Delta$" finite (a standard sufficient test; the paper does not spell out its implementation): write $\Pi=\begin{bmatrix}\Pi_{vv}&\Pi_{vw}\\ \Pi_{vw}^{\top}&\Pi_{ww}\end{bmatrix}$ and require $\Pi_{ww}\preceq0$. For fixed $v$ the form $v^{\top}[I;\Delta]^{\top}\Pi[I;\Delta]v$ is then a concave quadratic function of the diagonal entries of $\Delta$ (its Hessian is $2\,\mathrm{diag}(v)\Pi_{ww}\mathrm{diag}(v)\preceq0$), so it is smallest at a vertex of the box, and checking the LMI at the $2^n$ vertex matrices $\Delta$ (at $2^{2n}$ for $\Pi^{\rm cy}$) suffices. The number of vertices grows exponentially with $n$.

With the IQCs in hand, the stability test is a dissipation inequality for the plant plus filter states. Let $\eta=(\tilde x,\xi)$ stack the plant state and the state $\xi$ of the filters $\Psi$ driven by $(\tilde v,\tilde w)$, with realisation $\eta_{t+1}=A_{\rm tot}\eta_t+B_{\rm tot}\tilde w_t$, $r_t=C_{\rm tot}\eta_t+D_{\rm tot}\tilde w_t$.

Derivation — The plant-plus-filter realisation

Write the (stacked) filters as $\xi_{t+1}=A_\psi\xi_t+B_{\psi v}\tilde v_t+B_{\psi w}\tilde w_t$, $r_t=C_\psi\xi_t+D_{\psi v}\tilde v_t+D_{\psi w}\tilde w_t$, and substitute the loop equations $\tilde u=N_{ux}\tilde x+N_{uw}\tilde w$, $\tilde v=N_{vx}\tilde x+N_{vw}\tilde w$ and $\tilde x_{t+1}=A\tilde x_t+B\tilde u_t$:

$$A_{\rm tot}=\begin{bmatrix}A+BN_{ux}&0\\ B_{\psi v}N_{vx}&A_\psi\end{bmatrix},\quad B_{\rm tot}=\begin{bmatrix}BN_{uw}\\ B_{\psi v}N_{vw}+B_{\psi w}\end{bmatrix},\quad C_{\rm tot}=\begin{bmatrix}D_{\psi v}N_{vx}&C_\psi\end{bmatrix},\quad D_{\rm tot}=D_{\psi v}N_{vw}+D_{\psi w}.$$

For the FIR filters of this section these four matrices are fixed (shift registers); the multiplier coefficients sit only in the middle matrix $M$, which is why the LMI below is linear in $(P,M)$.

Theorem — Stability and ROA with general hard IQCs (Pauli et al., CDC 2021, Thm. 2 and Cor. 1; our notation)
Assumptions. The shifted network $\tilde\phi$ satisfies $\tilde\phi\in\mathrm{slope}[\mu,\nu]\cap\mathrm{sec}[\alpha,\beta]$ for all inputs in a box $[\underline d,\bar d]$, obtained by interval propagation from a first-layer box $\tilde v^1\in[-d^1,d^1]$. $M$ belongs to the chosen multiplier class (e.g. diagonal circle plus ZF, or the combination $\Pi^{\rm comb}$ of full-block and ZF classes). With $Q:=W_0C$ (the network input is $y=Cx$; $C=I$ for state feedback) and rows $Q_j$, there exist $P=\begin{bmatrix}P_x & P_{x\xi}\\ P_{\xi x} & P_\xi\end{bmatrix}\succ0$ and $M$ with $$\begin{bmatrix}I & 0\\ A_{\rm tot} & B_{\rm tot}\\ C_{\rm tot} & D_{\rm tot}\end{bmatrix}^{\top}\begin{bmatrix}-P & 0 & 0\\ 0 & P & 0\\ 0 & 0 & M\end{bmatrix}\begin{bmatrix}I & 0\\ A_{\rm tot} & B_{\rm tot}\\ C_{\rm tot} & D_{\rm tot}\end{bmatrix}\prec0,$$ $$\begin{bmatrix}(d^1_j)^2 & Q_j & 0\\ Q_j^{\top} & P_x & P_{x\xi}\\ 0 & P_{\xi x} & P_\xi\end{bmatrix}\succeq0\ \ (j=1,\dots,n_1).$$ Statement. For every $\tilde x_0\in\mathcal E(P_x,0)$ the loop is locally asymptotically stable, and $\mathcal E(P_x,x_*)$ is an inner approximation of the ROA of the original loop. Theorem 1 of Yin et al. is the special case of static multipliers.
In words: a quadratic storage for the plant and filter states that dissipates against the IQC supply, plus an ellipsoid that keeps the first layer inside its box, certifies that ellipsoid. Why the assumptions: the slope and sector bounds hold only inside the box, hence the containment LMI; the multiplier class must be valid for the shifted, in general non-repeated nonlinearity (coupled $M_j$ only within groups of neurons that carry the same shifted function and share their slope bounds, e.g. a bias-free network with $\varphi(0)=0$ at $x_*=0$ and one common interval per group; diagonal $M_j$ otherwise); the hard IQC requires the filter to start at $\xi_0=0$, which is why only that slice of $\mathcal E(P,0)$ is certified.
Proof sketch — Summed dissipation, and why "local" needs one more argument

Step 1. Multiply the first LMI by $(\eta_t,\tilde w_t)$: with $\eta_{t+1}=A_{\rm tot}\eta_t+B_{\rm tot}\tilde w_t$ and $r_t=C_{\rm tot}\eta_t+D_{\rm tot}\tilde w_t$ this is $\eta_{t+1}^{\top}P\eta_{t+1}-\eta_t^{\top}P\eta_t+r_t^{\top}Mr_t\le-\varepsilon\|\eta_t\|^2$.

Step 2 (sum, then use the hard IQC). Summing over $t=0,\dots,N$ gives $\eta_{N+1}^{\top}P\eta_{N+1}+\varepsilon\sum_t\|\eta_t\|^2\le\eta_0^{\top}P\eta_0-\sum_{t=0}^Nr_t^{\top}Mr_t\le\eta_0^{\top}P\eta_0$. With dynamic multipliers $V(\eta_t)$ need not decrease at every step; only $V(\eta_t)\le V(\eta_0)$ is guaranteed. That is enough to keep a trajectory that starts with $\xi_0=0$ and $\eta_0^{\top}P\eta_0\le1$ in the ellipsoid, and with the summable $\|\eta_t\|^2$ it gives convergence. It does not make the augmented ellipsoid invariant for restarts with a nonzero filter state: the hard IQC needs $\xi_0=0$. So the certified initial states are the slice $\{\tilde x_0:\ \tilde x_0^{\top}P_x\tilde x_0\le1\}$ of the ellipsoid at $\xi=0$, not its (larger) projection onto the $x$-coordinates.

Step 3 (box). The Schur complement of the second LMI gives $(Q_j\tilde x_t)^2\le(d^1_j)^2\eta_t^{\top}P\eta_t\le(d^1_j)^2\eta_0^{\top}P\eta_0\le(d^1_j)^2$ when $\xi_0=0$ and $\tilde x_0\in\mathcal E(P_x,0)$ (because then $\eta_0^{\top}P\eta_0=\tilde x_0^{\top}P_x\tilde x_0$).

Step 4 (local validity). A hard IQC is a statement about whole sequences, so the one-step induction of Section 2 is not available. Pauli et al. first assume the bounds hold globally and then argue that behaviour outside the box is irrelevant. One way to make this precise: replace $\tilde\varphi_i$ outside its interval $[\underline d_i,\bar d_i]$ by an affine extension with slope $\alpha_i$ (the tight lower sector bound, which lies in $[\mu_i,\nu_i]$ because it is a chord slope); every new chord slope and every new ratio $\hat\varphi_i(s)/s$ is a convex combination of old ones and $\alpha_i$, so the extension satisfies the bounds globally (for the repeated classes use one common interval, and hence common bounds, for all coupled neurons, so that the extension is again repeated); the theorem applies to it, trajectories from $\mathcal E$ never leave the box (Step 3), and on the box the extension equals the true activation, so these trajectories are trajectories of the real loop.

Why the extension keeps the bounds. Write the interval as $[\underline d,\bar d]\ni0$ and set $\hat\varphi(s)=\varphi(\bar d)+\alpha(s-\bar d)$ for $s\gt\bar d$, $\hat\varphi(s)=\varphi(\underline d)+\alpha(s-\underline d)$ for $s\lt\underline d$, and $\hat\varphi=\varphi$ in between. For $s\gt\bar d\gt0$, $\hat\varphi(s)/s=\frac{\bar d}{s}\cdot\frac{\varphi(\bar d)}{\bar d}+\big(1-\frac{\bar d}{s}\big)\alpha$, a convex combination of an old sector ratio and $\alpha$, hence in $[\alpha,\beta]$. A chord from $s_1\in[\underline d,\bar d]$ to $s_2\gt\bar d$ has slope $\frac{\bar d-s_1}{s_2-s_1}\cdot\frac{\varphi(\bar d)-\varphi(s_1)}{\bar d-s_1}+\frac{s_2-\bar d}{s_2-s_1}\,\alpha$, a convex combination of an inside chord slope and $\alpha$, hence in $[\mu,\nu]$. Chords with both ends outside on one side have slope $\alpha$, a chord across the whole interval is a convex combination of the same kind, and the left side is symmetric. $\blacksquare$

How much structure to exploit. The multiplier class is a dial between conservatism and cost (Pauli et al., Table I; $n$ neurons, FIR orders $\ell_-,\ell_+$):

Multiplier classEncodesValid forDecision variables
Diagonal circle $\Pi^d_{[\alpha,\beta]}$sectorany activations$n$
Full-block circle $\Pi_{[\alpha,\beta]}$sectorany activations$2n(2n+1)/2$
Circle + Yakubovich $\Pi^{cy}$sector and slope (one-step differences)any activations$4n(4n+1)/2$
ZF, diagonal $M_j$slope, across timeany activations$(\ell_-+\ell_++1)n$
ZF-RL, block-diagonal by layerslope, across time and neurons of a layerrepeated within each layer, one slope bound per layer (e.g. bias-free, $\varphi(0)=0$, $x_*=0$)$(\ell_-+\ell_++1)\sum_in_i^2$
ZF-R, full $M_j$slope, across time and all neuronsrepeated, one slope bound for all neurons (e.g. bias-free, $\varphi(0)=0$, $x_*=0$)$(\ell_-+\ell_++1)n^2$

Numbers from the paper. (1) A linearised inverted pendulum (see Primer D) with a bias-free $5$–$5$ tanh network trained on MPC data ($|u|\le1$): diagonal ZF ($\ell=1$) barely improves on the diagonal circle criterion, while the repeated classes ZF-RL and ZF-R give markedly smaller $\mathrm{trace}(P_x)$, i.e. larger certified ellipses. (2) The vehicle example of Yin et al. (a $32$–$32$ tanh policy with saturation and LTI uncertainty):

Multiplier$\min_\delta\mathrm{trace}(P_x)$$\delta_{\max}$decision variablessolve time
diagonal circle (static)3.8420.67980.88 s
causal ZF, $\ell=1$ (= off-by-one IQC)2.7261.472,754123.9 s
acausal ZF, $\ell=1$2.6961.519,4422,095 s
Key insight
Dynamic multipliers roughly double the admissible box and shrink the trace by 30% here; the step from causal to acausal is small in this example but costs 17 times the computation. The large gains come from encoding more true structure: slope information across time, and, when the network is bias-free with $\varphi(0)=0$ and analysed at the origin, the fact that all neurons share one function (used with common slope bounds). The price is that each extra piece of structure is only valid under its own hypothesis.
Caveat — Check every multiplier condition against its hypothesis
Zames–Falb conditions differ between monotone and odd nonlinearities (sign constraints versus absolute-value dominance), between continuous time ($\|h\|_1\lt m_0$) and the discrete FIR version above (row and column sums), and between causal and acausal classes. The repeated classes ZF-R and ZF-RL couple neurons, which is valid only between neurons that carry the same function and use the same slope bounds. After a shift to a nonzero equilibrium the neurons generally carry different functions (neurons with equal $v_{*,i}$ still share one), and per-neuron local slope bounds break the repetition as well; between such neurons only diagonal $M_j$ remain valid. The IQCs of this section are also only local (they need the box), and the Lyapunov argument needs them hard.

4. Offset-Free Setpoint Tracking With NN Controllers (Pauli et al.)

Stabilising one equilibrium is rarely the job; a controller must bring the output to a setpoint $r$, and every $r$ has its own equilibrium. Section 2 would have to be redone for each of them, and Yin et al. fixed the biases to zero precisely to keep $x_*=0$. Pauli, Köhler, Berberich, Koch & Allgöwer, L4DC 2021 add the oldest trick in control, an integrator, and obtain certificates that cover whole sets of references.

Definition — NN controller with integral action (Pauli et al., L4DC 2021, Sec. 2.1)
Plant $x_{t+1}=Ax_t+Bu_t$, $y_t=Cx_t$ with as many inputs as outputs ($n_u=n_r$). Controller $$\xi_{t+1}=\xi_t+r-y_t,\qquad u_t=k_\xi\,\xi_t+\kappa(x_t,r),$$ where $k_\xi$ is invertible and $\kappa$ is an $\ell$-layer network with input $w^0=H_xx+H_rr$ (state feedback: $H_x=I$, $H_r=0$; output-error feedback: $H_x=-C$, $H_r=I$). The augmented state $\chi=(x,\xi)$ obeys $$\begin{gathered}\chi_{t+1}=\tilde A\chi_t+\tilde B\,\kappa(x_t,r)+B_rr,\qquad y_t=\tilde C\chi_t,\\ \tilde A=\begin{bmatrix}A & Bk_\xi\\ -C & I\end{bmatrix},\ \tilde B=\begin{bmatrix}B\\ 0\end{bmatrix},\ B_r=\begin{bmatrix}0\\ I\end{bmatrix},\ \tilde C=\begin{bmatrix}C & 0\end{bmatrix}.\end{gathered}$$ Standing assumption: $A_a=\begin{bmatrix}A-I & B\\ C & 0\end{bmatrix}$ is square and nonsingular, so each $r$ has a unique steady state $\begin{bmatrix}x_*\\ u_*\end{bmatrix}=A_a^{-1}\begin{bmatrix}0\\ r\end{bmatrix}$ with $y_*=r$; then $\xi_*=k_\xi^{-1}\big(u_*-\kappa(x_*,r)\big)$ and $v_*,w_*$ follow by propagating $(x_*,r)$ through the network.
Why the integrator makes zero offset structural
At any equilibrium of the augmented loop $\xi_{t+1}=\xi_t$, hence $r-y=0$. Nothing about the weights and biases of $\kappa$ entered this argument; nor would a model error or a constant input disturbance, as long as the loop still converges to some equilibrium. The certificate therefore only has to prove convergence, and zero offset follows. Without the integrator, zero offset would require $\kappa(x_*(r),r)=u_*(r)$ exactly for every $r$, an equation a trained network satisfies only approximately and that any model error breaks (Exercise 14.5). As the paper puts it, no structural constraints on the weights and biases are needed, because any undesired offset is eliminated by the integrator.

In error coordinates $\tilde\chi=\chi-\chi_*$ the reference drops out of the linear part: $\tilde\chi_{t+1}=\tilde A\tilde\chi_t+\tilde B\big(\kappa(x_t,r)-\kappa(x_*,r)\big)$ and $y_t-r=\tilde C\tilde\chi_t$, since $B_rr$ is the same at every time. It still enters through the equilibrium $v_*(r)$ at which the activations are shifted. The activations are described by the slope QC around the equilibrium, $\alpha\le\frac{\varphi(v_{i})-\varphi(v_{*,i})}{v_i-v_{*,i}}\le\beta$ with $0\le\alpha\lt\beta$, stacked with a diagonal multiplier; the paper points out that the richer coupled multipliers once claimed for LipSDP are invalid, by the counterexample of Pauli et al., L-CSS 2022 (see Module 12).

Theorem — Global exponential stability and offset-free tracking (Pauli et al., L4DC 2021, Thm. 1)
Let $\tilde\kappa=\kappa(x,r)-\kappa(x_*,r)$ be the deviation of the network output, let $R_V$ map $(\tilde\chi,\tilde w)\mapsto(\tilde\chi,\tilde\kappa)$ and $R_\phi$ map $(\tilde\chi,\tilde w)\mapsto(\tilde v,\tilde w)$, built from the weights as in Section 1 (the integrator state does not enter the network). The input of $\tilde B$ is $\tilde\kappa$, not the total deviation $k_\xi\tilde\xi+\tilde\kappa$ of $u$: the integral feedback is already inside $\tilde A$ and would otherwise be counted twice. If there exist a diagonal $T\succeq0$ and $P\succ0$ with $$R_V^{\top}\begin{bmatrix}\tilde A^{\top}P\tilde A-P & \tilde A^{\top}P\tilde B\\ \tilde B^{\top}P\tilde A & \tilde B^{\top}P\tilde B\end{bmatrix}R_V+R_\phi^{\top}\begin{bmatrix}-2\alpha\beta T & (\alpha+\beta)T\\ (\alpha+\beta)T & -2T\end{bmatrix}R_\phi\ \prec\ 0,$$ then for every initial condition and every constant reference $r$ the steady state $\chi_*(r)$ is globally exponentially stable and $y_t\to r$.
Assumptions used: $A_a$ square and nonsingular (one steady state per $r$), $k_\xi$ invertible, $n_u=n_r$, and every activation slope-restricted on all of $\mathbb R$ with bounds $0\le\alpha\lt\beta$ (the paper uses one pair $[\alpha,\beta]$ as slope bounds and as the sector bounds they imply around every equilibrium). In words: one quadratic Lyapunov shape works around every equilibrium that a reference can induce, and the integrator turns convergence into zero offset.
Why one LMI covers all references. Slope restriction is incremental: the sector $[\alpha,\beta]$ holds around every point of the graph, so the QC around $v_*(r)$ has the same matrix for every $r$, and the LMI contains no $r$. The Lyapunov function $V_r(\chi)=(\chi-\chi_*(r))^{\top}P(\chi-\chi_*(r))$ changes with $r$, its matrix does not.

Networks trained on data from a bounded region are rarely globally stabilising (in the paper's example Theorem 1 is infeasible), so the paper localises the analysis in two ways.

Theorem — Local stability for one reference and for a set of references (Pauli et al., L4DC 2021, Thms. 3–4)
One reference. Choose $d\gt0$, the box $v^1\in[v^1_*-d,\ v^1_*+d]$, propagate it and compute tight local bounds $D_\alpha,D_\beta$. If the LMI of Theorem 1 holds with $D_\alpha,D_\beta$ in place of $\alpha,\beta$ and $\begin{bmatrix}d_j^2 & (N_0^1)_j\\ (N_0^1)_j^{\top} & P\end{bmatrix}\succeq0$ for every first-layer neuron $j$ (with $v^1-v^1_*=N^1_0\tilde\chi$), then for all $\chi_0\in\mathcal E(P,\chi_*)$ the steady state is exponentially stable and $y_t\to r$.
A set of references. The steady-state map is linear, $x_*(r)=\Gamma r$ (from $A_a^{-1}$). Centre the box at $v^1_*(r_{\rm nom})$ for a nominal $r_{\rm nom}$ and compute the bounds there. In the coordinates $e=\chi-\chi_*(r)$, $s=r-r_{\rm nom}$ the first-layer deviation is affine, $v^1-v^1_*(r_{\rm nom})=Le+Js$ with $L=W_0H_x[I\ \ 0]=N^1_0$ and $J=W_0(H_x\Gamma+H_r)$ (derivation below). If the same LMI holds and, for some $Q\succ0$, $$\begin{bmatrix}d_j^2&L_j&J_j\\ L_j^{\top}&P&0\\ J_j^{\top}&0&Q\end{bmatrix}\succeq0\qquad\text{for every first-layer neuron } j,$$ then for every pair $(\chi_0,r)$ in the joint certified set $$\begin{aligned}\mathcal E_{P,Q}(r_{\rm nom})=\Big\{(\chi,r):\ &(\chi-\chi_*(r))^{\top}P(\chi-\chi_*(r))\\ &+(r-r_{\rm nom})^{\top}Q(r-r_{\rm nom})\le1\Big\}\end{aligned}$$ the loop with the constant reference $r$ stays in its slice $\{\chi:(\chi,r)\in\mathcal E_{P,Q}\}$, converges exponentially to $\chi_*(r)$, and $y_t\to r$ (the local bounds must hold around every $v_*(r)$ involved; second remark below).
Why the assumptions: the tight local bounds (a larger lower bound than the global $0$) are what make the first LMI feasible where Theorem 1 fails, but they are true only on the box; the containment LMI keeps the state (and, in the second version, the reference) where they are true, and the proof is the one-step induction of Section 2. $Q\succ0$ makes the admissible set of references bounded, and $Q$ trades the size of the admissible initial states against the distance of $r$ from $r_{\rm nom}$.
Derivation — The joint containment LMI and the shape of the certified set

From $v^1=W_0(H_xx+H_rr)+b_0$ and $x_*(r)-x_*(r_{\rm nom})=\Gamma s$, $$v^1-v^1_*(r_{\rm nom})=W_0H_x\big(x-x_*(r)\big)+W_0H_x\Gamma s+W_0H_rs=Le+Js .$$ The Schur complement of the LMI is $L_jP^{-1}L_j^{\top}+J_jQ^{-1}J_j^{\top}\le d_j^2$, and Cauchy–Schwarz with the weight $\mathrm{blkdiag}(P,Q)$, as in Section 2, gives $|L_je+J_js|\le\big(L_jP^{-1}L_j^{\top}+J_jQ^{-1}J_j^{\top}\big)^{1/2}\big(e^{\top}Pe+s^{\top}Qs\big)^{1/2}\le d_j$ on $\mathcal E_{P,Q}$. At $e=0$ this also keeps every equilibrium input $v^1_*(r)$ of the set inside the box. For state feedback ($H_r=0$) the LMI is the paper's (9b), whose $M$ is our $\Gamma$. The set is an ellipsoid in $(e,s)$, but not in $(\chi,r)$: $x_*(r)$ is linear in $r$, while $\xi_*(r)=k_\xi^{-1}\big(u_*(r)-\kappa(x_*(r),r)\big)$ contains the network, so $(\chi,r)$-space holds a curved family of ellipsoidal slices. The LMIs stay convex because they are written in $(e,s)$; the governor problem below, a search over $r$, is in general not convex.

Two remarks make these statements usable. First, for output-error feedback the first-layer input at steady state is $W_0(r-Cx_*)+b_0=b_0$ for every $r$, so the bounds are the same for all references and the one-reference theorem already certifies every $r$ with an ellipsoid of fixed shape centred at $\chi_*(r)$: local stability depends only on the tracking error. Second (our remark; the paper computes the bounds for $r_{\rm nom}$ and leaves this implicit), for a set of references the sector bounds must hold around every equilibrium $v_*(r)$ in the set, not only around $v_*(r_{\rm nom})$; slope bounds over the box (smallest and largest derivative) have this property, because the joint containment LMI keeps $v_*(r)$ itself in the box.

Algorithm idea — Reference governor on the certified set (Pauli et al., Sec. 3.4)
At every time, replace the desired reference $r$ by the closest reference for which the current state is certified: $$\begin{aligned}\hat r_t=\operatorname*{arg\,min}_{\hat r}\ \|r-\hat r\|^2\quad\text{s.t.}\quad&(\chi_t-\chi_*(\hat r))^{\top}P(\chi_t-\chi_*(\hat r))\\ &+(\hat r-r_{\rm nom})^{\top}Q(\hat r-r_{\rm nom})\le1 .\end{aligned}$$ Provided this problem is feasible at $t=0$, the state always lies in the certified set of the reference actually applied, so the guaranteed ROA becomes the union of the slices over all admissible references. The paper adds that $\hat r_t$ reaches $r$ when possible and otherwise the closest admissible reference is tracked; its example shows this, but the fixed-reference theorems alone do not prove it (see the Fact below).
Fact — What the governor guarantees (derived here)
Assume the set-of-references certificate above (with bounds valid around every $v_*(r)$), a continuous network $\kappa$, and a feasible governor problem at $t=0$; let $r$ be constant and let each $\hat r_t$ be a global minimiser, applied for one sample. Then (i) the problem is feasible at every $t$, and (ii) the distance $\|r-\hat r_t\|$ is nonincreasing, hence convergent. Proof: a minimiser exists because the feasible set is closed ($\chi_*$ is continuous) and bounded ($Q\succ0$). With $\hat r_t$ held for one step, the fixed-reference theorem keeps $(\chi_{t+1},\hat r_t)$ in $\mathcal E_{P,Q}$, so the old reference is feasible at $t+1$ and the new minimiser is at least as close to $r$. This proves continued certification only: neither that the limit of $\|r-\hat r_t\|$ is the distance to the closest admissible reference, nor that $\hat r_t$ itself converges. A local solver keeps (i) and (ii) only if it falls back to $\hat r_{t-1}$ whenever it finds nothing better (the problem is not convex in general, but for a scalar reference a one-dimensional search suffices).

The example. A linearised inverted pendulum ($m=0.15$ kg, $L=0.5$ m, friction $0.5$) discretised with $T_s=0.02$ s, controlled by a state-feedback network with two tanh layers of 5 neurons, trained by supervised learning to imitate an MPC with $|u|\le1$ on trajectories from $x_0\sim\mathcal U(-0.5,0.5)$ (initial states drawn uniformly, i.e. each coordinate independently from $[-0.5,0.5]$; Primer C), and $k_\xi=1$. The global theorem is infeasible. With $d_i=0.345$ the one-reference theorem ($r=0$, 1.14 s) and the set-of-references theorem ($r_{\rm nom}=0$, 0.50 s) are feasible; every reference angle $r\in[-0.2,0.2]$ can be tracked, the union of ellipsoids is much larger than the single one, and a reference governor tracks the closest admissible point to $r=-1$, while the same loop without the governor diverges.

Caveat — This is an analysis paper
The network in the L4DC example is trained by plain supervised imitation of an MPC; no LMI is enforced during training, and the certificate is computed afterwards. Remark 2 of the paper only notes that the LMI of Theorem 1 can be imposed during training, as in Pauli et al.'s Lipschitz training (L-CSS 2022, an ADMM scheme; Module 12) or Revay et al.'s convex parameterisations. Training with such constraints is a separate step with its own cost and convergence issues (Section 6).

5. Dissipativity Analysis and Training of RNNs (Pauli et al.)

A recurrent network is itself a dynamical system, and its robustness is an input–output property: how much can a perturbation of the input sequence change the output sequence? Pauli, Berberich & Allgöwer, at – Automatisierungstechnik 70(8):730–739, 2022 give a comprehensive introduction to the analysis of RNNs with robust control and dissipativity theory, use the $\mathcal H_2$ performance and the $\ell_2$ gain as robustness measures with respect to input perturbations, and present LMI constraints for training robust RNNs. For an LTI system with impulse-response matrices $G_t$ the squared $\mathcal H_2$ norm is the energy of the impulse response, $\|G\|_{\mathcal H_2}^2=\sum_t\|G_t\|_F^2$: an average over frequencies, whereas the $\ell_2$ gain is a worst case over inputs (Primer D). For a nonlinear RNN an $\mathcal H_2$-type measure needs its own definition; the construction below, the standard one behind such results, certifies only the $\ell_2$ gain.

Caveat — Source access
The full text of the at-Automatisierungstechnik paper was not available when these notes were written; its content is summarised from the abstract. The LFT model, the $\ell_2$-gain LMI and its derivation below are the textbook construction (dissipativity plus the S-procedure with sector constraints), not a transcription of the paper's LMIs, and its $\mathcal H_2$ LMI is not reproduced.
Definition — RNN as a linear system in feedback with its activations
$$\begin{gathered}x_{t+1}=Ax_t+B_1w_t+B_2u_t,\qquad v_t=C_1x_t+D_{11}w_t+D_{12}u_t,\\ y_t=C_2x_t+D_{21}w_t+D_{22}u_t,\qquad w_t=\phi(v_t).\end{gathered}$$ A vanilla RNN $h_{t+1}=\phi(W_hh_t+W_uu_t)$, $y_t=C_hh_t$ has $x=h$, $A=0$, $B_1=I$, $B_2=0$, $C_1=W_h$, $D_{12}=W_u$, $C_2=C_h$ and $D_{11}=0$. A nonzero $D_{11}$ gives an implicit (equilibrium) layer, as in the REN.
Definition — Dissipativity and QSR supply rates (Willems 1972; ordering of Module 2)
The system is dissipative with respect to $s(u,y)$ if there is a storage $V\ge0$ with $V(x_{t+1})-V(x_t)\le s(u_t,y_t)$ along all trajectories (Willems, Arch. Rational Mech. Anal. 1972). A QSR supply is $s=\begin{bmatrix}y\\ u\end{bmatrix}^{\top}\begin{bmatrix}Q & S\\ S^{\top} & R\end{bmatrix}\begin{bmatrix}y\\ u\end{bmatrix}$; the $\ell_2$ gain $\gamma$ is $Q=-I$, $S=0$, $R=\gamma^2I$: summing from $x_0=0$ with $V(0)=0$ gives $\sum_t\|y_t\|^2\le\gamma^2\sum_t\|u_t\|^2$.
Key equation — $\ell_2$-gain LMI for an RNN with activations in the sector $[0,1]$
With $\zeta=(x,w,u)$, the certificate is $P\succeq0$, $T=\mathrm{diag}(\lambda)\succeq0$ and $\gamma$ such that $$\begin{aligned}&\begin{bmatrix}A & B_1 & B_2\\ I & 0 & 0\end{bmatrix}^{\top}\begin{bmatrix}P & 0\\ 0 & -P\end{bmatrix}\begin{bmatrix}A & B_1 & B_2\\ I & 0 & 0\end{bmatrix}\\ &+\begin{bmatrix}C_2 & D_{21} & D_{22}\\ 0 & 0 & I\end{bmatrix}^{\top}\begin{bmatrix}I & 0\\ 0 & -\gamma^2I\end{bmatrix}\begin{bmatrix}C_2 & D_{21} & D_{22}\\ 0 & 0 & I\end{bmatrix}\\ &+\begin{bmatrix}C_1 & D_{11} & D_{12}\\ 0 & I & 0\end{bmatrix}^{\top}\begin{bmatrix}0 & T\\ T & -2T\end{bmatrix}\begin{bmatrix}C_1 & D_{11} & D_{12}\\ 0 & I & 0\end{bmatrix}\preceq0 .\end{aligned}$$ For fixed weights it is linear in $(P,T,\gamma^2)$, so the smallest certified gain is an SDP.
Derivation — From the LMI to the gain bound

Step 1. Multiply the LMI by $\zeta_t=(x_t,w_t,u_t)$ on both sides. The three terms become, in order, $V(x_{t+1})-V(x_t)$ (because $x_{t+1}=Ax_t+B_1w_t+B_2u_t$), $\|y_t\|^2-\gamma^2\|u_t\|^2$ (because $y_t=C_2x_t+D_{21}w_t+D_{22}u_t$), and $2w_t^{\top}T(v_t-w_t)=2\sum_i\lambda_iw_{t,i}(v_{t,i}-w_{t,i})$.

Step 2. For activations in the sector $[0,1]$ with $\varphi(0)=0$ (tanh, ReLU) each $w_i(v_i-w_i)\ge0$, so the third term is nonnegative and $V(x_{t+1})-V(x_t)\le\gamma^2\|u_t\|^2-\|y_t\|^2$: dissipativity with the $\ell_2$-gain supply.

Step 3. Sum from $0$ to $N-1$ with $x_0=0$, $V(0)=0$ (as for this quadratic storage), and use $V\ge0$: $\sum_{t\lt N}\|y_t\|^2\le\gamma^2\sum_{t\lt N}\|u_t\|^2$ for every $N$. If $D_{11}\ne0$ the layer is implicit, and this argument presupposes that it is well posed (that $w_t$ exists and is unique); the sector bound alone does not give that. The $(w,w)$ block contains $TD_{11}+D_{11}^{\top}T-2T$. If this is negative definite (which forces $T\succ0$) and the activations are slope-restricted in $[0,1]$, as tanh and ReLU are, the implicit layer is well posed (Section 1; REN paper, eq. 12 under its Assumption 1). A sector bound is not enough: $\varphi(v)=v(1+\sin10v)/2$ lies in the sector $[0,1]$, and with $D_{11}=0.5$, $T=1$ one has $2T-TD_{11}-D_{11}^{\top}T=1\gt0$, yet $v=0.5\,\varphi(v)+0.39$ has three solutions ($v\approx0.409,\ 0.736,\ 0.779$).

Incremental version and contraction. The sector QC compares one trajectory with the equilibrium $0$. Replacing it by the slope QC applied to the difference of two trajectories, $(\Delta v,\Delta w)$, with a diagonal $T$, turns the same LMI into a bound on the incremental gain, $\sum_t\|\Delta y_t\|^2\le\gamma^2\sum_t\|\Delta u_t\|^2+V(\Delta x_0)$: the RNN is $\gamma$-Lipschitz as a map between input and output sequences, up to a term from the initial states, and if the LMI holds strictly it also forgets its initial conditions (contraction). This is exactly the certificate built into the recurrent equilibrium network (Revay, Wang & Manchester, TAC 2024, Thm. 1): an incremental Lyapunov function $V(\Delta x)=\|\Delta x\|_P^2$ with $V(\Delta x_{t+1})\le\alpha^2V(\Delta x_t)-\Gamma(\Delta v_t,\Delta w_t)$ for contraction at some rate $\alpha\lt\bar\alpha$ (their eq. 17; $\Gamma\ge0$ is the incremental QC with a diagonal multiplier), and an incremental dissipation inequality for the incremental IQC defined by $(Q,S,R)$ with $Q\preceq0$ (their Lipschitz case is $Q=-\gamma^{-1}I$, $S=0$, $R=\gamma I$, a rescaling of ours).

Derivation — From a strict LMI to forgetting the initial state

Compare two trajectories with the same input, $\Delta u_t=0$. A strict LMI stays negative definite after adding $\varepsilon I$ for some $\varepsilon\gt0$; multiplying by $(\Delta x_t,\Delta w_t,0)$ and dropping the nonnegative terms $\|\Delta y_t\|^2$ and the slope QC gives $V(\Delta x_{t+1})\le V(\Delta x_t)-\varepsilon\|\Delta x_t\|^2\le\big(1-\varepsilon/\lambda_{\max}(P)\big)V(\Delta x_t)$, using $V(\Delta x)\le\lambda_{\max}(P)\|\Delta x\|^2$ (and $P\succ0$). Iterating and using $\lambda_{\min}(P)\|\Delta x\|^2\le V(\Delta x)$ gives $\|\Delta x_t\|^2\le\frac{\lambda_{\max}(P)}{\lambda_{\min}(P)}\big(1-\varepsilon/\lambda_{\max}(P)\big)^t\|\Delta x_0\|^2$: the difference of initial states is forgotten exponentially. In the REN theorem, $\bar\alpha\in(0,1]$ is a prescribed bound on the contraction rate, unrelated to the sector bound $\alpha$ of this page.

Pitfall — Coupled multipliers and incremental properties
Multipliers designed for repeated nonlinearities (coupling different neurons) certify boundedness of signals but, as the REN paper notes (Remark 5), are not valid for incremental IQCs and contraction. This is the same error that invalidated LipSDP-Network (Module 12): an incremental constraint compares two input vectors, and the ordering between different neurons that a repeated nonlinearity provides for a single input is lost.

Training under the constraint. The weights $(A,B_1,C_1,D_{11},\dots)$ appear multiplied by $P$ and $T$, so the LMI is not jointly convex in weights and certificate (REN paper, Remark 4). Three remedies appear in this literature: fix the multipliers and enforce the LMI during training by ADMM (Pauli et al., L-CSS 2022) or by a barrier term in the loss (Pauli et al., CDC 2022, which reports that barrier training is much faster than ADMM); change variables so that the constraint is jointly convex (the convex parameterisations of the REN paper, Sec. IV); or parameterise the weights so that the LMI holds for every parameter value, which makes training unconstrained (REN direct parameterisation, Sec. V; Module 13).

6. Synthesis With Guarantees: Dissipativity-Constrained RL and Youla-REN

Analysis certifies a given network. Synthesis must search over networks while keeping the certificate, and the difficulty is algebraic: the closed-loop matrices are affine in the controller parameters, but the certificate multiplies them by $P$ and $T$, so the design condition is a bilinear matrix inequality (BMI). Two recent answers take opposite routes: convexify and project (Junnarkar, Arcak & Seiler), or build the certificate into the parameterisation (Youla-REN).

Projection-based RL under a closed-loop dissipativity LMI

Junnarkar, Arcak & Seiler, Automatica 193 (2026) 113185 synthesise neural controllers that maximise an RL reward subject to the hard constraint that the closed loop is dissipative. Their setting is continuous time: a dot means $d/dt$, and the one-step storage difference of the earlier sections becomes the derivative $\frac{d}{dt}(x^{\top}Px)=\dot x^{\top}Px+x^{\top}P\dot x=2x^{\top}P\dot x$ (product rule, Primer B). The plant is an LTI system $G_p$ in feedback with an uncertainty $\Delta_p$ (unmodelled dynamics and nonlinearities) described by a hard IQC, with disturbance $d$, performance output $e$, control $u$ and measurement $y$; the interconnection is assumed well posed (their eq. 1):

$$\begin{bmatrix}\dot x_p\\ v_p\\ e\\ y\end{bmatrix}=\begin{bmatrix}A_p&B_{pw}&B_{pd}&B_{pu}\\ C_{pv}&D_{pvw}&D_{pvd}&D_{pvu}\\ C_{pe}&D_{pew}&D_{ped}&D_{peu}\\ C_{py}&D_{pyw}&D_{pyd}&0\end{bmatrix}\begin{bmatrix}x_p\\ w_p\\ d\\ u\end{bmatrix},\qquad w_p=\Delta_p(v_p).$$

The controller is a recurrent implicit neural network (RINN); the subscript $k$ labels the controller, not a time step:

$$\begin{gathered}\dot x_k=A_kx_k+B_{kw}w_k+B_{ky}y,\qquad v_k=C_{kv}x_k+D_{kvw}w_k+D_{kvy}y,\\ u=C_{ku}x_k+D_{kuw}w_k+D_{kuy}y,\qquad w_k=\phi(v_k),\end{gathered}$$

with activations sector-bounded and slope-restricted in $[0,1]$. With $A_k,B_{kw},B_{ky},C_{kv},C_{ku}$ all zero it is a static implicit network, and a feedforward network is the case of a strictly block lower triangular $D_{kvw}$ (their Example 1). Performance is dissipativity with a quadratic supply $s(d,e)=\begin{bmatrix}d\\ e\end{bmatrix}^{\top}X\begin{bmatrix}d\\ e\end{bmatrix}$, $X=\begin{bmatrix}X_{dd}&X_{de}\\ X_{de}^{\top}&X_{ee}\end{bmatrix}$, and a quadratic storage $x^{\top}Px$ (stability: $s=0$ with $P\succ0$; $\ell_2$ gain: $\gamma^2\|d\|^2-\|e\|^2$; passivity: $d^{\top}e$). The activations enter through the sector QC with multiplier $\begin{bmatrix}0 & \Lambda\\ \Lambda & -2\Lambda\end{bmatrix}$, $\Lambda\succeq0$ diagonal, and the plant uncertainty through a static IQC with multiplier $M_{\Delta_p}=\begin{bmatrix}M_{\Delta_p,vv}&M_{\Delta_p,vw}\\ M_{\Delta_p,vw}^{\top}&M_{\Delta_p,ww}\end{bmatrix}$ (dynamic IQCs of the class below are first converted to this form).

Theorem — Convexification for full-order controllers (Junnarkar, Arcak & Seiler, Thm. 4)
Assumptions. The plant interconnection is well posed; the uncertainty IQC is static, or dynamic with a factorisation $\Psi=\mathrm{diag}(\Psi_1,\Psi_2)$, $M=\mathrm{diag}(I,-I)$, $\Psi$ stable with a stable proper inverse (their Lemma 3 transforms it to a static one; $n_p$ then counts the filter states too); $X_{ee}\preceq0$ (e.g. $\ell_2$ gain, passivity) and $M_{\Delta_p,vv}\succeq0$ (e.g. LTI uncertainty, sector $[0,1]$ nonlinearities); the controller state has the plant's dimension, $n_k=n_p$.
Statement. For fixed $M_{\Delta_p}$ and $X$, there exist a controller and a quadratic storage that satisfy their dissipation inequality (Lemma 2, the continuous-time analogue of the LMI of Section 5) with $P\succ0$ and $\Lambda\succ0$ if and only if there exist new variables $\hat\theta=\{S,R,N_A,N_B,N_C,$ $D_{kuw},\hat D_{kvy},\hat D_{kvw},\Lambda\}$ satisfying an LMI that is affine in $\hat\theta$ (their eq. 23, written out below), with $S\succ0$, $R\succ0$, $\Lambda\succ0$ and $\begin{bmatrix}R & I\\ I & S\end{bmatrix}\succ0$. The equivalence is between two certificates, not a description of all dissipative controllers.
The change of variables. Partition $P=\begin{bmatrix}S & U\\ U^{\top} & \star\end{bmatrix}$, $P^{-1}=\begin{bmatrix}R & V\\ V^{\top} & \star\end{bmatrix}$ ($\star$: blocks not needed by name), apply the congruence $\mathrm{diag}(Y,I)$ with $Y=\begin{bmatrix}R & I\\ V^{\top} & 0\end{bmatrix}$ and set $\hat D_{kvy}=\Lambda D_{kvy}$, $\hat D_{kvw}=\Lambda D_{kvw}$, $N_B=SB_{pu}D_{kuw}+UB_{kw}$, $N_C=\Lambda D_{kvy}C_{py}R+\Lambda C_{kv}V^{\top}$ and $$\begin{aligned}N_A&=\begin{bmatrix}N_{A11}&N_{A12}\\ N_{A21}&N_{A22}\end{bmatrix}=\begin{bmatrix}SA_pR & 0\\ 0 & 0\end{bmatrix}\\ &\quad+\begin{bmatrix}U & SB_{pu}\\ 0 & I\end{bmatrix}\begin{bmatrix}A_k & B_{ky}\\ C_{ku} & D_{kuy}\end{bmatrix}\begin{bmatrix}V^{\top} & 0\\ C_{py}R & I\end{bmatrix}.\end{aligned}$$ The controller is recovered from $\hat\theta$ by factorising $VU^{\top}=I-RS$ (for instance from an SVD, Primer A; invertible because $\begin{bmatrix}R & I\\ I & S\end{bmatrix}\succ0$) and solving these linear equations. This is the change of variables of Scherer, Gahinet and Chilali for $\mathcal H_\infty$ synthesis, extended to the multiplier $\Lambda$ of the network.
Why the assumptions: $X_{ee}\preceq0$ and $M_{\Delta_p,vv}\succeq0$ allow the factorisations $X_{ee}=-L_X^{\top}L_X$ and $M_{\Delta_p,vv}=L_{\Delta_p}^{\top}L_{\Delta_p}$, so a Schur complement removes the terms that are quadratic in the closed-loop matrices and leaves a condition that is bilinear only through $P$, $\Lambda$ and the controller; the plant-order controller makes the congruence $Y$ and the recovery of $(A_k,B_{kw},\dots)$ from $\hat\theta$ possible; the restriction on the IQC lets their Lemma 3 absorb the filters into an extended plant with a static IQC.
Derivation — The affine synthesis LMI written out, and recovering the controller

From the BMI to an LMI. Close the loop: state $x=(x_p,x_k)$, $w=(w_p,w_k)$, closed-loop matrices $\mathcal A,\mathcal B_w,\mathcal B_d,\mathcal C_v,\mathcal D_{vw},\dots$, affine in the controller. Lemma 2 asks that $2x^{\top}P\dot x+\begin{bmatrix}v\\ w\end{bmatrix}^{\top}M\begin{bmatrix}v\\ w\end{bmatrix}-s(d,e)\le0$ for all $(x,w,d)$, with $M$ collecting both QCs ($M_{vw}=\mathrm{diag}(M_{\Delta_p,vw},\Lambda)$, $M_{ww}=\mathrm{diag}(M_{\Delta_p,ww},-2\Lambda)$); integrating it, the hard IQC (zero filter state) makes the $M$-term's integral nonnegative, which leaves the dissipation inequality. Two terms are quadratic in the controller, $v_p^{\top}M_{\Delta_p,vv}v_p=\|L_{\Delta_p}v_p\|^2$ and $-e^{\top}X_{ee}e=\|L_Xe\|^2$; a Schur complement moves them into an off-diagonal block $G$ (their eq. 21). The substitution $x=Yz$ (the congruence) then turns every product of $P$ or $\Lambda$ with a controller matrix into one of the new variables (their eqs. 23–25). Order the columns as $(z,w_p,w_k,d)$, let $J_w$ pick $(w_p,w_k)$ and $J_d$ pick $d$. The blocks are

$$\begin{gathered} H=Y^{\top}P\mathcal AY=\begin{bmatrix}A_pR+B_{pu}N_{A21}&A_p+B_{pu}N_{A22}C_{py}\\N_{A11}&SA_p+N_{A12}C_{py}\end{bmatrix},\\ K_w=Y^{\top}P\mathcal B_w=\begin{bmatrix}B_{pw}+B_{pu}N_{A22}D_{pyw}&B_{pu}D_{kuw}\\SB_{pw}+N_{A12}D_{pyw}&N_B\end{bmatrix},\\ K_d=Y^{\top}P\mathcal B_d=\begin{bmatrix}B_{pd}+B_{pu}N_{A22}D_{pyd}\\SB_{pd}+N_{A12}D_{pyd}\end{bmatrix},\end{gathered}$$

and, for $a=v,e$, the plant output $a$ as a function of $(z,w_p,w_k,d)$ (use $N_{A22}=D_{kuy}$ and $N_{A21}=C_{ku}V^{\top}+D_{kuy}C_{py}R$):

$$\begin{aligned}O_a=\big[\,&C_{pa}R+D_{pau}N_{A21}\ \ \ \ C_{pa}+D_{pau}N_{A22}C_{py}\\ &D_{paw}+D_{pau}N_{A22}D_{pyw}\ \ \ \ D_{pau}D_{kuw}\ \ \ \ D_{pad}+D_{pau}N_{A22}D_{pyd}\,\big].\end{aligned}$$

The multiplier cross terms, the second row being $\Lambda v_k$ in the new variables, are

$$Z=\begin{bmatrix}M_{\Delta_p,vw}^{\top}O_v\\ \big[N_C\ \ \ \hat D_{kvy}C_{py}\ \ \ \hat D_{kvy}D_{pyw}\ \ \ \hat D_{kvw}\ \ \ \hat D_{kvy}D_{pyd}\big]\end{bmatrix}.$$

With $\operatorname{He}(H)=H+H^{\top}$:

$$\begin{gathered} F=\begin{bmatrix}\operatorname{He}(H)&K_w&K_d\\K_w^{\top}&0&0\\K_d^{\top}&0&0\end{bmatrix}+J_w^{\top}Z+Z^{\top}J_w+J_w^{\top}M_{ww}J_w\\ {}-J_d^{\top}X_{dd}J_d-J_d^{\top}X_{de}O_e-O_e^{\top}X_{de}^{\top}J_d,\\ G=\begin{bmatrix}L_{\Delta_p}O_v\\L_XO_e\end{bmatrix},\qquad \boxed{\begin{bmatrix}F&G^{\top}\\G&-I\end{bmatrix}\preceq0}.\end{gathered}$$

Every entry is affine in $\hat\theta$ (and in $M_{\Delta_p,ww}$, $X_{dd}$, which may therefore also be decision variables, e.g. to minimise $\gamma$). By the Schur complement the boxed LMI is $F+G^{\top}G\preceq0$, the congruence-transformed Lemma 2 inequality.

Recovering the controller. The upper-left block of $PP^{-1}=I$ reads $SR+UV^{\top}=I$, so $VU^{\top}=I-RS$. The LMI $\begin{bmatrix}R&I\\ I&S\end{bmatrix}\succ0$ gives $R-S^{-1}\succ0$ (Schur complement), so $S^{1/2}RS^{1/2}\succ I$ and $I-RS=S^{-1/2}\big(I-S^{1/2}RS^{1/2}\big)S^{1/2}$ is invertible. Factor it, e.g. with an SVD $I-RS=\Omega_1\Sigma\Omega_2^{\top}$, as $V=\Omega_1\Sigma^{1/2}$, $U=\Omega_2\Sigma^{1/2}$; both are invertible, hence so is $Y$, and $P=\begin{bmatrix}I&S\\ 0&U^{\top}\end{bmatrix}Y^{-1}$ (check: $PY=\begin{bmatrix}SR+UV^{\top}&S\\ U^{\top}R+\star V^{\top}&U^{\top}\end{bmatrix}$ and $U^{\top}R+\star V^{\top}=0$ is the lower-left block of $PP^{-1}=I$). Then $D_{kvy}=\Lambda^{-1}\hat D_{kvy}$, $D_{kvw}=\Lambda^{-1}\hat D_{kvw}$, $(A_k,B_{ky},C_{ku},D_{kuy})$ from $N_A$ by inverting its two invertible outer factors, $B_{kw}=U^{-1}(N_B-SB_{pu}D_{kuw})$ and $C_{kv}=\Lambda^{-1}(N_C-\hat D_{kvy}C_{py}R)V^{-\top}$. With the zero supply the nonstrict LMI gives a nonincreasing storage (stability), not convergence.

Algorithm 2: RL with a dissipativity-enforcing projection (Junnarkar, Arcak & Seiler, Alg. 1)
  1. Initialise the controller parameters $\theta$ arbitrarily (randomly, or warm-started), and $P,\Lambda\leftarrow I$ (or a known certificate of $\theta$). // $\theta$ need not be certified yet
  2. For $i=1,2,\dots$:
  3. $\theta'\leftarrow$ one reinforcement-learning step on $\theta$ // reward only; the training environment may differ from the design model
  4. If some $(P',\Lambda')$ certifies $\theta'$ (an LMI in $(P,\Lambda)$ for fixed $\theta'$): $\theta,P,\Lambda\leftarrow\theta',P',\Lambda'$.
  5. Else: build $\hat\theta'$ from $(\theta',P,\Lambda)$; project it onto the convex set of Theorem 4 (first the minimum Frobenius distance $\delta^*$, a projection as in Primer B; then a re-solve that maximises $\epsilon_{RS}$ in $\begin{bmatrix}R&I\\ I&S\end{bmatrix}\succ\epsilon_{RS}I$ subject to distance $\le\beta\delta^*$, $\beta\ge1$, which improves the conditioning of $I-RS$);
  6. extract $P,\Lambda$ from the projection and set $\theta\leftarrow\arg\min\|\theta-\theta'\|$ over controllers certified by this $(P,\Lambda)$ // again an SDP

Every iterate that leaves the loop is certified for the design model; the initial $\theta$ is not, and $P,\Lambda=I$ only serve to build the first $\hat\theta'$ (the projection then produces a valid $(P,\Lambda)$). Well-posedness of the implicit layer is added as $\Lambda D_{kvw}+D_{kvw}^{\top}\Lambda-2\Lambda\prec0$ (their Remark 7), which is $\hat D_{kvw}+\hat D_{kvw}^{\top}-2\Lambda\prec0$, affine in the new variables. The authors estimate the cost of the dissipativity check as $\mathcal O\big((n_p^2+n_\phi)^2(n_p+n_\phi)^2\big)$ and of each projection as $\mathcal O\big((n_p+n_\phi)^6\big)$ flops ($n_p$ plant states, $n_\phi$ neurons), which is why the check runs first and the projections only when it fails. Examples: an inverted pendulum and a flexible rod on a cart. Reward maximisation is the soft objective and dissipativity the hard constraint, a certificate-level analogue of the constrained RL of Module 8.

Stable by design: the Youla-REN

Wang, Barbara, Revay & Manchester, IEEE L-CSS 7:91–96, 2023 take the opposite route. For a partially observed linear plant $x_{t+1}=Ax_t+Bu_t+d^x_t$, $y_t=Cx_t+d^y_t$ (stabilisable and detectable, Primer D), pick gains with $A-BK$ and $A-LC$ stable and augment the observer-based controller (see Primer D) with a nonlinear Youla parameter $\mathcal Q$:

$$\hat x_{t+1}=A\hat x_t+Bu_t+L\tilde y_t,\qquad u_t=-K\hat x_t+\mathcal Q(\tilde y)_t,\qquad \tilde y_t=y_t-C\hat x_t .$$
Theorem — Nonlinear Youla parameterisation (Wang et al., L-CSS 2023, Props. 1–2)
Write $d=(d^x,d^y)$ for the disturbances and $z=(x,u)$ for the performance output. (1) For every contracting and Lipschitz $\mathcal Q$, the closed loop is contracting and the map $d\mapsto z$ is Lipschitz. (2) Conversely, every output-feedback controller (with a locally Lipschitz state-space realisation) that achieves a contracting and Lipschitz closed loop can be written in this form with a contracting and Lipschitz $\mathcal Q$. The closed-loop response is $z=T_0d+T_1\mathcal Q(T_2d)$ with stable LTI systems $T_0,T_1,T_2$ (from zero initial states; nonzero ones add decaying free responses). (3) (Prop. 2) Let the REN activation be non-polynomial and slope-restricted in $[0,1]$. For every contracting and Lipschitz $\mathcal Q$ with a locally Lipschitz state-space realisation and all $\bar y,\epsilon\gt0$ there is a REN $\widetilde{\mathcal Q}$ with sufficiently many states and neurons such that $\|\widetilde{\mathcal Q}(\tilde y)-\mathcal Q(\tilde y)\|_\infty\le\epsilon$ for all inputs with $\|\tilde y\|_\infty\le\bar y$, where $\|\tilde y\|_\infty=\sup_t|\tilde y_t|$ bounds every sample (the proof handles initial states by fading memory, starting both systems in the infinite past). So the Youla-REN can approximate every stabilising controller uniformly on bounded inputs; this is an existence statement, not a promise that training finds those parameters.
Why it works. Subtracting the observer equation from the plant, the observer error obeys $(x-\hat x)_{t+1}=(A-LC)(x_t-\hat x_t)+d^x_t-Ld^y_t$, and $\tilde y_t=C(x_t-\hat x_t)+d^y_t$; neither equation contains $u$. Hence the innovation $\tilde y=T_2d$ does not depend on what $\mathcal Q$ outputs, and $\mathcal Q$ sits outside every feedback loop: stability reduces to a cascade of stable systems (see Primer D), and no LMI has to be solved during learning. Stabilisability and detectability are exactly what guarantee such gains $K,L$.

Because the REN is directly parameterised (every parameter vector gives a contracting, Lipschitz $\mathcal Q$; Module 13), learning is unconstrained: exact gradients through a differentiable model, or random search with a zeroth-order oracle (cost values only, no gradients; Primer B) as in RL. The paper reports faster convergence to better controllers than competing policy classes, with stability preserved during the learning transients. The guarantee is for the nominal linear model; with model error the innovation depends on $u$ again and robustness must be argued separately (for instance by small gain with the Lipschitz bound of $\mathcal Q$). The survey of Manchester, Wang & Barbara, Annu. Rev. 2026 collects these model classes and their use as observers and policies. A faster recurrent class from the same group, R2DN (Barbara, Wang & Manchester, CDC 2026), replaces the REN's equilibrium layer by a 1-Lipschitz feedforward network in feedback with an LTI system and reports up to an order of magnitude faster training and inference.

RouteWhat is guaranteed, whenPlant classComputation
Analysis LMI (Sections 2–4)stability / ROA for the trained network, after trainingLTI, perturbations via IQCsone SDP per certificate
Constrained training (ADMM, barrier; Module 12)the LMI at the end of training (ADMM); at every iterate with projection or a log-det barrierLTI + QCsSDP-type steps or barrier gradients
Projection-based RL (Junnarkar et al.)closed-loop dissipativity of every accepted iterateLTI + IQC uncertainty, full-order controllercheck LMI, convex projection when it fails
Youla-RENcontraction and Lipschitz closed loop for every parameter valueknown partially observed LTInone beyond gradient steps

7. Reachability of NN Loops and Verified Lyapunov Controllers

Stability says where trajectories end; safety asks where they go on the way. Two very different certificate styles address this: SDP relaxations built from the same QCs as above, and complete verification by bound propagation and branch-and-bound. Only the interface of the latter is needed here (Module 15 develops the algorithms). A verifier receives a bounded input region and a statement that must hold at every point of it. Bound propagation pushes sound lower and upper bounds on every intermediate value through the network; when they are inconclusive, branch-and-bound splits the problem, typically by fixing an undecided ReLU $w=\max(z,0)$ to its active piece ($z\ge0$, $w=z$) or its inactive piece ($z\le0$, $w=0$), and repeats on the pieces. The outcome is a proof, a counterexample, or "unresolved" when the time budget runs out.

Theorem — Soundness and completeness of ReLU branch-and-bound (Wang et al., NeurIPS 2021, Thm. 3.3)
For a ReLU network, a bounded $\ell_\infty$ input box $\mathcal C$ (including the paper's $\ell_\infty$ ball) and a linear specification $f(x)\gt0$ for all $x\in\mathcal C$, $\beta$-CROWN combined with branch-and-bound on ReLU splits is sound (a piece is declared verified only when its lower bound proves the specification) and complete (it decides whether the specification holds). Why: once every undecided ReLU of a piece is split, the network is affine there and $\beta$-CROWN returns the exact minimum of $f$ on it, detecting infeasible pieces through an unbounded dual. What it does not say: $m$ undecided ReLUs can require $2^m$ pieces, so completeness gives no useful time bound, and it does not carry over to smooth activations or nonlinear dynamics, which need sound nonlinear bounds and are then verified incompletely.

Reach-SDP

Hu, Fazlyab, Morari & Pappas, CDC 2020 over-approximate the forward reachable sets of a linear time-varying plant $x_{t+1}=A_tx_t+B_tu_t+c_t$ with a projected ReLU policy $u_t=\mathrm{Proj}_{\mathcal U_t}(\pi(x_t))$, $\mathcal U_t$ a box, so that the projection clips each input coordinate to its interval (Primer B). The exact one-step image of a set $\mathcal X$ is $\mathcal R_t(\mathcal X)=\{A_tx+B_t\mathrm{Proj}_{\mathcal U_t}(\pi(x))+c_t:\ x\in\mathcal X\}$, and the goal is a simple set $\bar{\mathcal R}$ that contains all of it. The projection is itself written as two extra ReLU layers, $x^{\ell+1}=\max(W_\ell x^\ell+b_\ell-\underline u,0)+\underline u$ and $x^{\ell+2}=-\max(\bar u-x^{\ell+1},0)+\bar u$, so the whole loop map is a ReLU network plus linear algebra. Three families of QCs describe it, all in the basis $\boldsymbol\xi=(x_0,\text{all neurons},1)$, whose constant entry lets a quadratic form carry affine and constant terms, $[x;1]^{\top}\begin{bmatrix}G&g\\ g^{\top}&g_0\end{bmatrix}[x;1]=x^{\top}Gx+2g^{\top}x+g_0$. The initial set: for a polytope $\{Hx\le h\}$ (Primer B) every product of two entries of $s=Hx-h\le0$ is nonnegative, so $s^{\top}\Gamma s\ge0$ for every symmetric $\Gamma$ with nonnegative entries (entrywise, not positive semidefinite). Each ReLU: $y_i\ge x_i$, $y_i\ge0$, $y_i^2=x_iy_i$, which is exact, because $y_i(y_i-x_i)=0$ leaves $y_i=0\ge x_i$ or $y_i=x_i\ge0$, i.e. $y_i=\max(x_i,0)$; and the coupled repeated-ReLU constraints $\lambda_{ij}\big[(y_j-y_i)(x_j-x_i)-(y_j-y_i)^2\big]\ge0$ with $\lambda_{ij}\ge0$. The candidate output set: $\{y:[y;1]^{\top}S[y;1]\le0\}$.

Theorem — SDP for the one-step reachable set (Hu et al., CDC 2020, Thm. 1)
If the initial set satisfies the QCs defined by $M_{\rm in}(\cdot)$, the projected network those defined by $M_{\rm mid}(\cdot)$, and $M_{\rm out}(S)$ encodes the candidate set $\bar{\mathcal R}$, then feasibility of $$M_{\rm in}(P)+M_{\rm mid}(Q)+M_{\rm out}(S)\preceq0$$ for some admissible multipliers $(P,Q,S)$ implies $\mathcal R(\mathcal X_0)\subseteq\bar{\mathcal R}(\mathcal X_0)$. Proof. Multiply by the basis vector $\boldsymbol\xi$ generated by any $x_0\in\mathcal X_0$: the first two quadratic terms are $\ge0$, so the third is $\le0$, which says that the successor state lies in $\bar{\mathcal R}$. $\blacksquare$ For a polytope with fixed facet normals $a_i$ one solves $\min b_i$ for each facet; for an ellipsoid $\{y:\|Ey+f\|_2\le1\}$, $\min-\log\det E$ (Sec. 4.3; convex after a Schur complement, derived below). Multi-step sets are obtained recursively, $\bar{\mathcal R}_{t+1}=\mathrm{ReachSDP}(\bar{\mathcal R}_t)$, so over-approximation errors accumulate with the horizon.
Derivation — The minimum-volume outer ellipsoid as a convex SDP

Take $E\succ0$ symmetric. The map $y\mapsto Ey+f$ sends the ellipsoid onto the unit ball, so its volume is the unit-ball volume divided by $\det E$; minimum volume means minimising $-\log\det E$, a convex function ($\log\det$ is concave, Primer B). Write the successor as $y=F\boldsymbol\xi$ and use $e_0^{\top}\boldsymbol\xi=1$ for the constant entry: then $Ey+f=\Omega\boldsymbol\xi$ with $\Omega=EF+fe_0^{\top}$, affine in $(E,f)$, and $\|Ey+f\|^2-1=\boldsymbol\xi^{\top}(\Omega^{\top}\Omega-e_0e_0^{\top})\boldsymbol\xi$, so $M_{\rm out}(S)=\Omega^{\top}\Omega-e_0e_0^{\top}$. This is quadratic in $(E,f)$, but by the Schur complement (with $-I\prec0$) $$\begin{gathered}M_{\rm in}(P)+M_{\rm mid}(Q)-e_0e_0^{\top}+\Omega^{\top}\Omega\preceq0\\ \iff\begin{bmatrix}M_{\rm in}(P)+M_{\rm mid}(Q)-e_0e_0^{\top}&\Omega^{\top}\\ \Omega&-I\end{bmatrix}\preceq0,\end{gathered}$$ which is affine in $(P,Q,E,f)$. Contrast Section 2: there the ellipsoid is an inner set in which $P$ enters linearly, and maximising its volume means minimising $\log\det P$, a concave function; here the outer set has a convex volume objective.

The paper verifies a double integrator (sampling time 1 s, $|u|\le1$) under a ReLU network with hidden layers of 10 and 5 neurons trained on 2420 samples of a linear MPC: every $x_0\in[2.5,3]\times[-0.25,0.25]$ reaches $[-0.25,0.25]^2$ within 6 steps without violating the state constraint, and the polytopic over-approximations are tight. A second example certifies a network that imitates a nonlinear MPC for a 6-D quadrotor.

Connection to Module 12
The coupled terms $\lambda_{ij}$ are valid here: they relate two neurons evaluated at the same input, i.e. two points $(x_i,y_i)$ and $(x_j,y_j)$ of one and the same scalar ReLU, so the chord between them has slope in $[0,1]$. The same coupling applied to increments $\varphi(x_i)-\varphi(x'_i)$ and $\varphi(x_j)-\varphi(x'_j)$ of two different input vectors has no such interpretation, which is why it is invalid in LipSDP (Module 12) and in incremental analysis (Section 5).

Formally verified neural Lyapunov controllers

Yang, Dai, Shi, Hsieh, Tedrake & Zhang, ICML 2024 drop the linear plant and the quadratic Lyapunov function. For $x_{t+1}=f(x_t,u_t)$ with continuous $f$ they train a controller $u=\mathrm{clamp}(\phi_\pi(x)-\phi_\pi(x_*)+u_*,u_{\rm lo},u_{\rm up})$, for output feedback also a neural observer $\hat x_{t+1}=f(\hat x_t,u_t)+\phi_{\rm obs}(\hat x_t,y_t-h(\hat x_t))-\phi_{\rm obs}(\hat x_t,0)$, and a Lyapunov function $V$ that is positive definite by construction, on the internal state $\xi$ ($\xi=x$, or $\xi=(x,\hat x-x)$ for output feedback).

Theorem — Verifiable ROA condition (Yang et al., ICML 2024, Thm. 3.3)
Let $\mathcal B$ be a compact "region of interest" (see Primer 0) containing the equilibrium $\xi_*$, $V$ continuous with $V(\xi_*)=0$ and $V\gt0$ elsewhere, $\kappa\gt0$ a decay rate, $\rho\gt0$, $\mathcal S=\{\xi\in\mathcal B:V(\xi)\lt\rho\}$ and $F(\xi)=V(f_{\rm cl}(\xi))-(1-\kappa)V(\xi)$. If ($\wedge$ = and, $\vee$ = or; Primer 0) $$\big(F(\xi)\le0\ \wedge\ f_{\rm cl}(\xi)\in\mathcal B\big)\ \vee\ V(\xi)\ge\rho\qquad\forall\,\xi\in\mathcal B,$$ i.e. every $\xi\in\mathcal B$ with $V(\xi)\lt\rho$ has decrease and a successor in $\mathcal B$, then $\mathcal S$ is forward invariant, $V$ decays exponentially on it, and $\mathcal S$ is an inner approximation of the ROA. Proof: for $\xi\in\mathcal S$ the second alternative is false, so $V(\xi^+)\le(1-\kappa)V(\xi)\lt\rho$ and $\xi^+\in\mathcal B$, i.e. $\xi^+\in\mathcal S$. Iterating, $V(\xi_t)\le(1-\kappa)^tV(\xi_0)\to0$. The step to $\xi_t\to\xi_*$, which the paper leaves implicit, uses compactness: for each $r\gt0$ the continuous $V$ has a positive minimum on the compact set $\{\xi\in\mathcal B:\|\xi-\xi_*\|\ge r\}$, so $V(\xi_t)\to0$ forces $\|\xi_t-\xi_*\|\lt r$ eventually. (Exponential decay of $\|\xi_t-\xi_*\|$ itself would need bounds of $V$ by multiples of $\|\xi-\xi_*\|$.)
Why the disjunction. Earlier works demanded decrease on all of $\mathcal B$ and then took the largest sublevel set inside $\mathcal B$, which wastes the part of $\mathcal B$ outside $\mathcal S$ (not invariant anyway) and yields a smaller certified set. Here decrease is required only where $V\lt\rho$, and the condition is still stated on the explicitly defined box $\mathcal B$ that a verifier can handle.

The condition is checked by $\alpha,\beta$-CROWN (bound propagation plus branch-and-bound; Xu et al., ICLR 2021, Wang et al., NeurIPS 2021), extended to the nonlinear operations of the dynamics, and the largest verifiable $\rho$ is found by bisection. Training alternates adversarial search for counterexamples with a loss that also enlarges $\mathcal S$. The authors report larger verified regions than prior neural Lyapunov methods and, to their knowledge, the first formally verified neural output-feedback controllers with Lyapunov certificates. Context: Module 11.

Two certificate cultures
LMI certificates are cheap, compositional and scale polynomially, but need a linear (or IQC-bounded) plant and a quadratic storage, and their conservatism is structural. Bound-propagation verification accepts nonlinear plants and neural Lyapunov functions, and can be complete for specified problem classes, but branch-and-bound is exponential in the worst case (complete verification of ReLU networks is NP-hard, Primer 0: unless P $=$ NP, no algorithm is fast on every instance, although many instances are easy) and in practice limited to low-dimensional state spaces. SDP-CROWN-type methods (Module 15) try to import SDP tightness into bound propagation.

8. Towards Scale: ReLU IQCs and Incremental Analysis (2025–2026)

Two recent preprints from the Seiler–Hu–Dullerud collaboration push the IQC programme in the two directions that matter most: tighter descriptions of the activation, and certificates whose size does not grow with depth. Both are arXiv preprints; statements may change between versions.

ReLU-specific hard IQCs. Vahedi Noori, Hu, Dullerud & Seiler (arXiv, Nov. 2025) analyse internal stability of a discrete-time feedback system with a ReLU nonlinearity, motivated by recurrent networks. After reviewing static QCs for slope-restricted nonlinearities they derive hard IQCs for the scalar ReLU from FIR filters and structured matrices, combine them with a dissipation inequality into an LMI, prove that the new IQCs form a superset of the Zames–Falb IQCs for slope-restricted nonlinearities, and report less conservative stability margins than Zames–Falb multipliers and static QCs, sometimes dramatically so.

Intuition
Zames–Falb multipliers use only monotonicity and slope bounds, so they hold for every function in $\mathrm{slope}[0,1]$. ReLU has more structure: $w=\max(v,0)$ gives $w\ge0$, $w-v\ge0$ and $w(w-v)=0$, the constraints DeepSDP and Reach-SDP use at a single time. Products of nonnegative quantities at different times are again valid quadratic constraints, e.g. $w_t(w_s-v_s)\ge0$ and $w_tw_s\ge0$ for all $t,s$; none of them holds for a general slope-restricted function (tanh takes negative values). The IQCs below are finite-memory combinations of exactly such products, which is why their class can be strictly larger than the Zames–Falb class.
Theorem — ReLU FIR hard IQCs and internal stability (Vahedi Noori et al. 2025, Lemma 4 and Thm. 2)
Let $w_t=\max(v_t,0)$ be a scalar ReLU and fix a window length $L$. The FIR filter $\Psi_L$ outputs $r_t=(v_t,\dots,v_{t-L},w_t,\dots,w_{t-L})$ from zero initial conditions (samples before $t=0$ are $0$). Take coefficients $m^1_i,m^2_i\ge0$ for $0\le i\le L$ and $m^3_i$ for $-L\le i\le L$ with $m^3_i\ge0$ for $i\ne0$ ($m^3_0$ is free). Let $M_1,M_2$ be the $(L+1)\times(L+1)$ matrices with first row and first column $(m^j_0,\dots,m^j_L)$ and zeros elsewhere, and $M_3$ the one with first row $(m^3_0,m^3_1,\dots,m^3_L)$, first column $(m^3_0,m^3_{-1},\dots,m^3_{-L})$ and zeros elsewhere. Then $$\begin{gathered}M=\begin{bmatrix}M_1&-M_3^{\top}-M_1\\ -M_3-M_1&M_1+M_2+M_3+M_3^{\top}\end{bmatrix}\quad\text{satisfies}\\ \sum_{t=0}^{N}r_t^{\top}Mr_t\ge0\quad\text{for every }N,\end{gathered}$$ a hard IQC. If an LTI plant is in well-posed feedback with this ReLU and, with $\Psi_L$ appended as in Section 3, the dissipation LMI of the theorem there holds strictly with this $M$ for some $P\succeq0$, then the loop is internally stable: every state trajectory converges to $0$ (the proof uses $P+\epsilon I\succ0$ and sums, as in Section 3).
Why the IQC holds. Stack the samples up to time $N$. The sum becomes a quadratic form in $a=w-v\ge0$ and $b=w\ge0$, namely $a^{\top}Q_1a+b^{\top}Q_2b+2b^{\top}Q_3a$ with banded Toeplitz matrices $Q_j$ built from the $m^j_i$. The first two terms add nonnegative coefficients times nonnegative products; in the third, every product $b_ta_s$ with $t\ne s$ has a coefficient $\ge0$, and the products with $t=s$ vanish because $w_t(w_t-v_t)=0$. Zames–Falb as a special case: $m^1=m^2=0$ and $m^3_i=-m_i$ reproduce every Zames–Falb multiplier of the same order ($m_i\le0$ for $i\ne0$, $\sum_im_i\ge0$), and the ReLU class does not even need the sum condition.

Depth-independent incremental certificates. Wang, Seiler, Dullerud & Hu (arXiv, Sept. 2026) consider an LTI plant $G$ in feedback with a deep feedforward network and an uncertainty $\Delta_U$ described by incremental $\rho$-hard IQCs. With nonzero biases the equilibrium may move or not exist, so they study pairs of trajectories (incremental convergence and incremental $\ell_2$ gain) instead of an equilibrium.

Definition — Incremental $\rho$-hard IQC (Wang et al. 2026, Def. 2 and Remark 1)
Let $\rho\in(0,1]$, $\Psi$ a stable filter and $M=M^{\top}$. A causal $\Delta_U$ with $\Delta_U(0)=0$ satisfies the incremental $\rho$-hard IQC defined by $(\Psi,M)$ if for any two inputs $p,p'$ with outputs $q=\Delta_U(p)$, $q'=\Delta_U(p')$ the filtered differences $\delta r=\Psi(p-p',q-q')$, computed from zero filter state, satisfy $$\sum_{k=0}^{K}\rho^{-k}\,\delta r_k^{\top}M\delta r_k\ge0\qquad\text{for every }K\ge0 .$$ $\rho=1$ is the incremental version of the hard IQC of Section 3. The weights $\rho^{-k}\ge1$ grow, so a smaller $\rho$ is a stronger requirement: the $\rho$-hard IQC implies the $\rho'$-hard one for every $\rho'\in[\rho,1]$, by summation by parts (with $S_k$ the weighted partial sums, $\sum_{k\le K}\delta r_k^{\top}M\delta r_k=\rho^KS_K+\sum_{k\lt K}(\rho^k-\rho^{k+1})S_k\ge0$ for $\rho'=1$). The certificates below, like the paper's Theorems 1–2, use only this $\rho=1$ consequence and therefore claim convergence without a rate; $\rho\lt1$ enters the paper's boundedness result (Thm. 3).

For layers $h^l=\phi(W_lh^{l-1}+b_l)$, $l=1,\dots,N$ (their indexing, $h^0=v$ the network input, the last layer's output $h^N=w$ taken directly as the network output; a linear readout is absorbed into the plant), activations slope-restricted in $[0,1]$ with $\phi(0)=0$, the whole network satisfies the $\Lambda$-incremental QC

$$\begin{bmatrix}\delta v\\ \delta h\end{bmatrix}^{\top}\underbrace{\begin{bmatrix}0 & W_1^{\top}\Lambda_1 & & \\ \Lambda_1W_1 & -2\Lambda_1 & W_2^{\top}\Lambda_2 & \\ & \Lambda_2W_2 & \ddots & \ddots\\ & & \ddots & -2\Lambda_N\end{bmatrix}}_{M_{\rm NN}(\Lambda)}\begin{bmatrix}\delta v\\ \delta h\end{bmatrix}\ge0,\qquad \Lambda_l\ \text{diagonal},\ \Lambda_l\succeq0,$$

where $\delta v=\delta h^0$ is the difference of the two network inputs (not the stack of pre-activations of Section 1) and $\delta h=(\delta h^1,\dots,\delta h^N)$, because the form equals $\sum_l2(\delta h^l)^{\top}\Lambda_l(W_l\delta h^{l-1}-\delta h^l)=\sum_{l,i}2\lambda_{l,i}\,\delta h^l_i(\delta z^l_i-\delta h^l_i)$, with $\delta z^l=W_l\delta h^{l-1}$ the pre-activation increment (the bias cancels), and slope $[0,1]$ gives $\delta h_i(\delta z_i-\delta h_i)\ge0$ neuron by neuron. The full-order SDP built from it (their Thm. 1) grows with the total number of neurons. The paper decomposes it and combines the decomposition with scalable Lipschitz estimation for the inner layers, so that the remaining control-analysis LMIs depend only on the widths of the last two layers and not on the depth (computing the inner-layer certificates still visits every layer); the incremental small-gain test is recovered as a special case, and a multi-round alternating update of the coupling variables reduces conservatism. They also show that incremental convergence implies convergence of all trajectories to a common limit point under an extra assumption (existence of one convergent trajectory), with a counterexample without it.

Derivation — From incremental convergence to a common limit

If $\|x_t-y_t\|\to0$ for any two trajectories and one trajectory satisfies $y_t\to y_*$, then $\|x_t-y_*\|\le\|x_t-y_t\|+\|y_t-y_*\|\to0$ (triangle inequality): every trajectory has the same limit. Without a convergent reference, differences can vanish while every trajectory drifts, as the sequences $x^{(a)}_t=t+a\,2^{-t}$ show: $x^{(a)}_t-x^{(b)}_t=(a-b)2^{-t}\to0$, yet each tends to $\infty$. The paper's counterexample (Thm. 4) is a time-varying interconnection that satisfies the LMI although its trajectories converge to no fixed point.

Theorem — Depth-independent reduced certificate (Wang, Seiler, Dullerud & Hu 2026, Lemma 1 and Thm. 2)

Setting. The loop is well posed; $\Delta_U$ is causal, $\Delta_U(0)=0$, and satisfies incremental $\rho$-hard IQCs with filters started at zero; the $N\ge2$ layers have activations slope-restricted in $[0,1]$ with $\phi(0)=0$; the network input is a linear function of the plant-and-filter state $\zeta$, of $q$ and of $d$ (no direct dependence on $w$), and this map is absorbed into the first weight, giving $\widetilde W_1$. Stack $z=(\delta\zeta,\delta q,\delta d)\in\mathbb R^m$ and $w=\delta h^N\in\mathbb R^{n_N}$. For a storage $V=\delta\zeta^{\top}P\delta\zeta$ with $P\succ0$, IQC weights $\tau\ge0$ (their $\alpha$) and $\gamma\gt0$, collect the plant part of the dissipation inequality in the matrix $\begin{bmatrix}\mathscr A&\mathscr B\\ \mathscr B^{\top}&\mathscr C\end{bmatrix}$ whose quadratic form in $(z,w)$ is $V(\delta\zeta^+)-V(\delta\zeta)+\delta r^{\top}M(\tau)\delta r+\|\delta e\|^2-\gamma\|\delta d\|^2$; it is affine in $(P,\tau,\gamma)$.

Statement. Let diagonal $\Lambda_1,\dots,\Lambda_{N-1}\succeq0$ and a symmetric $Y_1$ make the network QC matrix of the first $N-1$ layers (built with $\widetilde W_1$) plus $\mathrm{diag}(Y_1,0)$ negative definite, a Lipschitz-type certificate that scalable estimators (ECLipsE and its variants) compute layer by layer. Define

$$\begin{gathered}\Xi_1=Y_1,\qquad\Gamma_1=I_m,\\ \Xi_2=-2\Lambda_1-\Lambda_1\widetilde W_1\Xi_1^{-1}\widetilde W_1^{\top}\Lambda_1,\qquad\Gamma_2=-\Lambda_1\widetilde W_1\Xi_1^{-1},\\ \Xi_l=-2\Lambda_{l-1}-\Lambda_{l-1}W_{l-1}\Xi_{l-1}^{-1}W_{l-1}^{\top}\Lambda_{l-1}\quad(3\le l\le N),\\ \Gamma_l=-\Lambda_{l-1}W_{l-1}\Xi_{l-1}^{-1}\Gamma_{l-1}\quad(3\le l\le N),\\ D=\sum_{l=1}^{N-1}\Gamma_l^{\top}\Xi_l^{-1}\Gamma_l\end{gathered}$$

(every $\Xi_l\prec0$ by successive Schur complements, and $D\prec0$ because its first term is $Y_1^{-1}$). If some $P\succ0$, $\tau\ge0$, $\gamma\gt0$, diagonal $\Lambda_N\succeq0$, symmetric $Y_2$ and $Y_3\in\mathbb R^{m\times n_N}$ satisfy

$$\begin{gathered}\begin{bmatrix}\mathscr A-Y_1&\mathscr B-Y_3\\ \mathscr B^{\top}-Y_3^{\top}&\mathscr C-Y_2\end{bmatrix}\prec0,\\ \begin{bmatrix}\Xi_N&W_N^{\top}\Lambda_N+\Gamma_NY_3&0\\ \Lambda_NW_N+Y_3^{\top}\Gamma_N^{\top}&Y_2-2\Lambda_N&Y_3^{\top}\\ 0&Y_3&D^{-1}\end{bmatrix}\prec0,\end{gathered}$$

then the full-order certificate holds, so for $d=0$ any two state trajectories converge to each other, and $\sum_{k\le K}\|\delta e_k\|^2\le\gamma\sum_{k\le K}\|\delta d_k\|^2+V(\delta\zeta_0)$ for every $K$: incremental $\ell_2$ gain at most $\sqrt\gamma$ (here, as in the paper, $\gamma$ bounds the squared gain, unlike the notation box).

Why it works. The full LMI is a plant part plus a network part. Subtracting a coupling matrix $Y=\begin{bmatrix}Y_1&Y_3\\ Y_3^{\top}&Y_2\end{bmatrix}$ from the first and adding it to the second changes nothing when the two are added back, and the full LMI holds if and only if both split LMIs hold for some $Y$ (their Lemma 1). Eliminating the inner network blocks by successive Schur complements (the $\Xi_l,\Gamma_l$ recursion) turns the network part into the second LMI above. The two LMIs have sizes $m+n_N$ and $m+n_{N-1}+n_N$, whatever the depth.

Derivation — The alternating coupling updates (their Sec. V-B and Alg. 1)

Fix $\gamma$ and $Y_1\prec0$, compute $\Lambda_1,\dots,\Lambda_{N-1}$ with a scalable estimator, and minimise $s$ over $(P,\tau,\Lambda_N,Y_2,Y_3)$ subject to the first reduced LMI relaxed to $\preceq sI$ and the second kept strict. If $s\lt0$ (with a numerical margin), the certificate is found. Otherwise fix all $\Lambda_l$, $Y_2$, $Y_3$ and eliminate the network blocks from the output layer backwards ($2\le l\le N$):

$$\begin{gathered}\Xi'_1=Y_2-2\Lambda_N,\qquad\Gamma'_1=I_{n_N},\\ \Xi'_l=-2\Lambda_{N-l+1}-W_{N-l+2}^{\top}\Lambda_{N-l+2}(\Xi'_{l-1})^{-1}\Lambda_{N-l+2}W_{N-l+2},\\ \Gamma'_l=-\Gamma'_{l-1}(\Xi'_{l-1})^{-1}\Lambda_{N-l+2}W_{N-l+2},\\ K=\widetilde W_1^{\top}\Lambda_1+Y_3\Gamma'_N,\\ \Theta=\sum_{l=1}^{N-1}(Y_3\Gamma'_l)(\Xi'_l)^{-1}(Y_3\Gamma'_l)^{\top}+K(\Xi'_N)^{-1}K^{\top}.\end{gathered}$$

With all $\Xi'_l\prec0$, the network LMI is equivalent to $Y_1\prec\Theta$, so one minimises $s$ over $(P,\tau,Y_1)$ subject to $Y_1\prec\Theta$ and the plant LMI $\preceq sI$. The old $Y_1$ is feasible, so this step cannot increase $s$. Then the upstream certificate is recomputed for the new $Y_1$ (which must stay $\prec0$ for the estimators) and the rounds repeat up to a limit. Recomputing the upstream certificate carries no such monotonicity guarantee, and running out of rounds returns "unknown", not "infeasible".

Open problems (as of 2026)
  • Richer incremental multipliers. The repeated-nonlinearity (coupled) multipliers fail incrementally; which larger classes are valid for Lipschitz and contraction certificates, and how tight can they be?
  • Synthesis without BMIs beyond full-order changes of variables and direct parameterisations: reduced-order and structured controllers, output feedback with uncertain models, dynamic multipliers optimised jointly with the network.
  • Local certificates with constraints: saturation, state constraints and safe sets (Module 10) together with regions of attraction, for moving references and in the presence of disturbances.
  • Scale and architecture: convolutional and state-space layers (the 2-D systems view of Module 12), deep networks without layer-wise blow-up, and nonlinear plants beyond sector-bounded IQC models.
  • Learned plants: combining statistical model-error bounds (Modules 3 and 11) with IQC certificates for the loop.

Walkthrough: Stability LMI for a Linear Plant With a One-Layer NN

The smallest loop in which every step of Section 2 is visible has one state and one neuron. The six steps derive the stability LMI, solve it in closed form, and turn the insight box of Section 1 into an exact statement: with a global sector, what is certified is exactly what the two extreme members of the sector allow, and only the local sector can certify an open-loop unstable plant.

What carries over to the general case
Steps 2–4 are Theorem 1 of Yin et al. with every matrix of size one: the decrease form is $R_V^{\top}[\cdot]R_V$, the QC is $R_\phi^{\top}Q_{\alpha\beta}(T)R_\phi$, and homogeneity plus the containment LMI is how the ellipsoid gets its size. What does not carry over is the closed-form answer: with several states and neurons the S-procedure combines many constraints and is only sufficient, and the explorer below has to solve the SDP numerically.

Interactive: A Neural Controller in Closed Loop

The plant is a linearised pendulum, discretised by Euler's method with $h=0.1$: $x_{t+1}=\begin{bmatrix}1&h\\ ah&1-0.5h\end{bmatrix}x_t+\begin{bmatrix}0\\ bh\end{bmatrix}u_t$ (angle and rate; unstable for $a\gt0$). The controller is a bias-free tanh network with two neurons, $u=W_1\tanh(W_0x)$ with $W_0=w_0\begin{bmatrix}1&0.5\\ 0&1\end{bmatrix}$ and $W_1=w_1\begin{bmatrix}-6&-1.5\end{bmatrix}$: a saturating PD law (see Primer D) with linearisation $u\approx W_1W_0x$ and $|u|\le7.5\,w_1$. For every setting the page solves the SDP of Section 2 in the browser: a log-barrier interior-point method for the $4\times4$ LMI in $(P,\lambda_1,\lambda_2)$ and the two $3\times3$ containment LMIs, with a margin $\mathcal M\preceq-5\cdot10^{-5}\,\mathrm{trace}(P)\,I$. Feasibility is decided by a phase-I problem: with $\mathrm{trace}(P)=1$, minimise $s$ subject to $\mathcal M\preceq sI$ and $P\succeq-sI$. The stability LMI is homogeneous in $(P,T)$, so every certificate can be scaled to $\mathrm{trace}(P)=1$, and $s\lt0$ gives $P\succeq-sI\succ0$ and $\mathcal M\preceq sI\prec0$: $s\lt0$ is exactly a certificate, which is afterwards scaled up (shrinking the ellipse) until the containment LMIs hold. The page accepts only $s\lt-2\cdot10^{-4}$. Missing that threshold is not treated as a proof that no certificate exists: the panel shows the signed value of $s$ and claims impossibility only when a linear law inside the sector is unstable beyond rounding error (a spectral radius within $10^{-9}$ of $1$ is reported as too close to decide). In local mode it bisects for the largest box $|v_i|\le\delta$ on which the sector $[\tanh(\delta)/\delta,1]$ still admits a certificate and then minimises $\mathrm{trace}(P)$ over $\delta$ by golden-section search, as Pauli et al. do. The returned $(P,T)$ is re-checked with a Jacobi eigenvalue computation (plane rotations that drive the off-diagonal entries of a symmetric matrix to zero, so that the diagonal holds its eigenvalues up to rounding). Independently of all that, the true nonlinear loop is simulated, on a grid and from points on the certified ellipse.

Derivation — The explorer's LMI in its five unknowns

The unknowns are $(p_{11},p_{12},p_{22},\lambda_1,\lambda_2)$, with $P=\begin{bmatrix}p_{11}&p_{12}\\ p_{12}&p_{22}\end{bmatrix}$ and $T=\mathrm{diag}(\lambda_1,\lambda_2)$. With $G=[A\ \ BW_1]$ (so $x^+=G(x,w)$) and $F=\mathrm{diag}(W_0,I_2)$ (so $(v,w)=F(x,w)$), the stability matrix is $\mathcal M=G^{\top}PG-\mathrm{diag}(P,0)+F^{\top}Q_{\alpha\beta}(T)F$ with $\beta=1$, a $4\times4$ matrix affine in the unknowns. No separate constraint $\lambda_i\ge0$ is needed: the $(w_i,w_i)$ diagonal entry of $\mathcal M$ is $(BW_1)_i^{\top}P(BW_1)_i-2\lambda_i$ ($(BW_1)_i$ the $i$-th column), and it is negative when $\mathcal M\prec0$, so $\lambda_i\gt\frac12(BW_1)_i^{\top}P(BW_1)_i\ge0$ for every accepted point.

Derivation — Why a small enough box always admits a certificate here

Assume the linearised loop $A_c=A+BW_1W_0$ is Schur stable. On the box the sector is $[\alpha,1]$ with $\alpha=\tanh(\delta)/\delta\to1$ as $\delta\downarrow0$. Write each $w_i=mv_i+dq_i$ with $m=(1+\alpha)/2$ and $d=(1-\alpha)/2$; the sector says exactly $|q_i|\le|v_i|$, and with $T=\tau I$ the sector QC equals $2\tau d^2(\|v\|^2-\|q\|^2)$ because $w-\alpha v=d(v+q)$ and $v-w=d(v-q)$. Choose $P\succ0$ with $A_c^{\top}PA_c-P=-I$. At $d=0$ the loop is $x^+=A_cx$, and the stability form in $(x,q)$ plus $\lambda(\|W_0x\|^2-\|q\|^2)$ has the blocks $-I+\lambda W_0^{\top}W_0$ and $-\lambda I$, negative definite for $0\lt\lambda\lt1/\|W_0\|^2$. The matrix depends continuously on $d$, so it stays negative definite for all small $d\gt0$, i.e. for all small boxes; the change of variables $w=mW_0x+dq$ turns it into the LMI of Section 2 with $T=\frac{\lambda}{2d^2}I$. Scaling $(P,T)$ up then shrinks the ellipse into the box. The browser can still miss such a certificate: its smallest box is $\delta=0.02$ and it demands a margin.

Numerically certified ellipse vs finite simulation
Computing…
certified set $\mathcal E(P,0)$ (numerical check) / / simulation near zero / criterion not met or escaped / undecided within the step budget box $|v_i|\le\delta$ where the local sectors hold / trajectories from inside / outside the certificate
Caveat — Reading the explorer
The shading is a finite simulation of the true nonlinear loop from a $60\times40$ grid (600 steps per point, up to 6000 when the loop decays slowly). Green: $|x_1|+|x_2|$ fell below $5\%$ of the window size within that budget. Red: the state grew beyond $1000$ times the window size, or the linearised loop is not Schur stable and the state did not come close. Grey: neither, within the budget (only when the linearised loop is Schur stable). No colour is an infinite-time statement. The ellipse is a numerical certificate; were it exact, it could contain no red point, since its points never leave it and a certificate implies a Schur-stable linearisation ($w=v$ lies in every sector used). In slow, nearly marginal settings part of it may be grey, because the decay guaranteed by the certificate can take thousands of steps. The largest ratio $V(x_{t+1})/V(x_t)$ in the panel is measured on the simulated starts only; decrease everywhere on the ellipse comes from the matrix inequality. The converse is not required, and the gap between the ellipse and the green region is the conservatism of a quadratic $V$, of diagonal static multipliers and of the box. The certificate is for the model: an Euler-discretised linear plant with exactly these weights. The margin and the eigenvalue re-check reduce sensitivity to rounding, but do not bound floating-point error: this is not a verified-arithmetic proof.

From the mathematics to a real decision

Learning objectives

A commissioning decision

The enclosure's temperature reference is fixed at 40 degrees Celsius. Let $e_t=(T_t-40)/10$ be normalized temperature error, with one-minute sampling. A centered heater command $u_t$ is measured in kilowatts; negative command means reducing power relative to the equilibrium feedforward setting. Assume the plant and controller satisfy

$$e_{t+1}=0.9e_t+0.2u_t+w_t,\qquad u_t=-1.5\tanh(e_t).$$

The centered actuator permits $|u_t|\le0.8\,\mathrm{kW}$. The desired operating region is $|e_t|\le0.5$, corresponding to 35–45 degrees. Disturbance contributes at most $|w_t|\le0.02$ per update, or 0.2 degrees Celsius. Assume the plant equation is exact inside this region, the equilibrium feedforward really cancels the fixed ambient load, the controller has no delay, and the sensor initially reports $e_t$ exactly.

This is a deliberately simple neural nonlinearity. It isolates the feedback calculation without attributing a deployment guarantee to a trained network. The actuator requirement also makes the certificate local in its physical applicability even though the mathematical tanh expression is defined globally.

Worked decision, with its limits

Close the loop. The disturbance-free map is $F(e)=0.9e-0.3\tanh(e)$. Its derivative is $F'(e)=0.9-0.3\operatorname{sech}^2(e)$, which lies between 0.6 and 0.9. Since $F(0)=0$, the mean-value theorem gives $|F(e)|\le0.9|e|$. Thus a Euclidean storage $V(e)=e^2$ contracts by a factor at most $0.9^2=0.81$ in the nominal loop.

Add the disturbance before claiming containment. The physical loop obeys $|e_{t+1}|\le0.9|e_t|+0.02$. At the proposed boundary 0.5, the right side is 0.47, inside the region. Induction establishes sampled state containment from every allowed initial error. The limiting bound is $0.02/(1-0.9)=0.2$, or 2 degrees, rather than convergence to the reference under every persistent disturbance.

Check the actuator on that same region. For $|e|\le0.5$, the command magnitude is at most $1.5\tanh(0.5)\approx0.693176\,\mathrm{kW}$, below 0.8. Thus saturation does not change the assumed closed-loop equation on the invariant region. A global claim based on the unsaturated tanh controller would require more care, because its limiting command magnitude is 1.5 kilowatts.

Inspect a physical starting point. At 44 degrees, $e_0=0.4$. The nominal next error is $0.36-0.3\tanh(0.4)\approx0.246015$, giving temperature $42.460153$ degrees. Its energy changes from 0.16 to approximately 0.060524. With the worst positive disturbance, the next error is 0.266015, still covered. The uniform 0.47 boundary calculation is stronger evidence for the whole region than this one appealing trajectory.

The commissioning conclusion is sampled containment and an ultimate disturbance bound for the stated fixed reference. Continuous-time temperatures, setpoint changes, sensor errors, and controller timing are not silently included. Each can be incorporated, but the equations and resulting certificate must then be revised.

A tempting wrong approach

Common mistake — A controller gain decides stability

The controller alone has sensitivity at most 1.5 kilowatts per normalized error unit. Calling that gain safe or unsafe without the plant ignores both the negative feedback sign and the plant's 0.2 coefficient. The closed map has derivative at most 0.9. Conversely, an actuator saturation or delay can change that map even if the same network weights are deployed.

Transfer the argument

Exercise 14.B1 — Medium: Carry a bounded sensor bias through feedback

The sensor reports $e+b$ with $|b|\le0.03$, corresponding to 0.3 degrees. Keep the disturbance bound 0.02. Derive a sufficient ultimate error bound, test invariance of $|e|\le0.5$, and check the actuator limit.

Review: Containment of a certified region.

Show hint

Tanh is 1-Lipschitz. Treat the change from $\tanh(e)$ to $\tanh(e+b)$ as an additional bounded input to the closed map.

Show worked solution

The extra next-error magnitude is at most $0.3|b|\le0.009$. Therefore $|e_{t+1}|\le0.9|e_t|+0.029$, with ultimate bound 0.29, or 2.9 degrees. At the boundary the bound is $0.45+0.029=0.479\le0.5$, preserving invariance. The command magnitude is at most $1.5\tanh(0.53)\approx0.728072$ kilowatts, still below 0.8. Constant bias need not permit zero tracking error; the correct conclusion is the revised bound.

Exercise 14.B2 — Hard: A stable delayed loop can leave the temperature box

Consider instead the linear delayed controller $u_t=-1.5e_{t-1}$ with zero disturbance. Use augmented state $(e_t,e_{t-1})$ and matrix $A_d=\left[\begin{smallmatrix}0.9&-0.3\\1&0\end{smallmatrix}\right]$. Test all initial augmented states in the square, allowing controller memory to be initialized independently of the current temperature. Show that this linear loop is asymptotically stable but the square $|e_t|,|e_{t-1}|\le0.5$ is not invariant. No prior actuator-bounded history is required in this initialization model.

Review: The full closed-loop state.

Show hint

Compute the roots of $\lambda^2-0.9\lambda+0.3=0$. Then try histories with current and previous error of opposite signs.

Show worked solution

The roots are $0.45\pm0.312250i$, each with magnitude $\sqrt{0.3}\approx0.547723\lt1$. Hence all augmented trajectories converge to zero. Yet history $(0.5,-0.5)$ gives next error $0.9(0.5)-0.3(-0.5)=0.6$, outside the temperature limit. The command is 0.75 kilowatts and respects the 0.8 limit, so saturation does not explain the violation. Stability in the augmented state does not imply containment in the requested square. A valid delay-aware design needs an invariant region in both state coordinates whose projection satisfies the physical temperature limit.

Synthesis and bridge

A feedback certificate ties together an equilibrium, a plant, a controller, a state description, and an operating domain. The same weights can support a different conclusion after sensor bias or timing is changed. Storage decrease and physical containment should be read together, with the disturbance terms kept visible.

The final verification chapter returns to the neural map and asks whether a specified output inequality holds over every allowed input. Such a verifier can supply the missing local bounds used in a feedback argument, but the distinction between a function specification and an entire trajectory remains essential.

Exercises

Readiness check

Find an equilibrium, calculate a scalar Lyapunov decrease, and check an ellipsoid against a pre-activation box. Explain why local sector validity and invariance must be proved together.

If this is difficult, revisit the earlier prerequisite, work the Easy questions, and return to the linked section. You can postpone the advanced extensions while building confidence with the core certificate.

Graded practice: build the calculation, then audit the claim

These twelve new questions each include an independent hint and a fully worked solution. The difficulty measures the amount of reasoning, not the amount of notation. The original research exercises follow below.

Easy: read definitions and compute

Easy 1 — Find the equilibrium before shifting

For $x_{t+1}=0.5x_t+u_t$ and $u_t=0.25x_t+1$, find $(x_*,u_*)$ and the deviation dynamics. Is the origin an equilibrium of the unshifted system?

Review if needed: Primer D: discrete-time state equations; this module's relevant section.

Hint

Substitute the controller into the plant, then solve $x_*=x_{t+1}=x_t$.

Worked solution

Step 1. The closed loop is $x_{t+1}=0.75x_t+1$. At equilibrium $0.25x_*=1$, so $x_*=4$ and $u_*=0.25(4)+1=2$.

Step 2. Set $\tilde x_t=x_t-4$. Subtracting $4=0.75(4)+1$ gives $\tilde x_{t+1}=0.75\tilde x_t$.

Step 3. The origin is not an equilibrium: it moves to $1$ in one step. Stability must be studied around the actual equilibrium; subtracting the correct equilibrium removes the affine constants.

Easy 2 — A Lyapunov difference is a difference of energies

For $x^+=0.8x$ and $V(x)=2x^2$, compute $V(x^+)-V(x)$. Evaluate $V$ before and after one step from $x=2$.

Review if needed: Primer D: Lyapunov stability and invariant sets; this module's relevant section.

Hint

Square the factor $0.8$ when substituting into $V$.

Worked solution

Step 1. $V(x^+)=2(0.8x)^2=1.28x^2$. Thus $\Delta V=(1.28-2)x^2=-0.72x^2$, strictly negative for nonzero $x$.

Step 2. At $x=2$, $V=8$. The next state is $1.6$ and $V(x^+)=2(1.6)^2=5.12$, a decrease of $2.88$.

Step 3. The nonnegative energy decreases and the exact trajectory is $x_t=0.8^tx_0\to0$. The negative coefficient in the Lyapunov difference certifies convergence in this scalar example.

Easy 3 — An ellipsoid inside a pre-activation interval

For $P=\operatorname{diag}(4,1)$, describe the semi-axes of $x^\top Px\le1$. If a pre-activation is $v=2x_1+x_2$, find its maximum absolute value on this ellipsoid. Is the box $|v|\le1$ large enough?

Review if needed: Primer A: quadratic forms and PSD order; this module's relevant section.

Hint

The sharp squared bound is $aP^{-1}a^\top$ for the row $a=(2,1)$.

Worked solution

Step 1. The inequality is $4x_1^2+x_2^2\le1$, with semi-axes $1/2$ and $1$. The inverse is $P^{-1}=\operatorname{diag}(1/4,1)$.

Step 2. $aP^{-1}a^\top=4(1/4)+1=2$, so $|v|\le\sqrt2$. Equality occurs at $x=P^{-1}a^\top/\sqrt2=(1/(2\sqrt2),1/\sqrt2)^\top$.

Step 3. The interval $[-1,1]$ misses that attainable value. A local sector stated only there cannot be applied on the whole ellipsoid; the required half-width is at least $\sqrt2$.

Easy 4 — Two shifted ReLUs need not be repeated

Define $\tilde\varphi_1(s)=\operatorname{ReLU}(s+1)-1$ and $\tilde\varphi_2(s)=\operatorname{ReLU}(s-1)$. Check both values at $s=0$ and $s=1/2$. Are these the same scalar function?

Review if needed: Module 12: valid and invalid activation multipliers; this module's relevant section.

Hint

Both shifts remove the value at the basepoint, but the basepoints lie on different sides of the kink.

Worked solution

Step 1. At $s=0$, the values are $1-1=0$ and $\operatorname{ReLU}(-1)=0$. Both shifted maps vanish at the origin.

Step 2. At $s=1/2$, the first value is $1.5-1=0.5$, whereas the second is $\operatorname{ReLU}(-0.5)=0$. Hence the maps differ.

Step 3. Each remains slope-restricted in $[0,1]$. Individual sector/slope QCs survive the shift, but an argument that assumes one identical repeated function across these two channels does not.

Medium: connect two or three steps

Medium 1 — A sector bound and a slope bound differ

For tanh on $[-1,1]$, compute the centred lower sector bound $\alpha=\tanh(1)$ and lower slope bound $\mu=1-\tanh^2(1)$. Evaluate the $[0,1]$ sector form at $v=1,w=\tanh(1)$. Why does substituting $\mu$ for $\alpha$ lose information?

Review if needed: Primer B: derivatives and directional slopes; this module's relevant section.

Hint

A sector compares a point with the fixed equilibrium; a slope bound compares any two points.

Worked solution

Step 1. $\alpha\approx0.76159416$ is the smallest chord slope from the origin on this interval. The smallest derivative is $\mu\approx0.41997434$, attained at the endpoints.

Step 2. The global $[0,1]$ QC is $2w(v-w)$. Substitution gives $2\tanh(1)(1-\tanh(1))\approx0.363137\ge0$.

Step 3. The slope bound implies the weaker sector $[\mu,1]$, but its smaller lower endpoint admits more nonlinearities. The larger centred chord bound $\alpha$ is legitimate for the fixed equilibrium and can improve the local stability certificate.

Medium 2 — Propagate an affine interval with signed weights

A hidden-output box has centre $c=(1,-1)$ and radius vector $r=(0.2,0.5)$. The next pre-activation is $v=2w_1-3w_2+1$. Find its exact interval over this box and the corner attaining the upper endpoint.

Review if needed: Primer 0: maxima and bounds; this module's relevant section.

Hint

The centre is $ac+b$ and the radius is $|a|r$.

Worked solution

Step 1. The centre is $2(1)-3(-1)+1=6$. The radius is $2(0.2)+3(0.5)=1.9$. Thus $v\in[4.1,7.9]$.

Step 2. To maximize $v$, take $w_1=1.2$ because its coefficient is positive and $w_2=-1.5$ because its coefficient is negative. This gives $2.4+4.5+1=7.9$.

Step 3. This endpoint is exact over the independent-coordinate box. If the true hidden outputs are correlated, that corner may be unreachable; the interval remains sound but may be conservative.

Medium 3 — Why an unstable plant needs a local lower sector

Consider $x^+=1.2x-0.5w$ and the uncertainty family $w=dx$ for $d\in[\alpha,1]$, $0\le\alpha\le1$. Find when every member is asymptotically stable. Compare $\alpha=0$ and $\alpha=0.5$.

Review if needed: Primer D: discrete-time state equations; this module's relevant section.

Hint

The scalar multiplier $1.2-0.5d$ must lie strictly between $-1$ and $1$.

Worked solution

Step 1. The multiplier decreases with $d$ and ranges from $0.7$ to $1.2-0.5\alpha$. Its smallest value is already positive and below one.

Step 2. Uniform strict stability therefore requires $1.2-0.5\alpha\lt1$, or $\alpha\gt0.4$. At $\alpha=0.4$ the endpoint has multiplier one, so asymptotic stability fails.

Step 3. For $\alpha=0$ the family includes $d=0$, the unstable open loop with multiplier $1.2$. For $\alpha=0.5$, all multipliers lie in $[0.7,0.95]$. A valid local sector can exclude a controller that is effectively switched off.

Medium 4 — A hard IQC can have a negative time sample

Let $w_t=v_t/2$, $a_t=v_t-w_t$, $b_t=w_t$, with $a_{-1}=0$ and $(v_0,v_1)=(2,1)$. Evaluate $b_t(a_t-a_{t-1})$ at both times and its partial sums. Show why nonnegative sums are possible without nonnegative individual terms.

Review if needed: Primer 0: sequences and geometric sums; this module's relevant section.

Hint

Here $a_t=b_t=v_t/2$. Use $2a_t(a_t-a_{t-1})=a_t^2-a_{t-1}^2+(a_t-a_{t-1})^2$.

Worked solution

Step 1. $a_0=b_0=1$ and $a_1=b_1=1/2$. The terms are $1(1-0)=1$ and $(1/2)(1/2-1)=-1/4$. Their partial sums are $1$ and $3/4$.

Step 2. For any finite sequence, summing the identity in the hint gives $\sum_{t=0}^N a_t(a_t-a_{t-1})=\tfrac12a_N^2+\tfrac12\sum_{t=0}^N(a_t-a_{t-1})^2\ge0$. The initial square vanishes because $a_{-1}=0$.

Step 3. This is a finite-horizon, or hard, IQC for this relation. A Lyapunov proof may use its summed nonnegativity; deleting a negative sample as though the IQC held pointwise would change the argument.

Hard: combine calculations with assumptions

Hard 1 — Close the local-sector invariance argument

A scalar map is $x^+=x/2$ for $|x|\le1$ and $x^+=2x$ for $|x|\gt1$. With $V=x^2$ prove that $\{V\le1\}$ is invariant and contained in the ROA of zero. Does the same local calculation prove convergence from $x_0=1.1$?

Review if needed: Primer 0: inequalities and proof steps; this module's relevant section.

Hint

Use the local rule at the current state, then prove the next state remains in its domain.

Worked solution

Step 1. If $V(x_t)\le1$, then $|x_t|\le1$, so the local rule applies and $V(x_{t+1})=V(x_t)/4\le1/4\le1$. This proves invariance by induction from any initial state in the set.

Step 2. Induction also gives $V(x_t)=4^{-t}V(x_0)$, hence $x_t\to0$. Thus the entire sublevel set is in the ROA, using only the rule valid there.

Step 3. At $x_0=1.1$ the exterior rule applies and $x_t=1.1\,2^t$, which diverges. A certificate based on local constraints needs the containment/invariance step; its conclusion extends only over the certified initial-state set.

Hard 2 — Check a stability LMI entirely by hand

For $x^+=1.2x-0.5w$ and the sector $w/x\in[0.5,1]$, use $V=x^2$ and multiplier $T=1$. Expand $\Delta V+2(w-0.5x)(x-w)$ as a quadratic form. Prove its matrix is negative definite and explain the stability conclusion when the sector is local.

Review if needed: Module 2: the S-procedure; this module's relevant section.

Hint

For a symmetric $2\times2$ matrix, a negative first diagonal entry and positive determinant imply negative definiteness.

Worked solution

Step 1. The Lyapunov difference is $0.44x^2-1.2xw+0.25w^2$. The QC is $-x^2+3xw-2w^2$. Their sum is $-0.56x^2+1.8xw-1.75w^2$.

Step 2. The matrix is $M=\begin{bmatrix}-0.56&0.9\\0.9&-1.75\end{bmatrix}$. Its determinant is $0.98-0.81=0.17\gt0$ and its first diagonal is negative, so $M\prec0$ (eigenvalues approximately $-2.23390,-0.0761001$).

Step 3. On the sector, the QC is nonnegative, so $\Delta V\le\Delta V+\mathrm{QC}\lt0$ away from zero. For a local sector, one must also exhibit a sublevel set contained in its pre-activation domain and use invariance; the LMI alone does not supply a global conclusion.

Hard 3 — A disturbance changes convergence into a bound

Suppose $V(x)=\|x\|_2^2$ and a closed-loop certificate proves $V_{t+1}\le0.5V_t+2d_t^2$ with $|d_t|\le0.1$. Starting with $V_0\le0.25$, derive a finite-time bound, prove invariance of $V\le0.25$, and state the limiting state-radius bound.

Review if needed: Primer D: Lyapunov stability and invariant sets; this module's relevant section.

Hint

Unroll the recursion and sum a geometric series. A persistent disturbance need not vanish.

Worked solution

Step 1. The disturbance term is at most $0.02$. Recursion gives $V_t\le0.5^tV_0+0.02\sum_{j=0}^{t-1}0.5^j=0.5^tV_0+0.04(1-0.5^t)$.

Step 2. If $V_t\le0.25$, then $V_{t+1}\le0.125+0.02=0.145\le0.25$. Hence the sublevel set is invariant for every allowed disturbance sequence.

Step 3. Taking a limit superior gives $\limsup_tV_t\le0.04$, so $\limsup_t\|x_t\|_2\le0.2$. This proves a disturbance-dependent ultimate bound, not convergence to zero under arbitrary persistent disturbance.

Hard 4 — Audit three tempting certificate shortcuts

A reviewer proposes: (i) couple all shifted ReLU channels with equilibrium pre-activations $(0,0,1)$ as one repeated nonlinearity; (ii) treat every time sample of a hard IQC as nonnegative; (iii) reuse a fixed-network stability certificate after a small unbounded-in-the-proof weight update. For each proposal identify the missing justification and a valid repair.

Review if needed: Module 12: valid and invalid activation multipliers; this module's relevant section.

Hint

Check which functions are identical, what time quantifier an IQC has, and which matrices the certificate actually covers.

Worked solution

Step 1. (i) Channels 1 and 2 have the identical shifted map $\operatorname{ReLU}(s)$; channel 3 has $\operatorname{ReLU}(s+1)-1$. At $s=-1/2$ their values are $0$ and $-1/2$. Coupling justified by repetition may be used within the equal-shift group, or replaced by per-channel diagonal constraints across all three.

Step 2. (ii) A hard IQC guarantees each finite-horizon sum is nonnegative. Individual terms may be negative. Sum the dissipation inequality with the IQC, retaining its filter state and initial-state assumptions, instead of asserting pointwise nonnegativity.

Step 3. (iii) The changed weights change the LMI and possibly local bounds/equilibria. Recompute and recheck the certificate, or establish a perturbation bound smaller than a proven feasibility margin. A small numerical update alone does not prove that the old certificate applies.

Original research exercises

Continue here when the core calculations and the certificate assumptions are clear. Use the graded questions above as a warm-up; the original derivations below remain available in full.

Exercise 14.1 — A two-layer network loop in LFT form

Plant $x_{t+1}=Ax_t+Bu_t$, $x\in\mathbb R^{n_x}$; controller $v^1=W_0x+b_0$, $w^1=\phi(v^1)$, $v^2=W_1w^1+b_1$, $w^2=\phi(v^2)$, $u=W_2w^2+b_2$ with widths $n_1,n_2$. (a) Write the loop in the form of Section 1 with $v=(v^1,v^2)$, $w=(w^1,w^2)$. (b) Why is it well posed? (c) Without biases, give $R_V$ and $R_\phi$ of Theorem 1 and the size of the stability LMI. (d) What changes for output feedback $u=\pi(Cx)$, and for nonzero biases?

Show answer
$$\begin{bmatrix}u\\ v^1\\ v^2\end{bmatrix}=\begin{bmatrix}0 & 0 & W_2 & b_2\\ W_0 & 0 & 0 & b_0\\ 0 & W_1 & 0 & b_1\end{bmatrix}\begin{bmatrix}x\\ w^1\\ w^2\\ 1\end{bmatrix}$$

(a) So $N_{ux}=0$, $N_{uw}=[0\ \ W_2]$, $N_{vx}=\begin{bmatrix}W_0\\ 0\end{bmatrix}$, $N_{vw}=\begin{bmatrix}0 & 0\\ W_1 & 0\end{bmatrix}$, $N_{ub}=b_2$, $N_{vb}=(b_0,b_1)$, and $x_{t+1}=Ax_t+BW_2w^2_t+Bb_2$. (b) $N_{vw}$ is strictly block lower triangular: $v^1$ depends on $x$ only, then $w^1$, then $v^2$ and $w^2$. Forward substitution computes $w$ for any activation; there is no algebraic loop.

(c) With $\zeta=(x,w^1,w^2)$, $R_V\zeta=(x,u)$ and $R_\phi\zeta=(v^1,v^2,w^1,w^2)$:

$$R_V=\begin{bmatrix}I & 0 & 0\\ 0 & 0 & W_2\end{bmatrix},\qquad R_\phi=\begin{bmatrix}W_0 & 0 & 0\\ 0 & W_1 & 0\\ 0 & I & 0\\ 0 & 0 & I\end{bmatrix}.$$

The first LMI has size $n_x+n_1+n_2$ and $n_x(n_x+1)/2+n_1+n_2$ variables; the local version adds $n_1$ containment LMIs of size $1+n_x$. (d) Output feedback replaces $W_0$ by $W_0C$ (Yin et al., Remark 2). With biases, first solve $x_*=Ax_*+B\pi(x_*)$ and shift: the bias column drops out, but neuron $i$ now carries $\tilde\varphi_i(s)=\varphi(s+v_{*,i})-\varphi(v_{*,i})$, so the sectors are offset sectors computed per neuron (Algorithm 1) and multipliers that rely on a repeated nonlinearity are in general no longer valid, except between neurons with equal $v_{*,i}$ and common slope bounds (Exercise 14.6, Section 3).

Exercise 14.2 — The sector QC of tanh, global and local (proof)

(a) Prove that $\tanh$ is slope-restricted and sector-bounded in $[0,1]$, and write the sector condition as a QC. (b) Assume $\bar v\gt0$. Prove that on $|v|\le\bar v$ the tightest centred sector is $[\tanh(\bar v)/\bar v,\,1]$ and the tightest slope bounds are $[1-\tanh^2\bar v,\,1]$. (c) Compare the lower bounds for $\bar v=1,2$; which one may Theorem 1 of Yin et al. use, and why?

Show answer

(a) $\tanh'(s)=1-\tanh^2s\in(0,1]$. By the mean value theorem every chord slope equals some $\tanh'(\xi)\in(0,1]$: slope $[0,1]$. The chord to the origin ($\tanh0=0$) gives $0\lt\tanh(v)/v\le1$: sector $[0,1]$. With $w=\tanh v$: $w(v-w)\ge0$, i.e. $\begin{bmatrix}v\\ w\end{bmatrix}^{\top}\begin{bmatrix}0 & 1\\ 1 & -2\end{bmatrix}\begin{bmatrix}v\\ w\end{bmatrix}\ge0$, the matrix of Section 1 with $\alpha=0$, $\beta=1$.

(b) $r(v)=\tanh(v)/v$ ($r(0)=1$) is even, and for $v\gt0$, $r'(v)=-h(v)/v^2$ with $h(v)=\tanh v-v(1-\tanh^2v)$. As $h(0)=0$ and $h'(v)=2v\tanh v\,(1-\tanh^2v)\gt0$, $h\gt0$ and $r$ decreases strictly on $(0,\infty)$. So $r(v)\in[r(\bar v),1)$ on $0\lt|v|\le\bar v$, with the lower end attained at $\pm\bar v$ and the upper approached as $v\to0$: the tightest sector is $[\tanh(\bar v)/\bar v,1]$. $\tanh'$ is even and decreasing in $|s|$, and a chord slope is the mean of $\tanh'$ over a subinterval of $[-\bar v,\bar v]$, so it lies in $[1-\tanh^2\bar v,1]$; short chords near $\pm\bar v$ and near $0$ approach both ends.

(c) $\bar v=1$: $0.762$ (sector) vs $0.420$ (slope); $\bar v=2$: $0.482$ vs $0.071$. The sector bound is always larger, since $\tanh(\bar v)/\bar v=\frac1{\bar v}\int_0^{\bar v}\tanh'(s)\,ds\ge\tanh'(\bar v)$. Theorem 1 uses a non-incremental QC relative to the equilibrium, so the tighter sector bound is valid there, but only on the box (hence the containment LMI). Slope bounds are needed wherever increments between two arbitrary points appear: Zames–Falb multipliers (Section 3), incremental gains and contraction (Sections 5 and 8), and certificates around a whole family of equilibria (Section 4).

Exercise 14.3 — Solve the walkthrough LMI numerically

(a) For $a=0.8$, $b=1$, $W_0=2$, $W_1=-0.6$: decide feasibility with Step 5, compute $\lambda^\star$ for $p=1$, $\mathcal M(1,\lambda^\star)$ and its eigenvalues, and all admissible $\lambda$. (b) Same network, $a=1.3$: show that the global LMI fails, find the largest interval certified with local sectors and check it by simulation. (c) Give an explicit multiplier for the box $\bar v=3$.

Show answer

(a) $g=bW_1=-0.6$, $c=gW_0=-1.2$; $|a|\lt1$ and $a+c=-0.4$: feasible. $\lambda^\star=(1-0.64+0.96)/4=0.33$ and

$$\mathcal M(1,0.33)=\begin{bmatrix}-0.36 & -0.48+0.66\\ -0.48+0.66 & 0.36-0.66\end{bmatrix}=\begin{bmatrix}-0.36 & 0.18\\ 0.18 & -0.30\end{bmatrix},$$

with eigenvalues $-0.148$, $-0.512$ and determinant $0.0756=\big(1.32^2-1.44\big)/4$, as Step 5 predicts. The determinant $-4\lambda^2+2.64\lambda-0.36$ is positive for $\lambda\in(0.193,\,0.467)$, where the $(2,2)$ entry $0.36-2\lambda$ is negative too: every such $\lambda$ (with $p=1$, or any positive multiple of the pair) certifies global exponential stability.

(b) The $(1,1)$ entry is $0.69p\gt0$: infeasible. The linearisation $a+c=0.1$ is stable and $(a-1)/|c|=0.25$, so $\tanh(\bar v^\star)/\bar v^\star=0.25$ gives $\bar v^\star=3.997$: every $|x_0|\lt\bar v^\star/|W_0|=1.9987$ is certified. Simulating $x^+=1.3x-0.6\tanh(2x)$: $x_0=1.9985$ converges and $x_0=1.999$ diverges, because the nonzero equilibria sit at $\pm1.9987$. The certificate is exact.

(c) $\alpha=\tanh(3)/3=0.332$, $a+\alpha c=0.902\lt1$. After the loop transformation $a'=0.902$, $c'=(1-\alpha)c=-0.802$, so $\lambda'^\star=(1-a'^2-a'c')/W_0^2=0.227$. Since $\hat w(v-\hat w)=(w-\alpha v)(v-w)/(1-\alpha)^2$, the multiplier in the original coordinates is $\lambda=\lambda'^\star/(1-\alpha)^2=0.509$. Check: $\begin{bmatrix}a^2-1-2\alpha\lambda W_0^2 & ag+(1+\alpha)\lambda W_0\\ \cdot & g^2-2\lambda\end{bmatrix}=\begin{bmatrix}-0.661 & 0.576\\ 0.576 & -0.658\end{bmatrix}$ has eigenvalues $-0.084$, $-1.236$ (any $\lambda\in(0.269,0.750)$ works). Rescaled to $p=W_0^2/\bar v^2=4/9$, $\lambda=0.226$, the containment LMI holds with equality and $|x|\le1.5$ is certified.

Exercise 14.4 — Why a local sector needs an invariant set (proof)

(a) A colleague argues: "The sector $[\alpha,1]$ holds on the box $\mathcal B=\{x:|v_i|\le\delta\ \forall i\}$ and the first LMI of Theorem 1 holds with it, so $V$ decreases on $\mathcal B$ and every $x_0\in\mathcal B$ converges." Find the gap. (b) Test the claim on the explorer's default loop ($a=2$, $b=1$, $w_0=w_1=1$, local mode, $\delta\approx2.87$) at $x_0=(-4.25,\,2.80)$. (c) Explain how the containment LMI and the induction of Section 2 close the gap, and why that is not circular.

Show answer

(a) "$V$ decreases on $\mathcal B$" means: if $x_t\in\mathcal B$ then $V(x_{t+1})\lt V(x_t)$. It does not give $x_{t+1}\in\mathcal B$. The box is a parallelogram, not a sublevel set of $V$, so a step can lower $V$ and still leave $\mathcal B$; outside, the sector is false and the LMI says nothing. The argument assumes what it must prove: that the state stays where the QC holds.

(b) $v_1=x_1+0.5x_2=-2.85$ and $v_2=2.80$, so $x_0\in\mathcal B$. With the explorer's $P\approx\begin{bmatrix}0.162 & 0.060\\ 0.060 & 0.144\end{bmatrix}$, however, $V(x_0)\approx2.63\gt1$: $x_0$ lies outside the certified ellipse. In simulation, $V$ decreases while the state is in $\mathcal B$, the state leaves $\mathcal B$ at $t=4$ ($|v_1|\gt\delta$), $V$ bottoms out near $1.46$ around $t=11$, and the trajectory then diverges. (It started beyond the stable manifold of the saddle equilibrium $(-2.985,0)$, where $2x_1=6\tanh x_1$: a saddle attracts along one direction and repels along another, and its stable manifold is the set of initial states that converge to it. Only the facts that $x_0$ lies in the box, outside the ellipse, and diverges matter here.) So a point of the box is not in the region of attraction; the box corners $(\mp4.31,\pm2.87)$ diverge as well.

(c) The containment LMI gives $\mathcal E(P,0)\subseteq\mathcal B$ (Cauchy–Schwarz, Step A of the proof in Section 2). Induction: $x_0\in\mathcal E$ by assumption; if $x_t\in\mathcal E$ then $x_t\in\mathcal B$, so the QC holds at time $t$, so $V(x_{t+1})\le V(x_t)-\varepsilon\|x_t\|^2\le1$ and $x_{t+1}\in\mathcal E$. Each step uses the QC only at the current state, whose membership in $\mathcal B$ was established one step earlier. The ellipse works because it is both a sublevel set of $V$ (decrease keeps you inside) and a subset of $\mathcal B$ (the QC is valid inside). In the scalar walkthrough the two sets coincide, which is why no gap appeared there.

Exercise 14.5 — Offset-free tracking needs integral action, even with a perfect network

Scalar plant $x_{t+1}=ax_t+bu_t+d$, $y=x$, nominal $a=0.5$, $b=1$, $d=0$. The network $\kappa(x,r)=(1-a)r/b-0.3\tanh(x-r)$ is "perfect" for the nominal model: $\kappa(x_*(r),r)=u_*(r)$ with $x_*(r)=r$ for every $r$. (a) Show that with $d=0.1$, or with the true gain $b=1.2$, the output settles at a wrong value; compute both offsets for $r=1$. (b) Add $\xi_{t+1}=\xi_t+r-y_t$ and $u=k_\xi\xi+\kappa(x,r)$ with $k_\xi=0.2$: show that every equilibrium has $y=r$ and check stability. (c) What must the certificate of Section 4 prove, and what not?

Show answer

(a) An equilibrium satisfies $x_\infty=ax_\infty+b\kappa(x_\infty,r)+d$; let $e=x_\infty-r$. Nominal $b$, $d=0.1$: $0.5e+0.3\tanh e=0.1$, so $e\approx0.1252$ (linearised: $d/(1-a+0.3b)=0.125$). True $b=1.2$, $d=0$: $(1-a)(r+e)=1.2\big[(1-a)r-0.3\tanh e\big]$, i.e. $0.5e+0.36\tanh e=0.1r$, so $e\approx0.1165$ for $r=1$. The loop converges (the map $x\mapsto ax+b\kappa(x,r)+d$ has slope in $[0.14,0.5]$), but to $y\ne r$: zero offset needs $\kappa(x_*(r),r)=u_*(r)$ for the true plant and disturbance, which no network fitted to a model can guarantee.

(b) At an equilibrium $\xi_{t+1}=\xi_t$, hence $r-y=0$; the argument uses neither $a$, $b$, $d$ nor $\kappa$. The equilibrium exists: $x=r$ and $\xi=\big[(1-a)r-b\kappa(r,r)-d\big]/(bk_\xi)$, e.g. $\xi=-0.833$ for $b=1.2$, $d=0.1$, $r=1$. The linearisation $\begin{bmatrix}a-0.3b & bk_\xi\\ -1 & 1\end{bmatrix}$ has spectral radius $0.632$ ($b=1$) and $0.616$ ($b=1.2$), so the equilibrium is locally exponentially stable (the closed-loop map is continuously differentiable); this says nothing about convergence from far away, which needs the certificate of Section 4. Simulated from rest, $y\to1.000000$ for $d=0.1$ with $b=1$ or $b=1.2$, and $y\to-0.5$ for $r=-0.5$.

(c) It must prove convergence of the augmented loop: for all references (Theorem 1 of Pauli et al., via slope restriction), for one reference (Theorem 3) or for a set of references (Theorem 4). Zero offset is then structural, and the network's steady-state map needs no constraint. With model error or disturbances, the stability certificate must cover the perturbed plant (e.g. with the IQC version of Section 2), but offset-freeness holds whenever the true loop converges.

Exercise 14.6 — Coupled multipliers after an equilibrium shift (counterexample)

Two tanh neurons of a biased network have equilibrium pre-activations $v_{*,1}=0$ and $v_{*,2}=2$. (a) Write the shifted activations; show that each is slope-restricted in $[0,1]$ and vanishes at $0$. (b) The repeated-nonlinearity class would admit the doubly hyperdominant $\mathsf M=\begin{bmatrix}1 & -1\\ -1 & 1\end{bmatrix}$, i.e. the QC $\tilde v^{\top}\mathsf M\tilde\phi(\tilde v)\ge0$. Show that it fails at $\tilde v=(-1,-1.5)$. (c) Why does it hold for a bias-free network at $x_*=0$, and which multipliers survive the shift?

Show answer

(a) $\tilde\varphi_1(s)=\tanh s$, $\tilde\varphi_2(s)=\tanh(s+2)-\tanh2$. Every chord of $\tilde\varphi_2$ is a chord of $\tanh$, so both are slope-restricted in $[0,1]$, and $\tilde\varphi_i(0)=0$.

(b) $\tilde v^{\top}\mathsf M\tilde\phi(\tilde v)=(\tilde v_1-\tilde v_2)\big(\tilde\varphi_1(\tilde v_1)-\tilde\varphi_2(\tilde v_2)\big)$. At $(-1,-1.5)$: $\tilde\varphi_1(-1)\approx-0.7616$ and $\tilde\varphi_2(-1.5)=\tanh0.5-\tanh2\approx-0.5019$, so the difference is approximately $-0.2597$ and the actual QC value is $0.5\big(\tanh(-1)-\tanh0.5+\tanh2\big)\approx-0.130\lt0$. A certificate that requires this QC for all shifted inputs has an invalid premise; this counterexample alone does not show that its final stability conclusion is false. Two different functions can be "out of order", which one monotone function never is.

(c) Bias-free at $x_*=0$ means $v_*=0$ for all neurons (because $\tanh0=0$, every layer's equilibrium is $0$), so $\tilde\varphi_1=\tilde\varphi_2=\tanh$ and $(\tilde v_1-\tilde v_2)(\tanh\tilde v_1-\tanh\tilde v_2)\ge0$ by monotonicity; with the sector terms this is the doubly hyperdominant lemma of Section 3. Being bias-free is sufficient, not necessary: neurons with equal $v_{*,i}$ would share one shifted function too, but here $v_{*,1}\ne v_{*,2}$. (With local slope bounds, a coupled multiplier also needs one common pair $(\mu,\nu)$ for the coupled neurons; Section 3.) Without additional shared-function assumptions, valid choices after the shift include the circle multipliers, which use each neuron's own sector (the diagonal one, here $\tilde v_1\tilde\varphi_1\approx0.762$ and $\tilde v_2\tilde\varphi_2\approx0.753$, with both actual products nonnegative, and also the full-block one, which is required to be nonnegative for every diagonal gain matrix in the sector box and therefore holds for any diagonal nonlinearity). The ordinary monotone Zames–Falb class with diagonal $M_j$ also remains valid under Section 3's coefficient signs and row/column sums, each channel's slope bounds, and zero initial filter state for the hard IQC. Time invariance keeps the same scalar function at different times; finite-horizon nonnegativity also needs these multiplier conditions. The enlarged odd-nonlinearity class additionally requires the shifted channel to be odd, which a nonzero tanh shift generally is not. The same failure of the coupled LipSDP-Network multiplier is why the offset-free analysis (Section 4) and the incremental certificates (Section 5, Module 12) use a diagonal $T$.

Key Papers

PaperVenueContributionWhy read it
Stability Analysis Using Quadratic Constraints for Systems With Neural Network Controllers (Yin, Seiler, Arcak)IEEE TAC 67(4), 2022Local offset sector QCs by interval propagation; stability LMI and ellipsoidal ROA inner approximation; IQC-robust versionThe template of Section 2 and of the walkthrough
Linear systems with neural network nonlinearities: Improved stability analysis via acausal Zames-Falb multipliers (Pauli, Gramlich, Berberich, Allgöwer)CDC 2021Hard IQCs for acausal FIR Zames–Falb multipliers; repeated-nonlinearity classes; larger certified ROAsDynamic multipliers for NN loops, and which class is valid when
Offset-free setpoint tracking using neural network controllers (Pauli, Köhler, Berberich, Koch, Allgöwer)L4DC 2021, PMLR 144Integral action plus slope QCs: one LMI for all references; local versions; reference governorStructure (an integrator) doing what training cannot
Robustness analysis and training of recurrent neural networks using dissipativity theory (Pauli, Berberich, Allgöwer)at – Automatisierungstechnik 70(8), 2022RNN robustness via dissipativity: $\mathcal H_2$ and $\ell_2$-gain measures, LMI constraints for trainingFrom feedforward Lipschitz bounds to recurrent models
Synthesizing Neural Network Controllers with Closed-Loop Dissipativity Guarantees (Junnarkar, Arcak, Seiler)Automatica 193, 2026Convexifying change of variables for full-order implicit NN controllers; RL with a dissipativity projectionSynthesis with a hard closed-loop constraint
Learning Over All Stabilizing Nonlinear Controllers for a Partially-Observed Linear System (Wang, Barbara, Revay, Manchester)IEEE L-CSS 7, 2023Nonlinear Youla parameterisation with a REN: every parameter value gives a contracting, Lipschitz closed loop, and large enough RENs approximate, on bounded inputs, every controller that achieves such a loopStability by construction instead of by constraint
Reach-SDP: Reachability Analysis of Closed-Loop Systems with Neural Network Controllers via Semidefinite Programming (Hu, Fazlyab, Morari, Pappas)CDC 2020Reachable-set over-approximation of NN loops by an SDP of QCs, incl. coupled repeated-ReLU constraintsSafety (where trajectories go), and where coupling is valid
Lyapunov-stable Neural Control for State and Output Feedback: A Novel Formulation (Yang, Dai, Shi, Hsieh, Tedrake, Zhang)ICML 2024Verifiable ROA condition on a box; neural Lyapunov functions and observers checked by α,β-CROWNThe bound-propagation alternative for nonlinear plants
Discrete-Time Stability Analysis of ReLU Feedback Systems via Integral Quadratic Constraints (Vahedi Noori, Hu, Dullerud, Seiler)arXiv 2025ReLU-specific hard IQCs containing the Zames–Falb classWhere multiplier design is heading
Scalable Incremental Robustness Analysis of Neural Network Feedback Systems (Wang, Seiler, Dullerud, Hu)arXiv 2026$\Lambda$-incremental QC; decomposition making the control LMIs depth-independentThe current answer to "does it scale?"
System analysis via integral quadratic constraints (Megretski, Rantzer)IEEE TAC 42(6), 1997The IQC framework and stability theoremThe language of Sections 2, 3 and 8
Dissipativity and Integral Quadratic Constraints: Tailored Computational Robustness Tests for Complex Interconnections (Scherer)IEEE CSM 42(3), 2022IQC theorems via dissipation and hard IQCs; tailored LMI testsThe proof route used by Yin et al. and Pauli et al.
Zames-Falb multipliers for absolute stability: From O'Shea's contribution to convex searches (Carrasco, Turner, Heath)European J. Control 28, 2016Survey of Zames–Falb multipliers and convex (FIR) searchesGet the multiplier conditions right
Neural Networks in the Loop: Learning with Stability and Robustness Guarantees (Manchester, Wang, Barbara)Annu. Rev. Control Robot. Auton. Syst. 9, 2026Survey: direct parameterisations, RENs, Youla-type policiesA map of the synthesis side of this module

Flashcards