PID Control: Theory, Tuning & Implementation
Why PID
The PID controller - three terms, two or three knobs, no model required - governs roughly 90% of industrial control loops.[fn::The oft-cited 90% figure traces to surveys by Desborough & Miller (2002) and Ender (1993); the precise number is fuzzy because "PID loop" is itself ill-defined - many loops run in PI-only mode with the derivative term switched off, and a nontrivial fraction are actually running in manual.] Its dominance is not evidence of optimality. It is evidence that most industrial plants are approximately first-order, that the engineering cost of modelling exceeds the performance gain from a model-based controller, and that PID is the most complex controller a technician can tune by hand without understanding the Laplace domain. See Control Theory: Signals & Laplace for the foundation.
The remaining 10% of loops - oscillatory plants, multivariable systems, plants with significant right-half-plane zeros - are where Bode plots, state-space methods, and LQR earn their keep. PID fails on these not because it is poorly tuned but because its structure is wrong: three parameters cannot represent the dynamics of a high-order or resonant plant.
Structure
The PID control law, in continuous time, is:
where =e(t) = r(t) − y(t)= is the tracking error, r the setpoint, y the process variable. The three terms:
- Proportional (P): Produces a control action proportional to the current error. Reduces rise time but never eliminates steady-state error for a Type-0 plant (one with no integrator in the loop).[fn::A Type-0 plant has finite DC gain. Under pure P control, the closed loop has finite DC gain < 1, so a step setpoint produces a persistent offset. The offset shrinks as Kp → ∞ but never reaches zero, and large Kp destabilises the loop.]
- Integral (I): Accumulates past error. Its transfer function
Ki/shas infinite DC gain, which forces the steady-state error to exactly zero for a step input. This is its sole raison d'être - not "correcting accumulated error" (a common but misleading textbook gloss), but injecting a pole at the origin into the loop transfer function, raising the system Type from 0 to 1. The cost is phase lag: the integrator contributes −90° at all frequencies, eroding phase margin and inviting oscillation.
- Derivative (D): Responds to the rate of change of error. In the Laplace domain it contributes
Kd·s, a zero at the origin - pure phase lead. It anticipates: if the error is closing fast, the D term brakes early, reducing overshoot. The cost is noise amplification: differentiation in time is multiplication bysin frequency, which boosts high-frequency content. Since all real signals carry noise (and noise is broadband), the raw D term is unusable without a filter. See Frequency Response: Bode & Nyquist.
Derivation from the Laplace Domain
The PID controller's transfer function is:
This is not arbitrary. Consider a general linear controller C(s) - any rational transfer function - expanded as a Laurent series about s = 0:
PID is the truncation that retains exactly three terms: the a₋₁/s (integral), a₀ (proportional), and a₁·s (derivative) coefficients. All higher-order dynamics are discarded. This is the deepest justification for PID's structure: it is the lowest-order approximation of a general controller that still guarantees (a) zero steady-state error, via the 1/s pole, (b) adjustable bandwidth, via the constant term, and (c) phase lead for damping, via the s term.[fn::This series-expansion view is sometimes attributed to Åström & Hägglund (2006, ch. 2), though the underlying idea - that PID is a Padé-like truncation - is older. The key insight is that retaining more terms (e.g., a double integrator 1/s², or a second derivative s²) would improve performance on higher-order plants, but at the cost of requiring more tuning parameters and more model knowledge, which defeats PID's purpose.]
The parallel between PID and a truncated series also explains why it works on first-order-plus-dead-time (FOPDT) plants: a FOPDT plant is itself well-approximated by a low-order Padé expansion, so a low-order controller suffices. When the plant has a resonant peak or a long dead time, the truncation error becomes significant and PID degrades. See Transfer Functions & Block Diagrams.
Tuning Methods
Ziegler–Nichols: Ultimate Sensitivity
The closed-loop (ultimate sensitivity) method: set Ki = Kd = 0, increase Kp until the loop sustains a stable oscillation. The gain at this point is Ku (ultimate gain); the period is Tu (ultimate period). Then:
The PID settings give a closed loop with approximately quarter-amplitude decay (each overshoot peak ≈ 1/4 the previous). This is aggressive by modern standards - quarter-amplitude decay implies ~25% overshoot and poor robustness margins.[fn::ZN was designed for self-regulating processes in the 1940s, when the objective was fast disturbance rejection on pressure and flow loops, not servo tracking. Applying it to temperature loops (where overshoot may damage product) is a category error. Tyreus–Luyben proposed more conservative settings: Kp = 0.45·Ku (same as ZN-PI), Ti = 2.4·Tu, Td = 0.4·Tu - roughly half the ZN aggressiveness.]
The open-loop (process reaction curve) variant: take the plant offline, apply a step, fit the response to a FOPDT model =G(s) = K·e^{-Ls}/(τs+1)=, then:
See Root Locus: Evans Construction for why these gains land where they do.
Cohen–Coon
An improvement on ZN open-loop that accounts for the dead-time-to-time-constant ratio L/τ. Cohen–Coon coefficients are tabulated but reduce to rational functions of L/τ; for small L/τ they converge to ZN, and for large L/τ (dead-time-dominant plants) they give more conservative gains. It is marginally better than ZN on paper and rarely worth the added complexity in practice.
Internal Model Control (IMC)
IMC reframes tuning as a design choice: pick a desired closed-loop time constant τc, then back-calculate the controller parameters from the plant model. For a FOPDT plant:
The single tuning parameter τc trades speed against robustness: small τc → fast but fragile; large τc → slow but robust. A common heuristic is =τc = max(τ/3, L)=. IMC is the modern default in process control because it makes the trade-off explicit rather than baking it into an opaque table. See Stability: Routh-Hurwitz for the stability analysis that constrains τc from below.
Tuning Trade-offs
No set of PID gains achieves fast response, high robustness, and zero overshoot simultaneously. This is not a failure of tuning method - it is a structural consequence of the waterbed effect (Bode integral): reducing sensitivity in one frequency band increases it in another.[fn::The Bode sensitivity integral states that ∫₀^∞ ln|S(jω)| dω = π·Σ Re(p_k), where the sum is over right-half-plane poles. For stable plants the integral is zero, so any frequency band where |S| < 1 (good disturbance rejection) is compensated by a band where |S| > 1 (amplified disturbances). Faster loops push the "bad" band to higher frequencies where it matters less, but never eliminate it.] Increasing Kp reduces rise time but erodes phase margin. Increasing Ki eliminates offset faster but adds oscillation. Increasing Kd damps overshoot but amplifies noise and can destabilise if the plant has high-frequency phase lag.
The practical takeaway: pick two of {fast, robust, no overshoot} and accept degradation on the third. Process control typically chooses robust + no overshoot (slow). Motion control typically chooses fast + no overshoot (fragile, requires careful anti-windup).
Anti-Windup
Integral windup is the #1 practical PID failure mode. When the actuator saturates (valve fully open, heater at max current), the loop is effectively open: the plant output stops responding to further increases in u. But the integrator keeps accumulating error, winding up to a large value. When the setpoint is reached or crossed, the wound-up integral must unwind - producing a massive, prolonged overshoot that can trip safety interlocks.
The mechanism: saturation breaks the feedback loop, but the integrator doesn't know. The fix is to make the integrator aware of saturation:
- Clamping (conditional integration): stop integrating when the controller output is saturated and the error is driving further into saturation. Simple, effective, the industrial default.
- Back-calculation: compute the difference between the requested output and the actual (saturated) output, feed it back to the integrator with a gain
Kt(tracking time constant). More graceful than clamping but requires tuningKt. - Tracking anti-windup: reset the integrator state to a value that makes the controller output sit exactly at the saturation limit. Most robust but most complex.
Any deployed PID must have anti-windup. A bare integrator in a loop with saturation is a bug, not a tuning problem. See Motors & Actuators for the actuator-saturation context.
Derivative Filter
The D term is Kd·s in the Laplace domain - a differentiator. In the frequency domain it has magnitude Kd·ω, rising at 20 dB/decade. Real measurement noise is broadband; the D term amplifies it without bound. A raw derivative term will make the control output jitter violently, wearing actuators and injecting noise into the plant.
The standard fix is to cascade the derivative with a first-order low-pass filter:
where Tf is the filter time constant. Equivalently, one specifies the filter ratio =N = Kd/(Kp·Tf)= (typically 5–20). This caps the high-frequency gain of the D term, trading derivative effectiveness against noise immunity. The filter is not optional. Most commercial PID controllers (Honeywell, Emerson, Siemens) include it by default and expose only N as a parameter.
A second, related issue: derivative kick. If the derivative acts on the error e(t), a step change in setpoint r(t) produces an impulse in de/dt, and thus a spike in the control output. The fix is process-variable derivative: compute the derivative of y(t) (the measurement), not e(t). Since y is continuous (the plant filters its own input), the kick disappears. This is standard in process control and should be the default everywhere.
Setpoint Weighting
The standard PID acts on the error =e = r − y=. But the P and D terms can be split between setpoint and measurement:
with weighting factors β, γ ∈ [0, 1]. Setting =β = γ = 1= gives the textbook error-based PID. Setting =β = γ = 0= gives a pure feedback controller with no setpoint feedforward - zero overshoot on setpoint steps but slower response. The integral term always uses the full error (weight 1) because it needs to see the true offset to eliminate it.
Typical choice: =β = 0.5–1.0=, =γ = 0= (derivative on measurement only, to avoid kick). Setpoint weighting decouples the servo (setpoint tracking) and regulator (disturbance rejection) behaviours, which a single-gain PID otherwise conflates.
Worked Example: Temperature Loop (PI, Ziegler–Nichols vs IMC)
Consider a heated water bath with FOPDT model (identified by step test):
Process gain =K = 0.8= (°C per % heater power), time constant =τ = 60 s=, dead time =L = 10 s=. The dead-time ratio =L/τ = 0.167= is moderate - PID should work well.
Ziegler–Nichols (open-loop) PI tuning:
This gives a closed loop with roughly 25–40% overshoot and a settling time of ~3–4·(τ + L) ≈ 210–280 s. The gain margin is thin (ZN targets quarter-decay, not robustness). On a water bath this is tolerable; on a batch reactor where overshoot spoils the product, it is not.
IMC PI tuning (τc = L = 10 s):
The IMC gain is about half the ZN gain. The closed loop will be slower (settling ~5–6·τ ≈ 300–360 s) but with <5% overshoot and far better robustness. Disturbance rejection is also slower - this is the explicit trade-off.
Critique:
ZN's aggressive gain (6.75) buys ~30% faster response than IMC (3.75) but at the cost of overshoot and fragility. If the bath's heat-loss coefficient drifts (e.g., ambient temperature changes), the ZN-tuned loop may go unstable; the IMC-tuned loop will merely slow down. For a =L/τ = 0.167= plant, either works; for L/τ > 0.5 (dead-time-dominant), ZN becomes dangerous and IMC (or a Smith predictor) is mandatory.
The deeper point: neither tuning is "correct" - they encode different assumptions about what matters (speed vs. robustness). The IMC formulation makes this explicit (τc is a design parameter); ZN hides it inside a table that was calibrated for 1940s flow loops. See Frequency Response: Bode & Nyquist.
When It Breaks
1. Integral windup (see above): the #1 failure mode. Any loop with actuator saturation and an unfiltered integrator will wind up. Deploy anti-windup before tuning.
2. Derivative kick on setpoint steps: if the D term acts on error, a setpoint step produces an output spike. Use process-variable derivative (=γ = 0=).
3. Noise sensitivity of D: the raw Kd·s term amplifies measurement noise without bound. Always filter (first-order, =N = 5–20=). On noisy loops (flow, pressure), disable D entirely - PI is usually sufficient.
4. Oscillatory plants: PID cannot control a plant with a significant resonant mode (e.g., a flexible shaft, a long slender column, a spring-mass system). The D term's phase lead is too coarse to cancel the resonance's phase drop; the loop will ring or go unstable. The fix is structural: add a notch filter at the resonance frequency, or use a model-based controller (LQR, H∞) that can place a zero near the resonance. PID's three parameters simply cannot represent the required dynamics. See Root Locus: Evans Construction.
5. Long dead time: when L/τ > 1, the phase lag from the dead time exhausts the phase lead available from the D term. A Smith predictor (which uses a model to cancel the dead time) or a model-predictive controller is needed. PID alone will be either sluggish (conservative gains) or unstable (aggressive gains).
6. Process-variable vs error derivative: always use PV derivative. This is not a tuning choice but an implementation correctness issue - error-based derivative is a latent bug that manifests as output spikes on every setpoint change.
7. Non-minimum-phase plants: if the plant has a right-half-plane zero (e.g., a boiler's "swelling" effect on drum level), the initial response is inverse - the output moves the wrong way before correcting. PID cannot handle this gracefully; the controller must be detuned until the inverse response is negligible, or a model-based feedforward path must cancel it.
Meta-Observation
PID's dominance is a consequence of good enough, not optimal. For the ~90% of industrial loops that are first-order-ish, mildly nonlinear, and tolerant of ±5% setpoint tracking, PID with hand-tuned or auto-tuned gains is the cheapest solution that works. The cost of building a plant model - step tests, identification, validation - exceeds the performance gain from a model-based controller, and the model degrades as the plant ages anyway.
The remaining 10% is where the theory earns its keep. Oscillatory plants need notch filters or LQR. Dead-time-dominant plants need Smith predictors or MPC. Multivariable plants need decoupling or state-space methods. Non-minimum-phase plants need careful feedforward. In each case, the failure of PID is structural - three parameters cannot represent the required controller dynamics - not a matter of tuning.
The scraped blog content this note replaced was symptomatic of a broader pattern: PID is widely taught as a recipe (ZN table → plug in numbers → ship), with the Laplace-domain justification either omitted or relegated to a footnote. This inverts the actual logic. PID is a constrained approximation of a general controller - a truncated Laurent series with three terms - and its tuning tables are empirical shortcuts for a design problem that IMC makes explicit. Understanding the constraint (why three terms, why these three) is what separates a controls engineer from a technician with a tuning app.
Related
- Optimal Control - LQR reduces to PID for low-order diagonal-weight plants
- Digital Control - the z-transform PID variant
- BMS Control - PID applied to battery charge limiting
- Control Theory: Signals & Laplace
- Transfer Functions & Block Diagrams
- Stability: Routh-Hurwitz
- Frequency Response: Bode & Nyquist
- Root Locus: Evans Construction
- Motors & Actuators