SOC drift: diagnosis and mitigation
The engineering problem
Persistent SOC error is not one failure mode with one remedy. A signed offset can accumulate from current measurement, survive because voltage provides weak SOC information, or be reinforced by model and estimator assumptions. A low nominal RMSE does not identify the cause, and changing Kalman covariances can hide a symptom without removing its source.
This page organizes the repository around four root-cause families:
- Measurement: current offset or gain error, voltage bias, coulomb-counting accumulation, quantization, communication, or data faults.
- Battery/model: OCV–SOC error, hysteresis, resistance/RC mismatch, temperature, capacity error, aging, or insufficient model order.
- Estimator: initial-state dependence, weak observability, inappropriate states, or numerical and structural assumptions.
- Tuning: poor \(Q/R\) balance, estimator overconfidence, a narrow useful covariance basin, or overfitting to a nominal objective.
When SOC drifts, do not start by tuning \(Q\) and \(R\). Test measurement, initialization, OCV/model residuals, temperature/history, and observability first. Covariance tuning changes how the filter reacts to uncertainty; it does not necessarily remove the physical source of persistent error.
A practical diagnostic sequence
| Step | Question | Diagnostic evidence | Engineering response |
|---|---|---|---|
| 1 | Is error drifting or merely initialized incorrectly? | Signed SOC error versus time; initial, mean, and final error | Estimate sign/rate and repeat from several initial SOC values. |
| 2 | Does error follow current measurement? | Zero-current offset, gain calibration, integrated Ah difference | Correct calibration or test a current-bias state. |
| 3 | Does voltage contain enough SOC information? | \(dOCV/dSOC\), current excitation, time spent on the LFP plateau | Avoid demanding correction that the available output cannot support. |
| 4 | Is residual structure model-related? | Innovation mean/autocorrelation versus SOC, current, temperature, direction, and rest/history | Revisit OCV, hysteresis, dynamics, capacity, temperature, or model order. |
| 5 | Is the filter fragile? | Covariance and initial-SOC sweeps | Select a broad robust region, not only the minimum. |
| 6 | Does the result transfer? | Additive-noise, gain/bias, and Hall-bias cases | Use robustness evidence in the estimator decision. |
| 7 | Is a persistent disturbance representable and observable? | Innovation plus direct disturbance-state recovery | Add a state or adaptive parameter only when it can be distinguished from other errors. |
What the current ATL20 P25 studies demonstrate
The canonical study uses models/ATL20model_P25.mat and the desktop_atl20_bss_v1 ESC-generated dataset. The promoted artifacts answer different diagnostic questions:
- The nominal benchmark measures the final tuned operating point.
- The initial-SOC sweep exposes convergence sensitivity over 0–100% initialization.
- The covariance sweep distinguishes a best point from a broad useful basin.
- The tuned-covariance and shared-covariance injection studies separate additive noise from persistent gain/offset faults.
The results do not collapse to one winner. For example, EacrSPKF is the nominal P25 leader at 0.497% SOC RMSE, while EaEKF is exceptionally flat in the covariance sweep (0.639% best, 0.640% mean, 0.647% worst). In the tuned initial-SOC sweep, EaEKF again has the strongest average robustness, whereas ESC-EKF reaches the best single initialization point but degrades much more at the extremes. These are different engineering properties.
Persistent bias is different from zero-mean noise
The P25 injection configuration includes three distinct disturbances:
- additive voltage/current noise;
- a combined current-gain/current-offset and voltage-gain/offset fault; and
- a
hall_biascase with constant current bias and clean voltage.
In the tuned-covariance study, the gain/bias case has 0.302 A current RMSE and the Hall-bias case has 0.430 A current RMSE. Selected SOC results show that the estimator ranking changes substantially:
| Estimator | Nominal SOC RMSE (%) | Gain/bias SOC RMSE (%) | Hall-bias SOC RMSE (%) | Hall-bias SOC mean error (%) |
|---|---|---|---|---|
EacrSPKF |
0.497 | 11.565 | 22.579 | -21.252 |
EbSPKF |
8.441 | 6.129 | 4.956 | -2.801 |
EsSPKF |
8.261 | 6.916 | 6.958 | -5.175 |
Em7SPKF |
7.593 | 7.555 | 10.282 | -8.860 |
The sustained signed errors in the Hall case are consistent with a persistent disturbance, and EbSPKF ranks first in both P25 persistent-bias cases. This is useful robustness evidence; ranking alone does not prove that its internal bias state recovered the injected current error. The disturbance also changes the voltage/SOC trajectory seen by each estimator, and the filters differ in states, dynamics, and tuning.
Estimating sensor bias as a state
For a current sensor with offset,
\[ I_{\mathrm{meas}} = I_{\mathrm{true}} + b_I + \nu_I, \]
an estimator can use
\[ I_{\mathrm{eff}} = I_{\mathrm{meas}} - \hat b_I, \qquad b_I(k+1) = b_I(k) + w_b. \]
Here \(b_I\) is an additional state. Its process noise controls how quickly the estimated offset may change. Battery-voltage and dynamic residuals provide the information used to update it. This idea is established in the battery state estimation literature: Zhao, Duncan, and Howey augment the battery model with sensor biases, derive nonlinear observability conditions, and compare nonlinear Kalman filters on cell data (Zhao et al. 2017). Malysz and coauthors use a KF current- bias estimator and show substantial recovery after injecting a 1 A bias on experimental drive-cycle data, while noting operating-condition dependence (Malysz et al. 2016).
Adding the state is not enough. SOC error, model error, resistance error, and sensor bias can produce similar voltage residuals. The augmented system must remain sufficiently observable.
What this repository implements
The implementation families are related but not equivalent:
| Estimator | Bias mechanism | Verified implementation status |
|---|---|---|
EbSPKF |
One current-bias state inside the main SPKF state vector | Active: prediction and output subtract the bias estimate, and the state follows a process-noise-driven random walk. |
EBiSPKF |
Separate two-stage bias filter using Bb, Cb, and V |
Hooks exist, but the default benchmark helper leaves these matrices at zero; under that wiring the bias estimate remains at its initial value. |
Em7SPKF |
The same external bias branch plus an R0 estimator |
The R0 branch is active, but the default bias branch has the same zero-matrix limitation as EBiSPKF. |
In iterEbSPKF.m, the state equation computes currentEff = current - xold(ibInd,:) + ..., uses that corrected current for RC, hysteresis, and SOC propagation, and evolves the bias as xnew(ibInd,:) = xold(ibInd,:) + .... EbSPKF therefore does not merely tune a conventional filter: it attempts to infer persistent current-sensor offset internally and removes the estimate before SOC propagation. The benchmark runner also records the bias estimate and its 3-sigma bound for all three named bias estimators.
.png)
Historical ATL diagnostic, not a current P25 recovery result. EbSPKF produces a changing estimate while the default EBiSPKF and Em7SPKF traces remain at zero. The plot does not overlay the true injected bias, so it is implementation evidence—not validation of bias accuracy.
See the estimator-design notes for the full assumptions and failure modes.
Why observability is central for LFP
Adding a bias state does not guarantee that the filter can distinguish sensor bias from SOC error or model mismatch. Identifiability depends on the battery model, OCV–SOC sensitivity, dynamic excitation, available measurements, covariance assumptions, and how many disturbances are estimated at once.
For LFP, broad flat OCV regions provide weak voltage information about SOC. During low excitation or long plateau residence, SOC error, OCV/model error, and accumulated current bias can therefore be difficult to separate. This is why the repository’s OCV decision, ESC model validation, injection tests, and estimator experiments belong in one causal story rather than independent leaderboards.
Direct bias-state validation — highest-priority experiment
The current benchmark already logs estimator bias traces and 3-sigma bounds, and the composite-injection dataset stores injected_current_bias_a. However, the promoted P25 summaries report resulting SOC/voltage metrics—not direct bias-recovery metrics—and the required multi-case recovery study has not been run. No recovery claim is made here.
The next reproducible study should compare true injected current bias against the estimated bias and its ±3-sigma bound, reporting:
- bias-estimation RMSE, mean error, and final error in amperes;
- convergence time and uncertainty-bound coverage;
- SOC RMSE with and without compensation; and
- behavior for constant positive and negative bias, random-walk bias, bias plus measurement noise, and low-excitation/LFP-plateau operation.
This requires a small promoted-summary/experiment layer around the existing logged outputs and several new cases. Implementing only a plot from the current constant-bias run would not satisfy that validation matrix, so it remains the highest-priority evidence gap rather than a fabricated result.
Model-side evidence and its boundary
The Middle OCV study addresses one model-side risk: finite-rate polarization and hysteresis should not be silently embedded in the static OCV backbone. The ESC validation summary also reports nonzero error on independent modelling datasets, which is relevant to potential model-driven drift.
The canonical desktop P25 dataset, however, is ESC-generated and reports zero application-dataset model error when replayed through the same model family. That controlled setup is excellent for isolating estimator architecture, covariance, initialization, and injected measurement faults. It cannot quantify real-cell plant/model mismatch.
The repository currently provides strong controlled evidence for estimator-, tuning-, initialization-, and measurement-bias-induced SOC error. Separating those effects from real battery-model mismatch requires evaluation against independent measured cell or BESS data. Direct recovery of the injected bias state also remains to be promoted and scored.
Engineering decision
The current ATL20 P25 selection combines nominal, covariance, initialization, and injection rankings. It selects Em7SPKF under the stated aggregate weighting, keeps EsSPKF as the balanced fallback, identifies EbSPKF as the bias-transfer specialist, and retains EacrSPKF as the nominal-performance reference. This is a portfolio of evidence, not proof that one filter removes every source of drift.