Diagnosing and mitigating SOC drift

An experimental framework for separating measurement, model, estimator, initialization, and tuning causes of persistent battery SOC error.

Why does this repository exist?

Reliable state of charge is difficult to maintain over long battery operation. Small errors in current measurement, battery-model parameters, OCV representation, initialization, and estimator assumptions can produce persistent SOC bias or drift. This repository is an experimental framework for diagnosing and mitigating SOC drift: it uses controlled studies to ask which mechanism creates an error and which mitigation actually addresses it.

Reference and estimated ATL battery SOC and terminal voltage over a desktop benchmark profile

Representative ATL desktop trajectory agreement. This is a historical ATL diagnostic figure; current numerical claims use the promoted ATL20 P25 result set.

Signed SOC estimation errors over the same ATL desktop benchmark profile

SOC-estimation error over the same benchmark. Persistent signed error and progressive separation are more informative for drift diagnosis than nominal trajectory agreement alone. This figure shares the historical provenance of the trajectory plot above.

ImportantWhat to do when SOC drifts

Do not start by tuning \(Q\) and \(R\). First determine whether the error follows measured-current bias, initialization, OCV/model residuals, temperature/history, or poor observability. Tune the filter after those causes have been checked. If a persistent physical disturbance can be represented and observed—such as current-sensor offset—estimate it explicitly as a state rather than forcing covariance tuning to absorb it.

SOC drift: diagnosis and mitigation

Step Question Evidence Likely action
1 Is the error actually drifting? SOC error versus time; mean and final error Characterize sign and rate.
2 Is measurement bias present? Calibration, zero-current offset, integrated Ah error Calibrate or estimate the bias.
3 Is initialization responsible? Initial-SOC sweep Improve initialization or estimator convergence.
4 Is OCV/model structure wrong? Residual versus SOC, direction, current, temperature, and history Improve OCV/hysteresis or re-identify dynamics.
5 Is SOC locally observable? \(dOCV/dSOC\), excitation, plateau residence Reduce reliance on voltage correction or change architecture.
6 Is tuning fragile? \(Q/R\) sweep and Bayes history Prefer a robust region over one minimum.
7 Does performance transfer? Noise, gain/bias, Hall-bias, and initial-SOC studies Use robustness-aware selection.
8 Does structured bias remain? Persistent innovation and SOC bias Add observable states, adaptive parameters, or a better model.

The detailed SOC drift technical study connects these tests to the current ATL20 P25 evidence, the bias-aware estimator code, and the highest-priority missing experiment.

Four root-cause families

Family Examples What the repository tests
Measurement Current offset/gain, voltage bias, accumulated coulomb-counting error, data faults Additive-noise, gain/bias, and Hall/current-bias injections
Battery/model OCV–SOC error, hysteresis, resistance/RC mismatch, temperature, capacity, aging, model order OCV inspection and ESC identification/validation
Estimator Initial-state dependence, weak observability, inappropriate states, numerical or structural assumptions Estimator comparisons, initial-SOC sweeps, bias-state filters
Tuning Poor \(Q/R\) balance, overconfidence, narrow covariance basin, nominal-objective overfitting Grid covariance sweeps and Bayesian optimization

Covariance tuning changes how an estimator reacts to uncertainty; it does not necessarily remove the physical source of a persistent error.

One connected evidence chain

WarningControlled evidence, not yet a real-cell mismatch study

The canonical ATL20 P25 desktop dataset is ESC-generated. It is strong controlled evidence for estimator-, tuning-, initialization-, and measurement-bias-induced SOC error. Because the synthetic plant and estimator can share the same model family, it does not by itself quantify real-cell plant/model mismatch. Separating those effects requires evaluation against independent measured cell or BESS data.

For implementation details, use the canonical ATL20 workflow and architecture policy.