Diagnosing and mitigating SOC drift
Why does this repository exist?
Reliable state of charge is difficult to maintain over long battery operation. Small errors in current measurement, battery-model parameters, OCV representation, initialization, and estimator assumptions can produce persistent SOC bias or drift. This repository is an experimental framework for diagnosing and mitigating SOC drift: it uses controlled studies to ask which mechanism creates an error and which mitigation actually addresses it.

Representative ATL desktop trajectory agreement. This is a historical ATL diagnostic figure; current numerical claims use the promoted ATL20 P25 result set.

SOC-estimation error over the same benchmark. Persistent signed error and progressive separation are more informative for drift diagnosis than nominal trajectory agreement alone. This figure shares the historical provenance of the trajectory plot above.
Do not start by tuning \(Q\) and \(R\). First determine whether the error follows measured-current bias, initialization, OCV/model residuals, temperature/history, or poor observability. Tune the filter after those causes have been checked. If a persistent physical disturbance can be represented and observed—such as current-sensor offset—estimate it explicitly as a state rather than forcing covariance tuning to absorb it.
SOC drift: diagnosis and mitigation
| Step | Question | Evidence | Likely action |
|---|---|---|---|
| 1 | Is the error actually drifting? | SOC error versus time; mean and final error | Characterize sign and rate. |
| 2 | Is measurement bias present? | Calibration, zero-current offset, integrated Ah error | Calibrate or estimate the bias. |
| 3 | Is initialization responsible? | Initial-SOC sweep | Improve initialization or estimator convergence. |
| 4 | Is OCV/model structure wrong? | Residual versus SOC, direction, current, temperature, and history | Improve OCV/hysteresis or re-identify dynamics. |
| 5 | Is SOC locally observable? | \(dOCV/dSOC\), excitation, plateau residence | Reduce reliance on voltage correction or change architecture. |
| 6 | Is tuning fragile? | \(Q/R\) sweep and Bayes history | Prefer a robust region over one minimum. |
| 7 | Does performance transfer? | Noise, gain/bias, Hall-bias, and initial-SOC studies | Use robustness-aware selection. |
| 8 | Does structured bias remain? | Persistent innovation and SOC bias | Add observable states, adaptive parameters, or a better model. |
The detailed SOC drift technical study connects these tests to the current ATL20 P25 evidence, the bias-aware estimator code, and the highest-priority missing experiment.
Four root-cause families
| Family | Examples | What the repository tests |
|---|---|---|
| Measurement | Current offset/gain, voltage bias, accumulated coulomb-counting error, data faults | Additive-noise, gain/bias, and Hall/current-bias injections |
| Battery/model | OCV–SOC error, hysteresis, resistance/RC mismatch, temperature, capacity, aging, model order | OCV inspection and ESC identification/validation |
| Estimator | Initial-state dependence, weak observability, inappropriate states, numerical or structural assumptions | Estimator comparisons, initial-SOC sweeps, bias-state filters |
| Tuning | Poor \(Q/R\) balance, overconfidence, narrow covariance basin, nominal-objective overfitting | Grid covariance sweeps and Bayesian optimization |
Covariance tuning changes how an estimator reacts to uncertainty; it does not necessarily remove the physical source of a persistent error.
One connected evidence chain
- Problem: the drift study and signed-error figures show why nominal overlap is insufficient.
- Model and observability: the Middle OCV decision and ESC validation evidence constrain what voltage can tell the estimator.
- Estimator and initialization: the initial-SOC sweep separates convergence behavior from nominal accuracy.
- Tuning: the covariance sweep and Bayesian tuning study distinguish robust regions from single best points.
- Deployment robustness: the P25 injection study tests zero-mean noise separately from persistent sensor errors.
- Decision: the current estimator selection combines nominal accuracy and robustness instead of selecting the lowest nominal RMSE alone.
The canonical ATL20 P25 desktop dataset is ESC-generated. It is strong controlled evidence for estimator-, tuning-, initialization-, and measurement-bias-induced SOC error. Because the synthetic plant and estimator can share the same model family, it does not by itself quantify real-cell plant/model mismatch. Separating those effects requires evaluation against independent measured cell or BESS data.
For implementation details, use the canonical ATL20 workflow and architecture policy.