Data-Driven Fault Diagnosis and Early Warning for EV Battery Packs

1. Introduction: Background and Research Motivation

As global population growth and economic development continue apace, the dual pressures of energy depletion and environmental degradation have become increasingly prominent. In response, governments around the world have accelerated the promotion of new energy vehicles as a cleaner transportation alternative. Over the past decade, electric vehicle sales have exhibited a robust upward trajectory, increasing from approximately 18,000 units in 2013 to over 3.8 million units in 2023. This remarkable growth has established the **EV battery pack** as an integral subsystem within contemporary automotive engineering.

Lithium-ion batteries have become the primary energy storage solution for electric vehicles, prized for their high energy density, low self-discharge rates, and exceptional cycling performance. However, the very factors that render Li-ion chemistry attractive also introduce considerable safety risks. Unlike conventional automotive components, lithium-ion cells operate within a narrow electrochemical window with respect to voltage and temperature. When these boundaries are exceeded, either through overcharging, over-discharging, or thermal stress, the resulting failures can have catastrophic consequences, including thermal runaway, fire, and explosion. Statistical records from the past decade indicate a pronounced upsurge in electric vehicle fire incidents, rising from just 5 cases in 2014 to approximately 464 cases in 2023, with a consistent upward trend in recent years. Such incidents not only pose severe threats to personnel safety and property but also undermine public confidence in the large-scale deployment of EVs.

A comprehensive understanding of the failure mechanisms of the **EV battery pack** is fundamental to the development of effective diagnostic strategies. The failure modes associated with lithium-ion batteries are often classified into two main categories: battery-internal faults and system-level faults. Battery-internal faults typically manifest as gradual degradation phenomena, such as capacity fade and internal resistance growth, or as sudden catastrophic events, including internal short circuits, thermal runaway, and electrolyte leakage. In contrast, system-level faults arise from issues with battery management systems, sensors, and interconnecting hardware components.

The failure of the **EV battery pack** traces its root to several complex triggers. Overcharge situations promote lithium dendrite formation and metal deposition at the anode, which can penetrate the separator and induce internal short circuits. Conversely, over-discharge leads to dissolution of transition metals from the cathode and irreversible structural damage. Both phenomena compromise the long-term stability of the power module and accelerate failure risk. Operation at excessive C-rates increases heat generation and contributes to uneven lithium-ion distribution and mechanical strain on electrode structures. Extreme temperatures have equally detrimental effects: high environmental heat accelerates parasitic chemical reactions that consume active material and produce gaseous byproducts; low temperatures render electrolytes viscous, hamper ion transport, and increase susceptibility to lithium plating during charging sessions. Additionally, cell-to-cell inconsistency remains an intrinsic challenge, originating from manufacturing tolerances and amplified through cycling and environmental disparities, ultimately driving voltage divergence and unequal state-of-charge distributions.

Traditional fault diagnosis methodologies for the **EV battery pack** generally fall into three categories: model-based, data-driven, and statistical analysis methods. Model-based approaches leverage equivalent circuit models or electrochemical models to estimate states and generate residuals for diagnosis, yet these approaches often exhibit poor generality across driving cycles and are challenged by their inherent difficulty in accurately modeling the complex and nonlinear dynamics of a real-world EV battery pack. Data-driven approaches use machine learning and deep learning algorithms to perform anomaly detection tasks on operational datasets, but they frequently suffer from significant computational overhead and intensive data requirement constraints. Statistical methods assess fault probability or hypothesis-test diagnostic outcomes but can struggle with accurate identification of fault characteristics and classification of types. Consequently, existing fault diagnosis methodologies continue to confront serious challenges when translating to actual vehicular deployments, including inefficient detection, weak robustness, insufficient sensitivity to subtle cell parameter shifts, and limitations in predicting imminent anomalies.

Given the regulatory mandate requiring EV installation with onboard data acquisition terminals that report to centralized monitoring platforms, a wealth of operational data is now routinely collected. However, these platforms currently serve a narrow monitoring function that is incapable of advanced failure anticipation. The present work aims to bridge this substantial gap by developing an integrated and comprehensive fault diagnosis and early warning methodology for the **EV battery pack** using a synergistic combination of data mining, statistics-based analytics, and machine learning. My primary research objectives can be summarized as implementing effective anomaly detection and localization strategies, diagnosing cell faults through probabilistic statistical methods, identifying specific fault types, and constructing a robust voltage-prediction model to achieve sufficiently early fault warning and comprehensive safety assessment of the **EV battery pack**.

2. Operational Data Source and Preprocessing Strategies

2.1 Working Principle of the Lithium-Ion Battery

The operating mechanism of a lithium-ion cell revolves around the intercalation and deintercalation of lithium ions between the anode and cathode materials. During the charging process, Li⁺ ions migrate from the positive electrode (typically lithium metal oxides) through the electrolyte to the negative electrode (commonly graphite), while the reverse movement occurs during discharging. The overall electrochemical process within a lithium cobalt oxide/graphite cell during operation can be expressed by the following reaction equations:

For the charging process, the positive electrode and negative electrode reactions are described as:

$$ \text{LiCoO}_2 \xrightarrow{\text{charge}} \text{Li}_{1-x}\text{CoO}_2 + x\text{Li}^+ + x e^- \tag{2.1} $$

$$ x\text{Li}^+ + x e^- + \text{C}_x \xrightarrow{\text{charge}} \text{Li}_x\text{C}_x \tag{2.2} $$

The reversible nature of these reactions is the underlying principle of battery operation. However, detrimental side reactions can occur under abusive conditions. During overcharge, cobalt dissolution and oxygen evolution lead to structural collapse. In cases of severe over-discharge, the copper current collector may dissolve, creating internal short-circuit pathways. These fundamentals emphasize the critical need for precise management of the **EV battery pack**.

2.2 Data Source and Battery Parameter Dynamic Analysis

The research dataset has been collected from a cloud monitoring platform established by a vehicle manufacturer, which covers a period of three full years of operational history. The platform collects over 50 distinct signals from vehicles in real time at a frequency of 1 Hz. These signals encompass alarms, battery measurements, vehicle locations, entire vehicle data, and electric motor parameters, including diagnostic trouble codes, total battery voltages and currents, individual cell voltages, internal temperatures, and SOC. To gain insight into the dynamic characteristics of vehicle operation, I examined a segment collected over a six-month cycle from one vehicle in the fleet. The data exhibit varied battery parameters: battery pack voltages range primarily from 332.8 V to 403.8 V, the SOC values are mainly between 29% and 100%, the highest temperature intervals fluctuate between 11°C and 37.5°C, and the speed range extends from 0 to 78.8 km/h. Individual cell voltages show a range from 3.502 V to 4.253 V. The analysis confirms that the EV battery pack encounters multiple operation modes—charging, discharging, static or dormant periods, and offline events—all contributing to highly dynamic and frequently random parameter fluctuations.

Through a series of observational inspections of various EV battery packs, I identified several typical anomalies within an **EV battery pack** by correlating vehicle status data with alarm messages. For example, at one specific timestamp, a temperature difference fault was reported when one temperature sensor demonstrated a temperature spread of about 6°C relative to other cells, implying cell inconsistency. A second event pinpointed a discharge overcurrent fault at a sudden discharge current drop to -43.7 A, which exceeds the battery pack’s respective safety upper limit of 25 A while driving under normal vehicle speed conditions. A third observation highlighted the occurrence of a sensor fault, where, despite all other parameters remaining within their normal ranges, a temperature value spiked and then reverted the following second. These analyses show the rich and variable nature of data available and underscore the significance of accurate fault diagnostics.

2.3 Data Preprocessing Methods

The reliability of subsequent analysis in fault diagnosis of the **EV battery pack** is heavily reliant on the quality and integrity of the operational data. Therefore, I first conducted comprehensive data preprocessing procedures. Because cloud platform data often include duplicate timestamps, I calculated the total percentage of repetitive time samples within the dataset. For instance, from an analysis sample, four particular vehicles had repetition rates of only 0.45%, 0.49%, 0.34%, and 0.48%, respectively, out of their total records. The method used to reduce redundancy in such scenarios was to retain the first record of each duplicated timestamp.

Concerning missing data, I evaluated whether the missing interval was longer than 1 minute. if the period of missing measurements exceeded one minute, information from that segment was deleted to eliminate potentially invalid interpolation. For missing data within short time spans, interpolation using cubic spline interpolation was implemented to preserve data dynamics without producing unrealistic oscillations. The interpolation efficacy is demonstrated by checking whether reconstructed data retained the essential patterns before and after missing samples.

Outlier detection procedures were based on the statistical distribution of parameters. Data collected from actual **EV battery pack** operations frequently show near-normal distributions across large sample counts if the dataset is sufficiently vast. In this context, the 3-sigma rule is used to detect outliers. For a variable \(X\), the mean \(\mu\) and standard deviation \(\sigma\) are calculated. Approximately \(99.7\%\) of instances fall within 3σ of the mean under normal assumptions; therefore, values extending beyond this threshold are classified as outliers. I applied this technique to parameters like pack voltage, electrical current, and SOC.

Not every relevant real-world measurement follows a normal distribution. Because factors like vehicle speed and driving mileage are often impacted by external traffic flow, multiple outliers, and traffic signals, these data showed skewness and non-normal behavior. For such parameters, the box-plot detection approach is more suitable. In the box-plot rule, values are considered outliers if they fall below the lower fence or above the upper fence, where the interquartile range is denoted by \(IQR\). The upper fence and lower fence are calculated as \(Q3 + 1.5\times IQR\) and \(Q1 – 1.5\times IQR\), respectively.

The significance of preprocessing in this study, specifically for fault detection of the EV battery pack, cannot be overemphasized. It cleans data, makes exploration possible, eliminates misleading artifacts, and guarantees appropriate data quality.

3. Data Mining Approaches for Fault Diagnosis and Anomaly Detection of the EV Battery Pack

3.1 Dimensionality Reduction and Visualization of Cell Voltage Data

To explore and analyze massive time series data that represent high-dimensional voltages from every cell of a vehicular battery, it was essential to apply dimension reduction technique. Dimensionality reduction simplifies datasets and enables potential data anomaly detection in pattern analysis. I examined several alternatives including PCA, autoencoder analysis, and t-SNE. PCA is computationally efficient but works best for linear structures. Autoencoders can capture nonlinear structures but require extensive data and hyperparameter tuning. For this dataset, t-SNE has an excellent ability to visualize complicated, nonlinear, and high-dimensional data. The t-SNE algorithm converts similarities between data points into joint probabilities and minimizes the Kullback-Leibler divergence between the high-dimensional and low-dimensional distributions. Mathematically, the similarity between two points \(x_i\) and \(x_j\) in high-dimensional space is given as:

$$ P_{j\mid i} = \frac{\exp\left(-\|x_i – x_j\|^2 / 2\sigma_i^2\right)}{\sum_{k \neq i} \exp\left(-\|x_i – x_k\|^2 / 2\sigma_i^2\right)} \tag{3.1} $$

where \(\sigma_i\) is the variance of the Gaussian kernel centered at \(x_i\). In low-dimensional space, the joint probability is represented using the Student t-distribution:

$$ Q_{ij} = \frac{(1 + \|y_i – y_j\|^2)^{-1}}{\sum_{k \neq l} (1 + \|y_k – y_l\|^2)^{-1}} \tag{3.2} $$

After visualizing the 30 individual cell voltage streams from a battery pack of an electrified passenger car using t-SNE, the transformed points formed multiple distinct data clusters and several isolated points. The visual clusters resemble normal variations of a series-connected EV battery pack. However, there were also points which deviated remarkably, suggesting potential abnormalities. I concluded that this behavior requires additional analysis, for example, by inference-based algorithms such as clustering.

3.2 Cell Anomaly Detection Using K-Means Clustering

After dimensionality reduction, I used K-means clustering algorithm to partition the t-SNE transformed voltage points into clusters for anomaly detection. K-means is an unsupervised classification procedure that assigns each point to the cluster with the nearest centroid, based on Euclidean distance. The objective is to minimize intra-cluster variance:

$$ J = \sum_{k=1}^{K} \sum_{x \in C_k} \|x – \mu_k\|^2 \tag{3.3} $$

where \(K\) represents number of clusters, \(C_k\) denotes the collection of assigned data points to the k-th cluster, and \(\mu_k\) represents the centroid of the cluster.

After clustering with the optimal number of clusters and a distance-sensitive criterion, K=3 typically revealed one or multiple dense clusters representing standard operation data and one isolated, sparse cluster representing abnormal voltage data. The anomaly cluster was sufficiently separated from the normal data, leading to the observation that certain cells occasionally behave differently under specific operational times.

After defining the clusters, each point’s distance from its cluster center was measured, and the threshold was based on mean plus a coefficient multiplied by standard deviation. This criterion is set as:

$$ P = \mu + k \cdot \sigma \tag{3.4} $$

where, based on data analysis, the coefficient value \(k = 1.5\). Anything further from the centroid than this threshold was taken to be abnormal.

3.3 Anomaly Localization Using Z-Score Statistical Approach

After flagging anomalous data points, my task then became identifying the specific batteries responsible. Since an **EV battery pack**s tends to evolve in gaussian-like distribution under big baseline conditions, we introduced a gaussian density function for each cell:

$$ f(x) = \frac{1}{\sqrt{2\pi}\sigma} \exp\left(-\frac{(x-\mu)^2}{2\sigma^2}\right) \tag{3.5} $$

where \(x\) is the target sample, \(\mu\) denotes average, and \(\sigma\) is deviation. I used the Z-score formula to compare each voltage measurement to the pack average:

$$ Z_i = \frac{V_i – \mu}{\sigma} \tag{3.6} $$

For improved detection of specific faulty cells, I derived the probability-based Z diagnostic coefficient:

$$ Z(t) = \sum_{i=1}^{k} \frac{P_i(t) – \frac{1}{k}\sum_{j=1}^{k} P_j(t)}{\sigma_p(t)} \tag{3.7} $$

where, \(P_i(t)\) is probability density at time t for cell i, k is the total number of cells (for this work k=30) and \(\sigma_p(t)\) is the standard deviation across all probability density values.

Investigating a specific fault on a particular date using voltage data from the pack, I identified cell number 7 as anomalous. The voltage profile for this cell displayed significantly lower voltage values compared to all other cells, and its voltage drop was fast, with maximum inter-cell difference reaching 236 mV and minimum cell voltage falling to 3067 mV. The Z-based anomaly index for cell 7 remained markedly higher and had wild fluctuations, providing conclusive evidence for the cell-level fault isolation.

3.4 Comprehensive Battery Performance Assessment Using Integrated Entropy Weight and Coefficient of Variation Approach

After detecting and isolating a specific cell, it becomes important to systematically evaluate the degree to which each cell in an **EV battery pack** has diverged. I used a combined weighting scheme merging entropy weight with coefficient of variation to assess individual cell behavior. This strategy accounts for the non-linear, multi-criteria nature of the problem.

The method starts from a voltage matrix \(\mathbf{U}\) of \(n\) cells at \(t\) time instants:

$$ \mathbf{U} = \begin{bmatrix} u_{11} & u_{12} & \cdots & u_{1t} \\ u_{21} & u_{22} & \cdots & u_{2t} \\ \vdots & \vdots & \ddots & \vdots \\ u_{n1} & u_{n2} & \cdots & u_{nt} \end{bmatrix} \tag{3.8} $$

For entropy weight computation, the share of each cell toward total at a given time instant is:

$$ p_{ij} = \frac{u_{ij}}{\sum_{i=1}^{n} u_{ij}} \tag{3.9} $$

The entropy of each time indicator \(j\) is computed via:

$$ e_j = -\frac{1}{\ln n} \sum_{i=1}^{n} p_{ij} \ln p_{ij} \tag{3.10} $$

The entropy-based weight is then derived as:

$$ W_j^{(1)} = \frac{1 – e_j}{\sum_{m=1}^{t} (1 – e_m)} \tag{3.11} $$

Regarding the coefficient of variation method, I first calculated the mean voltage among cells:

$$ A_j = \frac{1}{n}\sum_{i=1}^{n} u_{ij} \tag{3.12} $$

Standard deviation:

$$ S_j = \sqrt{\frac{1}{n}\sum_{i=1}^{n} (u_{ij} – A_j)^2} \tag{3.13} $$

Coefficient of variation:

$$ V_j = \frac{S_j}{A_j} \tag{3.14} $$

Then, the variation weight is:

$$ W_j^{(2)} = \frac{V_j}{\sum_{m=1}^{t} V_m} \tag{3.15} $$

The combined weights:

$$ W_j^c = \alpha W_j^{(1)} + (1-\alpha) W_j^{(2)} \tag{3.16} $$

The results from tuning suggested \(\alpha = 0.4\) for this work. Finally, each cell gets a composite score:

$$ S_i = \sum_{j=1}^{n} W_j^c u_{ij},\quad \bar{S} = \frac{\sum_{i=1}^{n} S_i}{n} \tag{3.17} $$

and the score difference with average is:

$$ \Delta S = S_i – \bar{S} \tag{3.18} $$

I computed upper and lower scores by noting the limits from cell voltage extremes. This gives view of potential spread. For 7th cell, its composite rating was the lowest among all cells, with \(\Delta S = 0.727\). The assessment also included weight values across each time metric, where the lowest computed weight was 0.0024 at the 28th second, confirming that the fault happened in a state of strong irregularity. I then compared cell 7 performance across three distinct operating periods: June 22, 2020, July 9, 2020, and July 14, 2020. The score difference of cell 7 shifts from 0.727 to 0.831 to 0.876 over these intervals. This increasing differential reflects performance degradation over the period, demonstrating that my integrated evaluation method captures trends in cell-level degradation within EV battery pack.

3.5 Fault Diagnosis Using the 3σ-MSS Multi-Level Screening Strategy

Although I could locate cells with divergent cell voltages, a deeper understanding of whether faults are systematic and steady or sudden and intermittent could only be accomplished with a probabilistic statistical analysis. Consequently, I made use of the 3σ-MSS criterion (multilevel screening strategy), whose iterative process aims to purify the center/cluster by eliminating multiple layers of outliers. Traditional ways of center-point estimation (arithmetic mean) are overly sensitive to outliers—so wrong conclusions may be drawn. This insight is core to fault classification. For each measured data set, the MSS procedure iteratively:

1. Calculates initial mean \(\mu_t^{(0)}\) and standard deviation \(\sigma_t^{(0)}\) of all the values present in the array of cell terminal voltages.
2. Evaluates the differences: \(\Delta V_i^{(0)} = V_i^{(0)} – \mu_t^{(0)}\).
3. Generates a new data matrix where those values that are out of \(3\sigma\) at stage \(m\) are left out, to compute the updated parameters \(\mu_t^{(m)}\), \(\sigma_t^{(m)}\).
4. Repeats steps \(m\) times so that center point is increasingly the actual center of the normal distribution. The stop criterion is given by:

$$ |\mu^{(m)} – \mu^{(m-1)}| < \mathcal{F} \tag{3.19} $$

where \(\mathcal{F}\) is a convergence threshold fixed at 0.001.

Then, the final fault threshold boundary for my application is:

$$ \mu^{(m)} \pm \beta \cdot \sigma^{(m)} \tag{3.20} $$

With \(\beta\) as an empirical coefficient chosen based on data quality. Values outside range are assigned 1 in fault matrix \(\mathbf{R}_{T_0,T_1}\); in-range values are treated as 0. Fault probability for any individual cell is calculated from fault matrix:

$$ P_{\text{fault}} = \frac{N_{\text{fault}}}{N_{\text{total}}} \times 100\% \tag{3.21} $$

Applying the process across a month’s data for multiple vehicles, the outcome yielded relatively limited fault probabilities in most cells, with variations helping classify the faults. I interpreted various patterns:

– For certain vehicles, cells exhibit short intervals with rapid transient voltage fluctuations at random positions, and fault probability for a subset goes above 15% at times while others remain below 2.5%. Those are considered sudden faults, usually triggered by unexpected events or external abuse.
– In another class of behavior, the fault probabilities are low (\(<2\%\)), stable, and appear primarily on same particular cells, i.e., those cells are consistently deviating over time. This was characterized instead as systematic fault, associated with deterioration, initial imbalance, design defect or fixed weaknesses.

This differentiation based on probability behavior helps make the practice of anomaly diagnosis for the **EV battery pack** far more valuable.

3.6 Comparative Analysis of Fault Diagnosis Algorithms

To verify the diagnostic efficacy of the presented 3σ-MSS algorithm, I benchmarked its results against two common algorithms used for outliers in the same scenario: LOF and COF. While LOF defines outliers through the concept of local density relative to its neighbors, and COF detects outliers based on connectivity distance relationships to cluster structures, my strategy estimates outlier values with statistical margins. I tested each on exactly the same input data set. Fault probability diagnosis results for three methods were in qualitative alignment. The 3σ-MSS method yielded a maximum fault probability among all cell elements of \(2.13\%\) and a minimum fault probability of \(1.53\%\). The COF results returned maximum probability of \(1.91\%\), but min value was only \(0.4\%\) and underdiagnoses low-level anomalies. The LOF method estimated \(2.47\%\) as maximum and \(1.43\%\) as min fault probability, and shows larger deviations under higher-probability conditions. The key outcome is that 3σ-MSS exhibited most stable and accurate output under both high and low fault rates, validating the rationale of using it for robust fault probability statistics.

3.7 Temporal and Seasonal Fault Frequency Variation of the EV Battery Pack

Then, I set out to examine the time–season variation of the fault probability from the **EV battery pack** dataset using the diagnosis method described. I grouped operational data that was taken from a fleet of ten pure electric passenger cars into four seasons (spring, summer, autumn, winter). The recorded fault probabilities for these seasons are as follows.

Table 1 presents the values of maximum and average fault probabilities in each season:

| Season | Max Cell Fault Probability (%) | Average Fault Probability (%) |
|——–|——————————-|——————————-|
| Spring | 1.99 | 1.54 |
| Summer | 4.95 | 4.31 |
| Autumn | 3.67 | 3.07 |
| Winter | 9.52 | 4.59 |

The results clearly demonstrate that the fault probabilities of the **EV battery pack** tend to be greatest for summer and winter conditions. High environmental temperatures during summer accelerate unwanted electrochemical deterioration and the battery’s cycle capacity; low temperature in winter can impede electrolyte transport and raise the chance of overdischarge. Also, there were isolated cell outliers in winter: cells 10 and 28 diverged by \(4.93\%\) and \(4.58\%\) from average, pointing toward random, sudden fault events.

In investigating three years of data from one selected EV, the first year showed minimal faults, and fault probabilities were mainly below 2%. The second year, however, a cell reaches \(12.69\%\) fault frequency during a single day, and in the third year, another cell peaks at \(14.14\%\), indicating gradual or unexpected age-related degradation in this particular EV battery pack.

These time-series-based analyses demonstrate practical value of data mining in the investigation of performance characteristics of an **EV battery pack** over long-term use.

4. Voltage Prediction and Fault Early Warning Using an IGWO-CNN-LSTM Model

4.1 Framework of Predictive Fault Warning

A single diagnostic model cannot achieve early prediction that prevents failures. Therefore, as a final important component, I used machine learning to anticipate future voltages of the EV battery pack to forecast fault occurrence early.

For monitoring an **EV battery pack**, I developed a predictive voltage model using a combined Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) model, with hyperparameters optimized using Improved Grey Wolf Optimizer. I then used prediction results to set thresholds and to send out warning levels and fault categories.

4.2 Correlation Analysis and Feature Selection

Before model training, I used both Pearson and Spearman correlation methods to decide the worth of candidate features for predicting terminal cell voltage. The Pearson coefficient is measured as:

$$ \rho = \frac{\sum_{t=1}^{n}(x_t-\bar{x})(y_t-\bar{y})}{\sqrt{\sum_{t=1}^{n}(x_t-\bar{x})^2 \cdot \sum_{t=1}^{n}(y_t-\bar{y})^2}} \tag{4.1} $$

The Spearman rank coefficient is given by:

$$ R = 1 – \frac{6\sum (r_{x_t}-r_{y_t})^2}{n(n^2-1)} \tag{4.2} $$

Based on the computed correlation strengths, the most relevant predictors of voltage are the total pack voltage, SOC, and the historical single-cell voltage. The correlation absolute values calculated by Pearson and Spearman methods both indicated \(|\rho|\) greater than 0.9 and \(|R|\) greater than 0.9 for those parameters. These features were then used as inputs into predictive framework. A table of correlation coefficients derived among voltage and other candidate variables is presented as:

| Parameter | Pearson Coefficient \( \rho \) | Spearman Coefficient \( R \) |
|———–|——————————-|——————————|
| Pack Voltage | 0.973 | 0.952 |
| SOC | 0.949 | 0.936 |
| Temperature | 0.438 | 0.407 |
| Current | 0.365 | 0.351 |
| Vehicle Speed | 0.126 | 0.114 |
| Cumulative Mileage | 0.218 | 0.229 |

From these values it is obvious that pack total voltage and SOC have the strongest link with cell voltage evolution, while others retain little meaningful relevance, thus I decided to use them in the model.

4.3 IGWO-CNN-LSTM Hybrid Neural Network

4.3.1 CNN for Feature Extraction

For spatiotemporal modeling, first I applied 1-D CNN to automatically extract the local features and patterns in the voltage series. A convolutional layer creates feature maps by convolving weight filters across the input:

$$ z^{(l)} = (W^{(l)} * a^{(l-1)}) + b^{(l)} \tag{4.3} $$

This is followed by batch normalization and ReLU activation, and max pooling is used for downsampling and selecting the most salient features. The CNN output is meant for subsequent sequential processing.

4.3.2 LSTM for Sequential Modeling

In LSTM, the memory cell and its gates are responsible for learning long-term dependencies. The three gates are calculated as:

Forget gate:

$$ f_t = \sigma(W_{xf} X_t + W_{hf} h_{t-1} + b_f) \tag{4.4} $$

Input gate:

$$ i_t = \sigma(W_{xi} X_t + W_{hi} h_{t-1} + b_i) \tag{4.5} $$

Output gate:

$$ o_t = \sigma(W_{xo} X_t + W_{ho} h_{t-1} + b_o) \tag{4.6} $$

Candidate memory:

$$ \tilde{C}_t = \tanh(W_{xc} X_t + W_{hc} h_{t-1} + b_c) \tag{4.7} $$

Memory update:

$$ C_t = f_t \odot C_{t-1} + i_t \odot \tilde{C}_t \tag{4.8} $$

Cell output:

$$ h_t = o_t \odot \tanh(C_t) \tag{4.9} $$

Where \(\sigma(\cdot)\) is sigmoid, \(W\) are weights, and \(\odot\) is element-wise multiplication. LSTM handles memory long enough to capture long term dependence in EV battery pack voltage dynamics, whereas the preceding CNN layer finds local characteristics. In my final structure, the CNN block applies convolution and pooling, then flattens results to the LSTM layer, followed by a fully connected layer to make final voltage predictions. The hyperparameters to tune include kernel count, mini-batch size, number of LSTM hidden units, dropout rate, learning rate, and optimizer selection.

Table shows the proposed layer structure for hybrid model:

| Layer Type | Configuration / Purpose |
|————|————————–|
| Input Layer | Sequential inputs of pack voltage, SOC, historical cell voltage |
| Convolution 2D | Convolution window [2,1], maps features |
| Batch Normalization | Normalize to accelerate convergence |
| ReLU Layer | Nonlinear activation |
| MaxPooling 2D | Downsampling to extract significant features |
| Flatten Layer | Prepare CNN output for LSTM layer |
| LSTM Layer | Model temporal long-term dependency |
| ReLU Layer | Nonlinear activation for LSTM output |
| Fully Connected | Output regression prediction layer |

4.3.3 Improved Grey Wolf Optimization

The standard GWO is based on hierarchical hunting behavior of grey wolves. In GWO, alpha (\(\alpha\)), beta (\(\beta\)), delta (\(\delta\)), and omega (\(\omega\)) represent four levels of hierarchy. Prey encircling can be written mathematically as:

$$ \vec{D} = |\vec{C} \cdot \vec{X}_p(t) – \vec{X}(t)|,\quad \vec{X}(t+1) = \vec{X}_p(t) – \vec{A}\cdot \vec{D} \tag{4.10} $$

The controlling parameters are \(\vec{A} = 2a \vec{r}_1 – a\) and \(\vec{C} = 2\vec{r}_2\), with \(a\) decreased linearly. Positions are updated by alpha, beta, delta wolves, as described in earlier sections.

In the standard formation, the coefficient \(a\) is reduced linearly from 2 to 0. The GWO approach can also suffer from premature convergence and stagnation in local optima. I therefore considered an improved grey wolf optimizer (IGWO) suited to optimise this deep learning architecture. The significant enhancements introduced are as follows:

1. Cauchy mutation in the search process: I added Cauchy distribution-based random mutations to chosen candidate wolves. Cauchy’s heavy-tail property allows large step sizes and enables wolves that move away from local minima in search space and improves global exploration capacities.

2. Control parameter tuned through a nonlinear function that adjusts exploration-exploitation balance over iterations:

$$ a(t) = 2 – 2\left(\frac{t}{T_{\max}}\right)^{1/k} \tag{4.11} $$

Here, \(k\) controls the decay rate and rises convergence rate.

3. Variable inertia weight factor, \(w\), decreasing dynamically in each iteration to accelerate convergence in late rounds and boost solution refinement.

Thus, IGWO offers better convergence property and more reliable global search. It was used to search the hyperparameter space of CNN-LSTM architecture. Each wolf represents combination of hyperparameters, and the fitness is evaluated by RMSE metric from validation predictions. IGWO iteratively guides towards parameter combinations with favorable performance. Once best set of hyperparameters is established, we train the hybrid model with those param values and test it offline.

4.4 Model Evaluation Metrics

Evaluation was made based on root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (\(R^2\)). The definitions are:

$$ \text{RMSE} = \sqrt{\frac{1}{n}\sum_{i=1}^n (u_i – \hat{u}_i)^2} \tag{4.12} $$

$$ \text{MAE} = \frac{1}{n}\sum_{i=1}^n |u_i – \hat{u}_i| \tag{4.13} $$

$$ R^2 = 1 – \frac{\sum_{i=1}^n (u_i – \hat{u}_i)^2}{\sum_{i=1}^n (u_i – \bar{u})^2} \tag{4.14} $$

Where \(u_i\) is observed value, \(\hat{u}_i\) is model output, and \(\bar{u}\) is mean target value.

4.5 Comparison of Different Deep Learning Models

I trained multiple models: BiLSTM, LSTM, CNN-LSTM, CNN-GRU, and IGWO-CNN-LSTM, all with similar settings and datasets. The prediction performance of the five architectures on the same dataset is presented below:

| Forecasting Model | MAE (mV) | RMSE (mV) | \(R^2\) |
|——————-|———-|———–|———|
| BiLSTM | 7.3997 | 9.6738 | 0.99647 |
| LSTM | 9.1646 | 11.594 | 0.99493 |
| CNN-LSTM | 6.2004 | 8.4264 | 0.99732 |
| CNN-GRU | 4.4459 | 5.4142 | 0.99889 |
| IGWO-CNN-LSTM | 3.5254 | 4.5086 | 0.99923 |

From Table, IGWO-CNN-LSTM achieves lowest values of MAE and RMSE (3.5254 mV and 4.5086 mV) and highest \(R^2\) of 0.99923. In comparison, basic LSTM has relatively weak forecasting accuracy and the hybrid networks work better. The application of IGWO has helped tune model so it achieves extremely high precision of cell voltage.

In terms of peak error values, the IGWO-CNN-LSTM prediction error peaks at 24 mV while standard CNN-LSTM error max height was 30 mV using a specific daily data test.

4.6 Fault Diagnosis Strategy and Warning Policy

For realistic monitoring for an EV battery pack, I built a diagnosis strategy using predicted voltages and alarm thresholds. The types of fault identified were over-voltage, under-voltage, rapid voltage rise or fall, and poor state-of-charge consistency.

Table below describes classification rules and alert level:

| Fault Type | Classification Rule | Alarm Level |
|————————–|————————-|————-|
| Cell Overvoltage | \(4.3\,\text{V} \le U \le 4.8\,\text{V}\) | Level 3 |
| Severe Cell Overvoltage | \(U > 4.8\,\text{V}\) | Level 2 |
| Cell Undervoltage | \(U \le 3.4\,\text{V}\) | Level 1 |
| Fast voltage rise | \( \Delta U \ge 0.4\,\text{V/s}\) | Level 1 |
| Fast voltage fall | \( \Delta U \le -0.4\,\text{V/s}\) | Level 1 |
| Poor voltage consistency | \( \partial \ge 0.5\,\text{V}\) | Level 1 |

A Level 1 warning is most severe, implying cell is at a dangerous condition, and a driver needs to stop cell operation and retreat. Level 2 means threshold is reached but still controllable if action is taken soon. Level 3 is minor, yet it suggests a potential condition that should be handled timely to avoid battery deterioration.

To verify the model’s capability for early prediction, I picked a dataset that contains an actual overvoltage event on one **EV battery pack**. From the model prediction, the cell voltage gradually rose above the upper threshold 4.3 V around sample point 4480–4500. Model predicted voltage profile leads the physical measurements by around one minute. The proposed method is verified to perform as a robust fault forecast tool.

4.7 Warning and Operational Maintenance Recommendation

Based on the severity of alarms, I recommend immediate response actions consistent with safety. For Level 1 warnings (danger), vehicle owner should stop quickly in safe area and avoid contact with high-voltage modules. In case of Level 2, one should not continue driving until inspection and determining root cause is complete. In case of Level 3 condition, to keep an **EV battery pack** functional for some extended period, immediate care is to reduce power demand, stop fast charging, check cell balance, schedule maintenance action and perform diagnostic service at earliest convenience to avoid progression of failure.

5 Conclusions and Future Prospects

This dissertation has systematically explored the anomaly detection, fault localization, fault diagnosis, and early warning for lithium-ion **EV battery packs**. It used cloud monitoring platform data from production pure electric passenger cars to create data-driven algorithms and produced the following findings.

First, using data from real vehicles, I explored how the parameters fluctuate over driving modes and identified different operational states. Data pre-processing including duplicate removal, missing value interpolation, and outlier screening using sigma-rule and box-plot methods succeeded in “purifying” raw vehicle data for next stages.

Second, in the analysis of the **EV battery pack**, t-SNE dimension reduction worked well for cell data and K-means clustering enabled visual separation of anomalous data. A fault detection approach based on Z-score analysis allowed exact anomaly detection and cell fault isolation. By combining entropy weights and coefficients of variation, complete battery performance assessment has become possible. Time evolution of cell scoring, such as the increase in the score difference of cell 7 from 0.727 to 0.876, made the validity evident.

Third, the 3σ-MSS diagnosis model characterized battery faults by statistically defining two groups: sudden and systematic. In a comparative analysis with LOF and COF methods, this approach demonstrates the most stable and accurate statistical diagnosis performance for battery packs. The diagnostic practices over different seasons and years prove its ability to estimate diverse fault probabilities and recognize relations with operational environment and battery age.

Fourth, in the voltage forecasting context, I used correlation analysis to define high-quality inputs. A hybrid CNN-LSTM block with improved hyperparameter update through IGWO was designed to capture local and temporal features from data. Models predictability on unseen testing data confirmed high precision: RMSE = 4.5086 mV, MAE = 3.5254 mV, and \(R^2 = 0.99923\). This was a significant improvement over other network models. Threshold-based early warning can predict a battery outlier event around 1 min before actual occurrence and can be used at cell-level in the **EV battery pack**.

Future work could include detailed failure-cause information to better differentiate fault categories. Also, multi-fault coupled scenarios are an important area yet to be fully investigated. To maximize safety, we could enhance prediction algorithms with neural architecture search and novel attention mechanisms, examine transfer abilities across vehicle models, and deploy online integrated diagnostics on edge hardware for active real-time protection.

In closing, the methods developed in this thesis prove that big data mining and deep learning for fault diagnosis and early warnings for **EV battery pack** not only are possible but when implemented with rigorous preprocessing, have the potential to significantly improve operational safety for electric vehicles.

Scroll to Top