Point Cloud Deep Learning for EV Battery Sealing Nail Defect Detection

The rapid expansion of the new energy vehicle industry has placed unprecedented demands on the safety, reliability, and production efficiency of automotive batteries, commonly referred to as EV batteries. As the core energy storage component of electric vehicles, the EV battery directly dictates the overall vehicle quality, driving range, and operational lifespan. The manufacturing process of an EV battery involves numerous critical steps, one of which is the welding of the sealing nail. This small metallic component is used to seal the electrolyte injection hole after the cell is filled. The integrity of this seal is paramount; any microscopic defect in the weld can lead to electrolyte leakage, short circuits, and catastrophic battery failure, posing severe risks to both property and personal safety.

However, the sealing nail is characterized by its diminutive size, typically just a few millimeters in diameter. During the laser welding process, various factors such as contamination of the welding nozzle, fluctuations in laser power, or uncleaned residues on the welding surface can lead to tiny surface defects. These defects, including burst points, pinholes, incomplete welds, and misalignments, exhibit complex morphological characteristics that are extremely difficult to detect using traditional methods. Conventional quality control often relies on manual visual inspection or basic 2D image analysis. Manual inspection is inherently inefficient, highly subjective, and susceptible to human error, leading to inconsistent quality control. While 2D machine vision offers a degree of automation, it fundamentally lacks the depth information required to accurately assess the geometric nature—such as volume, depth, and curvature—of these three-dimensional surface anomalies. This often results in low detection accuracy and a high rate of missed detections for small-scale defects.

To overcome these challenges, advanced three-dimensional (3D) point cloud analysis has emerged as a promising solution. By capturing the complete spatial geometry of the sealing nail, from point clouds generated by line laser scanners and structured light systems, a more comprehensive representation of the object’s surface is achieved. Nevertheless, applying conventional 3D point cloud processing algorithms presents its own set of hurdles, including low precision in complex scenarios and a tendency to overlook micro-scale irregularities that are characteristic of welding defects on EV battery components. Therefore, this research is dedicated to developing an intelligent, automated, and highly accurate defect detection method for EV battery sealing nails by harnessing the power of deep learning on 3D point cloud data. This study systematically investigates techniques ranging from point cloud preprocessing and segmentation to the design of sophisticated deep semantic segmentation networks, aiming to significantly improve the accuracy and reliability of automated inspections, ultimately contributing to the enhancement of EV battery production automation and quality assurance.

—

**1. Point Cloud Data Preprocessing and Welding Region Extraction**

Before conducting deep learning-based analysis, the raw point cloud data acquired from the EV battery sealing nails require meticulous preprocessing. The initial high-resolution data captured by the 3D scanner not only contains the sealing nail but also a significant amount of extraneous information from the battery top cover and substrate. Additionally, this raw data is often plagued by noise due to variations in surface reflectivity and environmental factors. Directly applying a deep learning model to such unrefined data would inevitably lead to suboptimal performance and excessive computational costs. My proposed preprocessing pipeline consists of two critical stages: statistical filtering for noise removal and an improved random sample consensus algorithm for precise extraction of the target welding region.

**1.1 Statistical Filtering for Point Cloud Denoising**

To mitigate the impact of sparse, outlying noise points—which arise from uneven surface roughness, material reflections, and minor vibrations of the scanning system—I employed a statistical filtering method. This method operates on the principle that points within a neighborhood typically conform to a Gaussian distribution of distances. For a given point \(p_i\) in the raw point cloud \(P = \{p_1, p_2, \ldots, p_n\}\), I first calculate its mean distance \(\tilde{d}_i\) to its \(k\) nearest neighbors \(p_{ij}\):
\[
\tilde{d}_i = \frac{1}{k} \sum_{j=1}^{k} \sqrt{(x_i – x_{ij})^2 + (y_i – y_{ij})^2 + (z_i – z_{ij})^2}
\]
Assuming the global mean distance \(\mu\) and standard deviation \(\sigma\) of the entire point cloud are computed, a point is deemed an outlier if its mean neighbor distance \(\tilde{d}_i\) exceeds a global threshold \(\varepsilon\). The threshold is defined as \(\varepsilon = \theta \sigma + \mu\), where \(\theta\) is a multiplier that controls the filtering strictness. Points such as \(p_i\), where \(\tilde{d}_i > \varepsilon\), are considered noise and are removed from the dataset. This process effectively smooths the point cloud, enhances signal-to-noise ratio, and ensures the subsequent segmentation stages are not influenced by data artifacts.

**1.2 Improved Random Sample Consensus (RANSAC) Algorithm**

Following the denoising, the next challenge is to isolate the sealing nail and its immediate welding zone from the battery top cover, as these non-target regions create significant interference for defect detection. Traditional segmentation algorithms, such as region growing or Euclidean clustering, often falter due to the uniform density and subtle geometric transitions between the nail and the surrounding cover. To accomplish this task effectively, I proposed an improved RANSAC algorithm. The core innovation of this algorithm lies in transforming the random sampling process of standard RANSAC into a localized and constrained search. The pseudo-code and steps are conceptually illustrated in my research, emphasizing two key improvements: a refined seed point selection strategy and a misclassification point correction mechanism.

The process begins with estimating the surface normals for all points in the denoised point cloud. For a specific point \(p_i\), I construct a covariance matrix \(C\) from its nearest neighbors \(X\), whose centroid is \(\bar{p}\):
\[
C = \frac{1}{n} \sum_{i=1}^{n} (p_i – \bar{p})(p_i – \bar{p})^T
\]
The normal vector at each point is then extracted from the eigenvector corresponding to the minimum eigenvalue of this symmetric positive semi-definite matrix. To identify ideal seed points for plane fitting—which primarily lie on the flat surfaces of the top cover and the sealing nail—I avoid purely random selection. Instead, I introduce a local sampling strategy. The algorithm randomly selects a point \(q_0\) from the target space and searches for its neighbors within a radius \(r\). It then calculates the average cosine similarity \(\gamma\) between the normal vectors \(\mathbf{n}_i\) of points in this neighborhood and a reference vertical vector \(\mathbf{n}_v\):
\[
\gamma = \frac{1}{n} \sum_{i=1}^{n} \mathbf{n}_v \cdot \mathbf{n}_i
\]
If the value of \(\gamma\) exceeds a pre-defined threshold (e.g., 0.9), the algorithm assumes this region exhibits a strong planar characteristic. From this curated set, three initial seed points are randomly selected and used to fit a plane model \(M\) with distance threshold \(\theta\). All points in the remaining cloud that fall within this threshold are considered inliers for that model. This local sampling approach significantly reduces the number of iterations required to find the best-fit plane, enhances computational efficiency, and minimizes the risk of generating spurious planes.

A frequent issue encountered during the sequential extraction of multiple planes is the misclassification of points on the boundary between the top cover and the sealing nail. To address this, I integrated a multi-step misclassification correction mechanism. After the initial plane segmentation, the algorithm calculates the perpendicular distance \(d_1\) of a point to the geometric center of a candidate patch, as well as its distance \(d_2\) from established plane models. Points are flagged for reassignment if they satisfy certain dual-distance criteria against multiple potential patches. For these flagged points, the algorithm facilitates a correction by consulting the spatial labels of their k-d tree neighbors. Specifically, the point is reclassified to the dominant label within its immediate neighborhood as detected by a radius search. Finally, to perfect the segmentation and eliminate over-segmentation artifacts, the algorithm employs a Euclidean clustering-based patch merging step. This step consolidates adjacent fragments arising from conservative threshold settings into coherent, complete surface patches, resulting in the isolation of the sealing nail patch and the battery cover patch.

To evaluate the effectiveness of my improved RANSAC algorithm for extracting the sealing nail region from EV battery top cover point clouds, extensive experiments were conducted on a dataset of 196 scan samples. The precision \(P\), recall \(R\), and F1 score (F1) were used as quantitative metrics, defined as:
\[
P = \frac{TP}{TP + FP}, \quad R = \frac{TP}{TP + FN}, \quad F1 = 2 \times \frac{P \times R}{P + R}
\]
Here, \(TP\) (true positive) indicates correctly segmented nail points, \(FP\) (false positive) represents cover points wrongly classified as nail, and \(FN\) (false negative) denotes nail points missed by the algorithm. The performance of my proposed method is benchmarked against three other conventional algorithms, as summarized in Table 1.

**Table 1: Performance Comparison of Different Segmentation Algorithms**
| Algorithm | Computation Time (s) | Precision (%) | Recall (%) | F1 Score (%) |
|———————–|———————-|—————|————|————–|
| Traditional RANSAC | 6.82 | 76.24 | 82.19 | 79.10 |
| Region Growing | 8.73 | 83.10 | 81.57 | 82.33 |
| Euclidean Clustering | 3.45 | 36.23 | 99.73 | 53.15 |
| **My Improved RANSAC**| **6.14** | **82.44** | **93.76** | **87.74** |

As depicted in Table 1, my proposed improved RANSAC algorithm achieves the highest F1 score of 87.74%, demonstrating its superior robustness and accuracy. It effectively balances precision and recall, mitigating the misclassification issues present in traditional approaches. While Euclidean clustering exhibits high recall, its precision is very low due to severe under-segmentation. Region growing achieves high precision but at a greater computational cost. The proposed algorithm excels by providing a an optimal trade-off, providing a trustworthy and reliable input for the downstream defect detection network.

—

**2. Deep Learning for Sealing Nail Defect Detection**

To address the persistent problem of false negatives, where existing models fail to identify defective regions in EV battery sealing nails, this chapter focuses on the development of an advanced deep learning network to enhance detection accuracy and robustness. I selected a Graph Attention Convolution Network (GACNet) as the foundational architecture for this task because of its explicit capability to handle unstructured point clouds and its unique dynamic attention mechanism. However, this base network demonstrates limitations when processing the intricacies of welding defects. To overcome these, I propose two key innovations: a Local Graph Attention Filter (LGAF) and a Spatial Attention Graph Pooling module (SAG Pooling).

**2.1 Overall Architecture of the Improved Network (LGANet)**

The overarching architecture of my improved network, LGANet, mirrors the graph pyramid structure of the original GACNet, which facilitates hierarchical feature learning from fine-grained local geometry to robust global context. The network operates by constructing a directed graph \(\mathcal{G} = (\mathcal{V}, \mathcal{E})\) from the point cloud data, where vertices \(V = \{1, …, n\}\) correspond to the input points and edges \(E\) are defined by the k-nearest neighbors of each point in the spatial domain. As illustrated in my research design, the feature extraction block at each scale of the graph pyramid was completely re-engineered. I replaced the original graph attention convolution with the more sophisticated LGAF module to fundamentally improve the way spatial and feature information are integrated in the encoding phase. Furthermore, I substituted the basic symmetric functions used for structure downsampling with my SAG Pooling module to ensure critical spatial characteristics are reinforced during the aggregation phase.

**2.2 Local Graph Attention Filter (LGAF)**

In the original GACNet, attention weights are generated by concatenating spatial displacement vectors \(\Delta p_{ij}\) with feature differences \(\Delta h_{ij}\). I argued that this direct concatenation in a common embedding space can lead to confusion between spatial and feature-domain semantics, degrading the model’s ability to precisely infer defect morphology. To resolve this, I introduced the LGAF module, which acts as a filter to explicitly model the relationship between spatial structure and feature similarity.

The attention weight calculation in LGAF is a two-fold process. First, I compute a feature-based correlation matrix \(C_{ij}\). To estimate the compatibility between a central point \(i\) and its neighbor \(j\), a subtraction mechanism is used in an embedded space rather than concatenation:
\[
C_{ij} = \phi_q(x_i) – \phi_k(x_i – x_j)
\]
In this formulation, \(x_i\) and \(x_j\) are the feature vectors of the central vertex and its neighbor, while \(\phi_q\) and \(\phi_k\) are learnable linear projections. This operation is exceptionally effective at evaluating the local feature gradient and emphasizing sharp transformations, which is imperative for detecting the tiny boundaries of pinholes or burst points. Concurrently, I model the geometric structural information using a normalized spatial distance matrix \(A’\). This matrix leverages the topological graph adjacency \(A_{ij}\) and the spatial distance between point pairs:
\[
\hat{A}_{ij} = A_{ij} \cdot e^{-\|p_i – p_j\|^2}, \quad A’ = D^{-1/2} \hat{A} D^{-1/2}
\]
Here, \(D\) represents the degree matrix used for normalization. In high-density point cloud regions, the distance constraint works to filter out semantically irrelevant distant neighbors, ensuring the convolution kernel adapts to the object’s local topology. Furthermore, a multi-layer perceptron generates a learnable spatial encoding \(\sigma_{ij} = M_p(p_i – p_j)\) to enrich the geometric context.

The final joint attention \(\alpha_{ij}\) is derived by first fusing these spatial and feature correlations through an MLP \(M_e\) and then applying a softmax function for normalization:
\[
e_{ij} = M_e(C_{ij} \cdot A’_{ij}), \quad \alpha_{ij} = \frac{\exp(e_{ij})}{\sum_{l \in \mathcal{N}(i)} \exp(e_{il})}
\]
Finally, the output feature \(y_i\) for each central point is obtained by aggregating the local features and neighbor features. Instead of directly aggregating raw neighbor features, I introduce a residual enhancement mapped by the value projection \(\phi_v\):
\[
y_i = \psi(x_i) + \sum_{j \in \mathcal{N}(i)} \alpha_{ij} \odot \left( \phi_v(x_i – x_j) + \sigma_{ij} \right)
\]
Where \(\psi\) is a nonlinear transformation and \(\odot\) denotes the Hadamard product. By combining the spatial vector difference with the feature difference and weighting them by the spatial-filtered alpha, the LGAF effectively emphasizes salient defect morphology, reducing the tendency of the network to smear boundaries between distinct surface areas.

**2.3 Spatial Attention Graph Pooling (SAG Pooling)**

In standard graph pyramid networks, downsampling is typically performed via farthest point sampling followed by symmetric operations like max-pooling on the aggregated feature sets. This coarse, generic pooling strategy often discards intricate spatial dependencies vital for accurate local classification. To overcome this issue and fully exploit the spatial information within the graph structure, I formulated the SAG Pooling module. Unlike standard pooling which treats neighboring features uniformly, SAG pooling uses spatial attention to compute a weighted sum of the neighbor features.

For the pooling operation, a similar attention mechanism is applied based on the spatial coordinates of the vertices \(p_i\) and its neighbors \(p_j\). Specifically, a learnable weight \(\hat{w}_{ij}\) is generated by an MLP \(M_w\), contrasting spatial coordinate projections:
\[
\hat{w}_{ij} = M_w \left( \phi_q(p_i) – \phi_k(p_j) \right)
\]
These weights are normalized across the neighbor set \(\mathcal{N}(i)\) using a softmax function to obtain \(r_{ij}\). The aggregated feature \(F_i\) for the newly sampled centroid at \(i\) is given by:
\[
r_{ij} = \frac{\exp(\hat{w}_{ij})}{\sum_{l \in \mathcal{N}(i)} \exp(\hat{w}_{il})}, \quad F_i = \sum_{j \in \mathcal{N}(i)} r_{ij} \odot \phi_v(F_{ij})
\]
Here, \(F_{ij}\) represents the feature vector of the previous graph layer at vertex \(j\). This method is superior to max-pooling as it does not suffer from information loss due to spatial translation invariance. Instead, it retains a high-fidelity representation of the surrounding geometry, ensuring that discriminative features are propagated up the pyramid.

**2.4 Optimization with Weighted Loss Function**

In the context of semantic segmentation for defect detection, the point cloud data is inherently imbalanced, as the vast majority of data points belong to the normal background class, while defect regions like pinholes constitute a minority. This poses a serious challenge, as a standard loss function will bias the model heavily towards classifying everything as background. To mitigate this, I designed a weighted cross-entropy loss function. The weight for each class \(k\) is carefully calibrated to counterbalance the frequency distribution. The weighting scheme is given by:
\[
\hat{w}_i = \left( \frac{n_{max}}{n_i} \right)^{1/2}
\]
where \(n_{max}\) is the number of data points in the largest class and \(n_i\) is the number of data points in class \(i\). These values are then normalized to force the sum of weights to one:
\[
w_i = \frac{\hat{w}_i}{\sum_{j=1}^{k} \hat{w}_j}
\]
Consequently, the loss \(L\) for a multi-class classification task becomes:
\[
L = -\sum_{i=1}^{k} w_i \cdot y_i \log \left( \frac{\exp(x_i)}{\sum_{l=1}^{k} \exp(x_l)} \right)
\]
Here, \(y_i\) is the true label for class \(i\), and \(x_i\) represents the network’s raw output score before normalization. By applying a larger weight to minority defect classes, the loss function forces the model to pay greater attention to these hard-to-classify points, discouraging the model from merely favoring the dominant background class. This optimization was instrumental in enhancing the recall rates across several categories of defects.

—

**3. Feature Fusion for Small-Scale Defect Detection**

While the LGANet architecture substantially improves general feature extraction, the segmentation accuracy for subtle, small-scale defects—such as pinholes, burst spots, and tiny molten beads—was still unsatisfactory. This challenge mostly stems from the information bottleneck created by high levels of downsampling in deep networks, where high-frequency, fine-grained geometric details are lost. Simply passing over high-frequency features via skip connections is insufficient because extracting dense point-level details from lower levels and correlating it with deeply extracted semantics in higher levels is non-trivial. To address this issue, I proposed an architectural extension called LGANet-CDF which incorporates a novel Cross-Attention-Based Cross-Density Feature Fusion module (CDF).

**3.1 Motivation and Overall Fusion Structure**

The LGANet-CDF adopts a feature pyramid architecture where features flow in three primary directions: horizontal, downward, and upward. The downward flow integrates high-resolution, high-density point details from shallow layers into the feature fusion node, while the upward flow introduces deep, sparse and abstract semantic features into the fusion node. This targeted connection allows the network to strengthen the spatial resolution and structural edge information present in deeper layers.

The core intuition behind the CDF module is to utilize shallow high-density data to restore detailed geometry and deep low-density data to provide contextual semantics for the current layer. I specifically designed the feature down-sampling and feature up-sampling modules to align the spatial resolution across different densities. The down-sampling module uses farthest point sampling to reduce a point count by 4x and performs feature mapping and neighborhood aggregation on the original feature map. The up-sampling module, conversely, increases the number of points to the same count as the adjacent shallower layer. Features at the target level \(l\) are projected to the density of level \(l+1\) via interpolation, and neighbor embeddings are retrieved using k-NN search and computed as a weighted sum.

**3.2 Cross-Attention-Based Cross-Density Feature Fusion (CDF)**

At the heart of this architecture lies the CDF module, which is responsible for adequately fusing features from different densities. Rather than concatenating features of varying semantics directly, which can deteriorate performance, I applied a cross-attention mechanism. In this module, the feature from the middle layer \(l\), \(A\), acts as the *Query*. The targeted features \(B\) from the deeper layer \(l+1\) and \(C\) from the shallower layer \(l-1\) are both converted into *Key* and *Value* pairs.

For the shallow, high-density feature \(C\), the network aims to preserve precise structural information to assist in edge segmentation. The output from this branch, \(y_{l-1, i}\), is calculated using vector dot products, which are better suited to capture local low-level similarity:
\[
y_{l-1, i} = \sum_{j \in \mathcal{N}_{l-1}(i)} \rho \left( \chi_q(x_l) \cdot \chi_k^{l-1}(x_{l-1}) \right) \cdot \chi_v^{l-1}(x_{l-1})
\]
Here, \(\chi_q\), \(\chi_k\), and \(\chi_v\) denote linear projection functions for query, key, and value, and \(\rho\) is the softmax normalization function.

For the deep, low-density feature \(B\), the goal is to extract high-level contextual semantics, which aids in identifying the defect class even when the spatial footprint is minimal. I found that using a subtraction operation within a multi-layer perceptron \(M\) followed by an element-wise multiplication is more suited to distill these rich abstract features. Thus, the contexture-rich output \(y_{l+1, i}\) is defined as:
\[
y_{l+1, i} = \sum_{j \in \mathcal{N}_{l+1}(i)} \rho \left( M \left( \chi_q(x_l) – \chi_k^{l+1}(x_{l+1}) \right) \right) \odot \chi_v^{l+1}(x_{l+1})
\]
After calculating both perspectives of enhanced features, I fuse the output by stacking the original feature with the enhanced cross-density features and subsequently applying an MLP to reduce the dimensions and generate a well-aggregated feature set \(F\). The design successfully enables the deep layers of the network to access the high-definition information from the earlier stages, which is empirically proven to significantly boost the recognition accuracy for tiny defects, while the deep context keeps the false-positive rate in check.

—

**4. Experiments, Datasets, and Results**

To rigorously validate the effectiveness and robustness of the proposed LGANet and LGANet-CDF models for EV battery defect detection, a series of comprehensive experiments were performed. This section details the experimental setup, baseline comparisons, and an extensive ablation study.

**4.1 Dataset Description and Augmentation**

I curated a specialized dataset by collecting point cloud samples from the sealing nail region of EV batteries using a high-precision line laser scanner. The original data was captured from 929 distinct physical samples. Given the typically limited amount of industrial data available to train deep models effectively, I implemented a suite of data augmentation techniques. These included random rotations around the Z-axis, random translation, random flipping, and random scaling of the point clouds to expand the diversity and size of the training set. Furthermore, during training, I introduced random coordinate jitter and feature map dropping to improve the model’s generalization and enhance its robustness to sensor noise.

After this process, the total dataset expanded to 2,787 samples. The ground-truth labels were assigned using the CloudCompare software, categorizing points into six classes: Background (normal area), Burst, Pit, Molten Beads, Warpage, and Pinhole. The data partition was strictly randomized, with 2,257 samples used for training, 251 for validation, and 279 for testing. Table 2 provides a comprehensive overview of the dataset distribution.

**Table 2: Distribution of Sample Counts Across Different Defect Classes**
| Class | Normal | Burst | Pit | Molten Beads | Warpage | Pinhole |
|————-|——–|——-|——|————–|———|———|
| Sample Count| 655 | 406 | 271 | 248 | 316 | 361 |

**4.2 Implementation Details**

The experiments were conducted on a workstation equipped with an NVIDIA RTX 4070 GPU and an Intel Core i5-13400F CPU, running on Ubuntu 22.04. The network was implemented using the PyTorch 2.2.1 framework with CUDA 12.1. The critical hyperparameters tuned during the training process are listed in Table 3.

**Table 3: Hyperparameters Configuration for Model Training**
| Training Parameter | Value |
|—————————–|——-|
| Epochs | 300 |
| Batch size | 8 |
| Points per sample | 16384 |
| Optimizer | AdamW |
| Initial Learning Rate | 0.01 |
| Learning Rate Decay Strategy| Multi-step decay |

**4.3 Evaluation Metrics**

To measure the semantic segmentation performance comprehensively, I used standard metrics: Overall Accuracy (OA), mean Class Accuracy (mAcc), and mean Intersection over Union (mIoU). These metrics are defined mathematically as:
\[
OA = \frac{\sum_{i=1}^k n_{ii}}{\sum_{i=1}^{k}\sum_{j=1}^{k} n_{ij}}, \quad
mAcc = \frac{1}{k}\sum_{i=1}^k \frac{n_{ii}}{\sum_{j=1}^k n_{ij}}, \quad
mIoU = \frac{1}{k}\sum_{i=1}^k \frac{n_{ii}}{\sum_{j=1}^k n_{ij} + \sum_{j=1}^k (n_{ji} – n_{ii})}
\]
where \(n_{ij}\) denotes the number of points belonging to class \(i\) but predicted as class \(j\). OA reflects the overall pixel-wise accuracy, while mAcc and mIoU provide a balanced viewpoint across all classes, crucial for imbalanced defect datasets.

**4.4 Comparison with State-of-the-Art Networks**

The experimental benchmark in Table 4 quantifies the performance of my improved network against several leading point cloud semantic segmentation architectures.

**Table 4: Quantitative Comparison on EV Battery Sealing Nail Defect Dataset**
| Model | OA (%) | mAcc (%) | mIoU (%) | Burst IoU (%) | Pit IoU (%) | Molten Beads IoU (%) | Warpage IoU (%) | Pinhole IoU (%) |
|—————|——–|———-|———-|—————|————-|———————–|—————–|—————–|
| PointNet | 82.66 | 72.46 | 45.74 | 28.33 | 30.01 | 21.53 | 62.64 | 34.95 |
| PointNet++ | 89.33 | 84.65 | 62.67 | 47.91 | 59.32 | 42.37 | 81.64 | 46.53 |
| PointTransformer v1| 98.77 | 90.96 | 72.83 | 68.54 | 62.83 | 58.37 | 89.67 | 58.37 |
| PTv3 | 99.03 | 94.96 | 75.13 | 77.32 | 67.98 | 60.44 | 89.17 | 56.21 |
| OA-CNN | 98.72 | 91.13 | 74.04 | 64.89 | 66.91 | 64.87 | 88.37 | 60.04 |
| **GACNet** | 96.68 | 87.96 | 70.85 | 56.88 | 67.52 | 54.64 | 87.93 | 59.37 |
| **LGANet (Ours)**| 99.47 | 92.68 | 79.23 | 78.17 | 72.17 | 71.33 | 91.61 | 64.95 |
| **LGANet-CDF (Ours)**| **99.13** | **96.20** | **81.25** | **80.06** | **71.91** | **75.99** | **89.89** | **70.03** |

The performance metrics in Table 4 highlight a consistent improvement in defect detection. My initial LGANet model, which only integrated the LGAF and SAG pooling modifications, already surpassed existing baselines, achieving a high mIoU of 79.23%. However, the greatest improvement is realized with the integration of the CDF module. The LGANet-CDF architecture demonstrates an outstanding increase in the average IoU for the three categories of small-scale defects previously identified as most challenging. The experimental data support the efficacy of each architectural module proposed.

**4.5 Ablation Study on Module Effectiveness**

To verify the distinct contribution of each module, a sequence of ablation experiments was conducted. The results are presented in Table 5, where check marks (✓) indicate that a particular module is incorporated into the baseline GACNet architecture.

**Table 5: Ablation Study of the Improved Network Modules**
| Model | LGAF | SAG Pooling | OA (%) | mAcc (%) | mIoU (%) |
|————————|——|————-|——–|———-|———-|
| GACNet | | | 96.68 | 87.96 | 70.85 |
| GACNet + LGAF | ✓ | | 97.38 | 90.14 | 77.17 |
| GACNet + SAG Pooling | | ✓ | 97.96 | 88.95 | 74.81 |
| **LGANet** | ✓ | ✓ | 99.47 | 92.68 | 79.23 |

The ablation study clearly demonstrates that each proposed component contributes positively to the final result. The inclusion of LGAF boosts the mIoU by over 6%, underscoring the importance of decoupling spatial structures from feature space semantics to reduce boundary confusion. Including SAG pooling increases robustness and provides a modest but consistent gain. Most importantly, the effective synergistic operation of both LGAF and SAG Pooling is confirmed by LGANet, which yields an mIoU improvement of 8.38% over the original GACNet and significantly suppresses the false negative rate.

In a separate ablation study, I analyzed the contribution of the upward and downward information flow within the context of the CDF module. The results show that the downward flow, which transports data from dense shallow layers, is particularly beneficial for detecting small surface features like pinholes and molten beads. The upward flow, carrying contextual information from deep layers, is highly useful for preventing local over-segmentation. The final model LGANet-CDF, integrating both flows, leverages the contributions from each and achieves a balanced improvement across all classes, confirming its superior capability.

—

**5. Development of a Visualization Software Platform**

To translate the research findings into a practical tool that can be seamlessly integrated into an industrial setting as a visual detection tool for EV battery parts, I designed a user-centric application powered by the PyQt framework. This software platform serves as an interface for the non-expert user to interact with the trained deep learning models, visualize the inspection results in real-time, and export crucial data logs for production optimization.

The developed software platform is visually presented in the interface. The core design of the platform revolves around an intuitive, interactive GUI comprising control button panels and result display windows. The interface operates through an orderly workflow:

1. **Model and Weight Selection**: Users can conveniently load pre-trained neural network weights (e.g., `.pth` files) through the designated dialog, enabling the software to adapt to different detection requirements.
2. **Data Loading and View**: The operator simply imports a target point cloud file, which the software processes and renders as a depth map for preliminary viewing, ensuring the system is inspecting the correct region of the EV battery.
3. **Welding Area Fitting**: To optimize processing speed and accuracy, the interface allows for manual calibration of the sealing nail’s inner and outer weld diameters based on the specific production batch. Upon clicking “Fit Weld,” the software calculates the geometric centroid of the nail and visually masks the precise welding region, preventing computational resources from being wasted on irrelevant background points from the EV battery top cover.
4. **Defect Detection and Visualization**: The “Detect” button activates the inference phase of the network. The trained deep learning network processes the preprocessed point cloud and labels the defect points. Detected anomalies are displayed clearly in distinct colors, overlaid on the grayscale depth map for immediate operator review.
5. **Data Saving and Reporting**: After inspection, the “Save Defect Data” function outputs a structured text file containing the defect types, counts, and specific dimensions. This seamless data recording and exportability provide a robust foundation for quality inspection logs and facilitate offline root-cause analysis on the EV battery production line.

The successful development of this software confirms the practical applicability of the proposed deep learning scheme. It bridges the gap between advanced algorithmic research and pragmatic, factory-floor quality control by delivering a fast, accessible, and reliable interface.

—

**6. Conclusion and Future Outlook**

This research conducted a thorough investigation into automated defect detection for EV battery sealing nails, presenting a comprehensive pipeline integrating classical segmentation and cutting-edge deep learning. In conclusion, within my study regarding point clouds, I put forward several important contributions that significantly enhanced the state of the art.

**(1) Enhanced Precision in Target Region Extraction:** My study addressed the critical need for isolating the sealing nail region in a complex production environment. The introduction of an improved RANSAC algorithm, equipped with a local sampling strategy and a misclassification correction mechanism, effectively segmented the welding region, achieving a high F1 score of 87.74%. This meticulous point cloud preprocessing establishes a solid basis for all downstream defect detection stages.

**(2) Significant Improvement in Detection Accuracy with LGANet:** To resolve the issue of missing defect areas in current methodologies, I systematically improved the GACNet architecture. By designing the Local Graph Attention Filter (LGAF) module and the Spatial Attention Graph Pooling (SAG Pooling) module, I successfully enabled the network to focus on morphological differences inherent to defects. The ablation studies confirm that these enhancements contributed to an 8.38% improvement in mIoU and a 4.72% improvement in mAcc compared to the baseline, effectively suppressing false negatives and boosting the boundary recognition ability.

**(3) Breakthrough in Small-Scale Defect Detection with LGANet-CDF:** Recognizing the persistent challenge of tiny defects on EV battery components, I proposed the Cross-Attention-Based Cross-Density Feature Fusion (CDF) module. This module was designed to weave together high-definition spatial details from lower layers with rich contextual semantic information from deeper layers. Experimental outcomes were very promising, as the final model resulted in substantial relative improvements in the IoU metrics for burst points, molten beads, and pinholes by 1.89%, 4.66%, and 5.08%, respectively, when compared to the LGANet model that did not use CDF.

**(4) Streamlined Industrial Usability:** Lastly, the development of a bespoke PyQt software platform encapsulated the proposed approach into a powerful, user-oriented package. This work addresses the growing industrial demand for automated inspection and delivers a simple, visual, and accessible tool for non-specialist workers. Its deployment capabilities allow for effective monitoring, minimal production downtime, and transparent data logging.

Looking toward the future, several avenues for extending this work have been identified. The generalizability of the model could be further improved by incorporating additional scarce defect classes and diverse noise conditions. In terms of practical deployment, the speed of the algorithm is currently bound by high-end GPU devices. Exploring lightweight model architectures, coupled with model compression techniques such as pruning and knowledge distillation, represents a promising direction for achieving real-time, high-throughput inspection on edge devices. Despite the promise shown by deep learning for 3D point cloud analysis in EV battery manufacturing, multimodal fusion of 2D texture and 3D geometry information is another path forward to inspect features that are only visible optically. With such developments, the integration levels of automated manufacturing and intelligent quality inspection on EV battery production lines can be significantly elevated.

Scroll to Top