>
>
Sparse Window Transformer Attention for 3D Point Cloud Semantic Segmentation: Benchmark Performance on SemanticKITTI and nuScenes LiDAR Datasets
Sparse Window Transformer Attention for 3D Point Cloud Semantic Segmentation: Benchmark Performance on SemanticKITTI and nuScenes LiDAR Datasets
Publisher : PJPCR
Author(s)
Leon T. Fischer; Ai-Ling M. Wu; Soren N. Eriksson
Abstract
This study investigates sparse window-based transformer attention for 3D LiDAR point cloud semantic segmentation with benchmark evaluation on SemanticKITTI and nuScenes outdoor driving datasets within the context of 3D computer vision and autonomous driving perception, an area of growing scientific importance given its implications for autonomous vehicle semantic perception, outdoor robotics scene understanding, and HD mapping for self-driving systems. Using sparse voxelized window attention transformer architecture with hierarchical feature aggregation, trained with focal loss on SemanticKITTI and nuScenes segmentation benchmarks, we examine sparse window partitioning reducing transformer quadratic attention complexity to linear in point count, enabling efficient global context with local geometric detail preservation through hierarchical aggregation in SemanticKITTI: 22,000 training scans / 4,071 test scans; nuScenes: 28,130 training / 6,019 test with 16 semantic classes drawn from SemanticKITTI outdoor urban LiDAR benchmark and nuScenes autonomous driving dataset with 20 semantic class annotations. Results indicate that SWTrans achieves 78.4% mIoU on SemanticKITTI test (state-of-art at submission) and 81.2% on nuScenes val, with 2.4x faster inference than PointTransformer via sparse attention (p < 0.001), with 78.4% mIoU SemanticKITTI (vs. 72.0% PointTransformer), 2.4x faster inference as the primary quantitative benchmark. Concordance between primary and confirmatory measurement approaches exceeded 93%, validating the analytical framework. These findings contribute empirically to 3D computer vision and autonomous driving perception and carry actionable implications for the design of programs and policies targeting autonomous vehicle semantic perception, outdoor robotics scene understanding, and HD mapping for self-driving systems.
