Generalization of Monocular Semantic BEV Prediction in Low Annotated Regimes for Unstructured Off-Road Environments
Hämtar...
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
Autonomous navigation in unstructured off-road environments presents distinct challenges
from urban driving. Terrain boundaries are ambiguous, annotated datasets
are scarce, and in military contexts active sensors such as LiDAR or radar carry
the risk of revealing a vehicle’s position. This thesis addresses the problem with a
weakly-supervised perception pipeline that relies on passive camera input at inference
time, using LiDAR only offline to provide geometric supervision during training.
The Vision Foundation Model CAT-Seg is used to generate semantic pseudo-labels
across five collapsed traversability classes. These are projected into Bird’s Eye
View space using LiDAR-assisted depth integration from DepthAnythingV2 calibrated
by LiDAR through RANSAC alignment, producing 100×100 BEV semantic
grids as training targets without any manual annotation. Three BEV architectures,
Lift Splat Shoot, FocusBEV and SimpleBEV, are trained under three supervision
regimes: manually annotated ground truth (≈1,000 frames), VLM-derived pseudolabels
(≈18,000 frames), and a pseudo-label model fine-tuned on annotated data.
The Great Outdoors dataset is used for in-distribution testing and RELLIS-3D for
out-of-distribution testing. In distribution, pseudo-label trained models perform
comparably to manually supervised ones despite lower label quality, with the large
unlabeled volume compensating for pseudo-label noise. Out of distribution, pseudolabel
training improves generalization across all three architectures, most clearly
for LSS (0.141 to 0.214 mIoU) and FocusBEV (0.151 to 0.199 mIoU), inverting
the in-distribution ranking where annotated training was equal or better. SimpleBEV
consistently achieves the strongest out-of-distribution performance across
training conditions in our experiments. These results suggest that VLM-produced
pseudo-labeling can be a viable alternative to manual annotation for off-road BEV
perception when generalization to new terrain is the primary concern.
Beskrivning
Ämne/nyckelord
off-road autonomy, LiDAR, UGV, weak-supervision, Vision Foundation Models, Bird’s Eye View, pseudo ground-truth, manual annotation
