Generalization of Monocular Semantic BEV Prediction in Low Annotated Regimes for Unstructured Off-Road Environments

dc.contributor.authorMagnusson, Max
dc.contributor.authorLind, Mattias
dc.contributor.departmentChalmers tekniska högskola / Institutionen för elektrotekniksv
dc.contributor.examinerHammarstrand, Lars
dc.contributor.supervisorBergsjö, Dag
dc.date.accessioned2026-07-06T08:45:08Z
dc.date.issued2026
dc.date.submitted
dc.description.abstractAutonomous navigation in unstructured off-road environments presents distinct challenges from urban driving. Terrain boundaries are ambiguous, annotated datasets are scarce, and in military contexts active sensors such as LiDAR or radar carry the risk of revealing a vehicle’s position. This thesis addresses the problem with a weakly-supervised perception pipeline that relies on passive camera input at inference time, using LiDAR only offline to provide geometric supervision during training. The Vision Foundation Model CAT-Seg is used to generate semantic pseudo-labels across five collapsed traversability classes. These are projected into Bird’s Eye View space using LiDAR-assisted depth integration from DepthAnythingV2 calibrated by LiDAR through RANSAC alignment, producing 100×100 BEV semantic grids as training targets without any manual annotation. Three BEV architectures, Lift Splat Shoot, FocusBEV and SimpleBEV, are trained under three supervision regimes: manually annotated ground truth (≈1,000 frames), VLM-derived pseudolabels (≈18,000 frames), and a pseudo-label model fine-tuned on annotated data. The Great Outdoors dataset is used for in-distribution testing and RELLIS-3D for out-of-distribution testing. In distribution, pseudo-label trained models perform comparably to manually supervised ones despite lower label quality, with the large unlabeled volume compensating for pseudo-label noise. Out of distribution, pseudolabel training improves generalization across all three architectures, most clearly for LSS (0.141 to 0.214 mIoU) and FocusBEV (0.151 to 0.199 mIoU), inverting the in-distribution ranking where annotated training was equal or better. SimpleBEV consistently achieves the strongest out-of-distribution performance across training conditions in our experiments. These results suggest that VLM-produced pseudo-labeling can be a viable alternative to manual annotation for off-road BEV perception when generalization to new terrain is the primary concern.
dc.identifier.coursecodeEENX30
dc.identifier.urihttps://hdl.handle.net/20.500.12380/311857
dc.language.isoeng
dc.setspec.uppsokTechnology
dc.subjectoff-road autonomy
dc.subjectLiDAR
dc.subjectUGV
dc.subjectweak-supervision
dc.subjectVision Foundation Models
dc.subjectBird’s Eye View
dc.subjectpseudo ground-truth
dc.subjectmanual annotation
dc.titleGeneralization of Monocular Semantic BEV Prediction in Low Annotated Regimes for Unstructured Off-Road Environments
dc.type.degreeExamensarbete för masterexamensv
dc.type.degreeMaster's Thesisen
dc.type.uppsokH
local.programmeComplex adaptive systems (MPCAS), MSc

Ladda ner

Original bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
Master_Thesis_2026_förbättrad.pdf
Size:
3.56 MB
Format:
Adobe Portable Document Format

License bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
license.txt
Size:
2.35 KB
Format:
Item-specific license agreed upon to submission
Description: