Scenario Classification via Imagery and Video Data for ADAS Trigger Analysis
| dc.contributor.author | Zhao, Weirui | |
| dc.contributor.department | Chalmers tekniska högskola / Institutionen för elektroteknik | sv |
| dc.contributor.examiner | Alvén, Jennifer | |
| dc.contributor.supervisor | Rahm, Robin | |
| dc.date.accessioned | 2026-09-17T12:01:41Z | |
| dc.date.issued | 2026 | |
| dc.date.submitted | ||
| dc.description.abstract | Advanced Driver Assistance Systems (ADAS) generate large volumes of trigger events during development and validation. For low-speed reversing functions such as Rear Auto Brake (RAB), these events must be reviewed to distinguish recurring traffic situations and valid interventions from ambiguous or undesired activations. Manual event-by-event inspection is difficult to scale, while extensive ground-truth labels are generally unavailable. This thesis therefore investigates whether pretrained visual representations and unsupervised clustering can structure industrial RAB trigger data for exploratory analysis. An offline, configuration-driven pipeline was developed to represent each event by one key frame, extract frozen embeddings using DINOv2 ViT-S/14 or ResNet50, optionally reduce their dimensionality to 128 principal components, and cluster them using HDBSCAN or K-Means. The evaluation comprises 18 controlled experiments on the same 8,017 RAB trigger images. Cluster count, noise ratio, Silhouette score, and Davies–Bouldin index are considered together with PCA and UMAP visualizations. DINOv2 combined with HDBSCAN produced the strongest internal cluster structure among the evaluated configurations, and PCA improved every paired DINOv2– HDBSCAN experiment. The best internal result reached a Silhouette score of 0.2709 and a Davies–Bouldin index of 1.4676, while assigning 46.8% of the samples to noise. DINOv2 also retained more events and separated them more strongly than ResNet50 in the two matched comparisons, whereas K-Means covered every event but produced substantially weaker separation. The high HDBSCAN noise ratios indicate considerable visual diversity and show that many events did not form sufficiently dense recurring groups. A manual annotation of 820 events drawn from one reference run examined whether the discovered groups carry consistent semantics. Weighted purity reached 0.931 for trigger validity, 0.947 for scenario cause, and 0.943 for fine object type, while individual clusters still combined opposing validity judgments and the sampled noise group consisted mostly of clear true-positive observations of ordinary static obstacles. Self-supervised embeddings and density-based clustering therefore expose measurable and partly interpretable structure in unlabeled RAB trigger data and can support grouped inspection. The approach remains an exploratory structuring tool rather than an automated scenario classifier, because the semantic evidence came from one run and one reviewer, single key frames omit motion and sensor context, and the effect on review effort was not measured. | |
| dc.identifier.coursecode | EENX30 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12380/312488 | |
| dc.language.iso | eng | |
| dc.setspec.uppsok | Technology | |
| dc.subject | ADAS | |
| dc.subject | rear auto brake | |
| dc.subject | representation learning | |
| dc.subject | DINOv2 | |
| dc.subject | unsupervised clustering | |
| dc.subject | HDBSCAN | |
| dc.subject | scenario discovery | |
| dc.subject | trigger analysis | |
| dc.title | Scenario Classification via Imagery and Video Data for ADAS Trigger Analysis | |
| dc.type.degree | Examensarbete för masterexamen | sv |
| dc.type.degree | Master's Thesis | en |
| dc.type.uppsok | H | |
| local.programme | Information and communication technology (MPICT), MSc |
