hh.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Li-ViP3D++: Query-Gated Deformable Camera-LiDAR Fusion for End-to-End Perception and Trajectory Prediction
Slovak University of Technology, Bratislava, Slovakia.ORCID iD: 0000-0002-8002-9887
Slovak University of Technology, Bratislava, Slovakia.ORCID iD: 0009-0008-4159-6767
Halmstad University, School of Information Technology. Karlsruhe Institute of Technology, Karlsruhe, Germany.ORCID iD: 0000-0003-4894-4134
Faculty of Informatics and Information Technologies, Slovak University of Technology, Bratislava, Slovakia.ORCID iD: 0000-0001-6622-526X
2026 (English)In: IEEE Access, E-ISSN 2169-3536, Vol. 14, p. 102999-103012Article in journal (Refereed) Epub ahead of print
Abstract [en]

End-to-end perception and trajectory prediction from raw sensor data is one of the key capabilities for autonomous driving. Modular pipelines restrict information flow and can amplify upstream errors. Recent query-based, fully differentiable perception-and-prediction (PnP) models mitigate these issues, yet the complementarity of cameras and LiDAR in the query-space has not been sufficiently explored. Models often rely on fusion schemes that introduce heuristic alignment and discrete selection steps which prevent full utilization of available information and can introduce unwanted bias. We propose Li-ViP3D++, a query-based multimodal PnP framework that introduces Query-Gated Deformable Fusion (QGDF) to integrate multi-view RGB and LiDAR in query space. QGDF 1) aggregates image evidence via masked attention across cameras and feature levels, 2) extracts LiDAR context through fully differentiable BEV sampling with learned per-query offsets, and 3) applies query-conditioned gating to adaptively weight visual and geometric cues per agent. The resulting architecture jointly optimizes detection, tracking, and multi-hypothesis trajectory forecasting in a single end-to-end model. On nuScenes, Li-ViP3D++ improves end-to-end behavior and detection quality, achieving higher EPA (0.505) and mAP (0.616) while substantially reducing false positives (FP ratio 0.069), and it is faster than the prior Li-ViP3D variant (139.82 ms vs. 145.91 ms). Additional experiments were performed to evaluate the impact of the reduced resolution of RGB inputs and missing HD maps on the behavior of the model. The results of the experiments indicate that query-space, fully differentiable camera–LiDAR fusion can increase the robustness of end-to-end PnP without sacrificing deployability. © 2026 The Authors.

Place, publisher, year, edition, pages
Piscataway: Institute of Electrical and Electronics Engineers (IEEE), 2026. Vol. 14, p. 102999-103012
Keywords [en]
Perception, perception and prediction, machine learning, computer vision, deep learning, multimodality, trajectory prediction
National Category
Computer graphics and computer vision
Identifiers
URN: urn:nbn:se:hh:diva-60074DOI: 10.1109/access.2026.3709080Scopus ID: 2-s2.0-105043674826OAI: oai:DiVA.org:hh-60074DiVA, id: diva2:2087083
Available from: 2026-07-17 Created: 2026-07-17 Last updated: 2026-07-17Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Vinel, Alexey

Search in DiVA

By author/editor
Halinkovic, MatejMasarykova, NinaVinel, AlexeyGalinski, Marek
By organisation
School of Information Technology
In the same journal
IEEE Access
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 24 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf