hh.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Semantics-aware Multi-modal Domain Translation: From LiDAR Point Clouds to Panoramic Color Images
Halmstad University, School of Information Technology, Halmstad Embedded and Intelligent Systems Research (EIS), CAISR - Center for Applied Intelligent Systems Research.ORCID iD: 0000-0002-8067-9521
Middle East Technical Univetsity, Ankara, Turkey.
Halmstad University, School of Information Technology, Halmstad Embedded and Intelligent Systems Research (EIS), CAISR - Center for Applied Intelligent Systems Research.ORCID iD: 0000-0002-5712-6777
2021 (English)In: 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Los Alamitos: IEEE Computer Society, 2021, p. 3032-3041Conference paper, Published paper (Refereed)
Abstract [en]

In this work, we present a simple yet effective framework to address the domain translation problem between different sensor modalities with unique data formats. By relying only on the semantics of the scene, our modular generative framework can, for the first time, synthesize a panoramic color image from a given full 3D LiDAR point cloud. The framework starts with semantic segmentation of the point cloud, which is initially projected onto a spherical surface. The same semantic segmentation is applied to the corresponding camera image. Next, our new conditional generative model adversarially learns to translate the predicted LiDAR segment maps to the camera image counterparts. Finally, generated image segments are processed to render the panoramic scene images. We provide a thorough quantitative evaluation on the SemanticKitti dataset and show that our proposed framework outperforms other strong baseline models. Our source code is available at https://github. com/halmstad-University/TITAN-NET. © 2021 IEEE.

Place, publisher, year, edition, pages
Los Alamitos: IEEE Computer Society, 2021. p. 3032-3041
Keywords [en]
Computer vision, Image segmentation, Laser radar, Three-dimensional displays, Image synthesis, Computational modeling, Semantics, Color
National Category
Computer graphics and computer vision
Identifiers
URN: urn:nbn:se:hh:diva-45850DOI: 10.1109/ICCVW54120.2021.00338ISI: 000739651103014Scopus ID: 2-s2.0-85122119016ISBN: 978-1-6654-0191-3 (electronic)ISBN: 978-1-6654-0192-0 (print)OAI: oai:DiVA.org:hh-45850DiVA, id: diva2:1609341
Conference
2021 IEEE/CVF International Conference on Computer Vision Workshops ICCVW 2021, Montreal, BC, Canada, Virtual, Online, 11-17 October, 2021
Projects
SHARPEN
Funder
Vinnova, 2018-05001Available from: 2021-11-08 Created: 2021-11-08 Last updated: 2025-10-01Bibliographically approved
In thesis
1. Semantics-aware Multi-modal Scene Perception for Autonomous Vehicles
Open this publication in new window or tab >>Semantics-aware Multi-modal Scene Perception for Autonomous Vehicles
2024 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Autonomous vehicles represent the pinnacle of modern technological innovation, navigating complex and unpredictable environments. To do so effectively, they rely on a sophisticated array of sensors. This thesis explores two of the most crucial sensors: LiDARs, known for their accuracy in generating detailed 3D maps of the environment, and RGB cameras, essential for processing visual cues critical for navigation. Together, these sensors form a comprehensive perception system that enables autonomous vehicles to operate safely and efficiently.

However, the reliability of these vehicles has yet to be tested when key sensors fail. The abrupt failure of a camera, for instance, disrupts the vehicle’s perception system, creating a significant gap in sensory input. This thesis addresses this challenge by introducing a novel multi-modal domain translation framework that integrates LiDAR and RGB camera data while ensuring continuous functionality despite sensor failures. At the core of this framework is an innovative model capable of synthesizing RGB images and their corresponding segment maps from raw LiDAR data by exploiting the scene semantics. The proposed framework stands out as the first of its kind, demonstrating for the first time that the scene semantics can bridge the gap across different domains with distinct data structures, such as unorganized sparse 3D LiDAR point clouds and structured 2D camera data. Thus, this thesis represents a significant leap forward in the field, offering a robust solution to the challenge of RGB data recovery without camera sensors.

The practical application of this model is thoroughly explored in the thesis. It involves testing the model’s capability to generate pseudo point clouds from RGB depth estimates, which, when combined with LiDAR data, create an enriched perception dataset. This enriched dataset is pivotal in enhancing object detection capabilities, a fundamental aspect of autonomous vehicle navigation. The quantitative and qualitative evidence reported in this thesis demonstrates that the synthetic generation of data not only compensates for the loss of sensory input but also considerably improves the performance of object detection systems compared to using raw LiDAR data only.

By addressing the critical issue of sensor failure and presenting viable solutions, this thesis contributes to enhancing the safety, reliability, and efficiency of autonomous vehicles. It paves the way for further research and developiment, setting a new standard for autonomous vehicle technology in scenarios of sensor malfunctions or adverse environmental conditions.

Place, publisher, year, edition, pages
Halmstad: Halmstad University Press, 2024. p. 40
Series
Halmstad University Dissertations ; 117
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:hh:diva-53115 (URN)978-91-89587-50-2 (ISBN)978-91-89587-51-9 (ISBN)
Public defence
2024-06-13, Wigforss, hus J, Kristian IV:s väg 3, Halmstad, 09:00 (English)
Opponent
Supervisors
Available from: 2024-05-07 Created: 2024-04-08 Last updated: 2025-10-01Bibliographically approved

Open Access in DiVA

fulltext(831 kB)749 downloads
File information
File name FULLTEXT01.pdfFile size 831 kBChecksum SHA-512
dd6f341d01e4d93f0655413fd4c43e77bf31cfaaea67769f9ba2167ba4fca7aea86921973c262bc5558c7bd68ecfc93f1c943ff6d2fbb14fc47ab4f1886e76dd
Type fulltextMimetype application/pdf

Other links

Publisher's full textScopus

Authority records

Cortinhal, TiagoAksoy, Eren

Search in DiVA

By author/editor
Cortinhal, TiagoAksoy, Eren
By organisation
CAISR - Center for Applied Intelligent Systems Research
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar
Total: 750 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 555 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf