hh.sePublications
Change search
Link to record
Permanent link

Direct link
Hernandez-Diaz, KevinORCID iD iconorcid.org/0000-0002-9696-7843
Publications (10 of 26) Show all publications
Norrby, H., Färm, G., Hernandez-Diaz, K. & Alonso-Fernandez, F. (2026). FGSSNet: Feature-Guided Semantic Segmentation of Real World Floorplans. In: Progress in Artificial Intelligence and Pattern Recognition: 9th International Congress, IWAIPR 2025, Varadero, Cuba, October 14–17, 2025, Proceedings. Paper presented at 9th International Congress, IWAIPR: International Congress on Artificial Intelligence and Pattern Recognition, 2025, October 14–17, 2025, Varadero, Cuba (pp. 39-51). Heidelberg: Springer
Open this publication in new window or tab >>FGSSNet: Feature-Guided Semantic Segmentation of Real World Floorplans
2026 (English)In: Progress in Artificial Intelligence and Pattern Recognition: 9th International Congress, IWAIPR 2025, Varadero, Cuba, October 14–17, 2025, Proceedings, Heidelberg: Springer, 2026, p. 39-51Conference paper, Published paper (Refereed)
Abstract [en]

We introduce FGSSNet, a novel multi-headed feature-guided semantic segmentation (FGSS) architecture designed to improve the generalization ability of wall segmentation on floorplans. FGSSNet features a U-Net segmentation backbone with a multi-headed dedicated feature extractor used to extract domain-specific feature maps which are injected into the latent space of U-Net to guide the segmentation process. This dedicated feature extractor is trained as an encoder-decoder with selected wall patches, representative of the walls present in the input floorplan, to produce a compressed latent representation of wall patches while jointly trained to predict the wall width. In doing so, we expect that the feature extractor encodes texture and width features of wall patches that are useful to guide the wall segmentation process. Our experiments show increased performance by the use of such injected features in comparison to the vanilla U-Net, highlighting the validity of the proposed approach. © 2026 The Author(s), under exclusive license to Springer Nature Switzerland AG

Place, publisher, year, edition, pages
Heidelberg: Springer, 2026
Series
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 16328
National Category
Signal Processing Computer Sciences
Identifiers
urn:nbn:se:hh:diva-58587 (URN)10.1007/978-3-032-11358-0_4 (DOI)2-s2.0-105029895778 (Scopus ID)978-3-032-11357-3 (ISBN)978-3-032-11358-0 (ISBN)
Conference
9th International Congress, IWAIPR: International Congress on Artificial Intelligence and Pattern Recognition, 2025, October 14–17, 2025, Varadero, Cuba
Funder
Swedish Research Council
Available from: 2026-03-23 Created: 2026-03-23 Last updated: 2026-04-22Bibliographically approved
Alonso-Fernandez, F., Hernandez-Diaz, K., Buades Rubio, J. M. & Bigun, J. (2026). Leveraging Large-Scale Face Datasets for Deep Periocular Recognition via Ocular Cropping. In: Gerhard Goos; Juris Hartmanis (Ed.), Progress in Artificial Intelligence and Pattern Recognition - 9th International Congress, IWAIPR 2025, Veradero, Cuba, October 14-17, 2025: Proceedings. Paper presented at 9th International Workshop on Artificial Intelligence and Pattern Recognition (IWAIPR 2025), Veradero, Cuba, October 14-17, 2025 (pp. 26-38). Cham: Springer
Open this publication in new window or tab >>Leveraging Large-Scale Face Datasets for Deep Periocular Recognition via Ocular Cropping
2026 (English)In: Progress in Artificial Intelligence and Pattern Recognition - 9th International Congress, IWAIPR 2025, Veradero, Cuba, October 14-17, 2025: Proceedings / [ed] Gerhard Goos; Juris Hartmanis, Cham: Springer, 2026, p. 26-38Conference paper, Published paper (Refereed)
Abstract [en]

We focus on ocular biometrics, specifically the periocular region (the area around the eye), which offers high discrimination and minimal acquisition constraints. We evaluate three Convolutional Neural Network architectures of varying depth and complexity to assess their effectiveness for periocular recognition. The networks are trained on 1,907,572 ocular crops extracted from the large-scale VGGFace2 database. This significantly contrasts with existing works, which typically rely on small-scale periocular datasets for training having only a few thousand images. Experiments are conducted with ocular images from VGGFace2-Pose, a subset of VGGFace2 containing in-the-wild face images, and the UFPR-Periocular database, which consists of selfies captured via mobile devices with user guidance on the screen. Due to the uncontrolled conditions of VGGFace2, the Equal Error Rates (EERs) obtained with ocular crops range from 9–15%, noticeably higher than the 3–6% EERs achieved using full-face images. In contrast, UFPR-Periocular yields significantly better performance (EERs of 1–2%), thanks to higher image quality and more consistent acquisition protocols. To the best of our knowledge, these are the lowest reported EERs on the UFPR dataset to date.

© 2026 The Author(s), under exclusive license to Springer Nature Switzerland AG

Place, publisher, year, edition, pages
Cham: Springer, 2026
Series
Lecture Notes in Computer Science ; 16328
Keywords
Periocular biometrics, Ocular Recognition, Partial face recognition, Ocular crops, Convolutional Neural Networks (CNNs), Transfer learning, VGGFace2 database, UFPR database
National Category
Signal Processing
Identifiers
urn:nbn:se:hh:diva-57693 (URN)10.1007/978-3-032-11358-0_3 (DOI)2-s2.0-105029898917 (Scopus ID)978-3-032-11357-3 (ISBN)978-3-032-11358-0 (ISBN)
Conference
9th International Workshop on Artificial Intelligence and Pattern Recognition (IWAIPR 2025), Veradero, Cuba, October 14-17, 2025
Funder
EU, Horizon Europe, PopEyeSwedish Research Council
Available from: 2025-10-30 Created: 2025-10-30 Last updated: 2026-05-11Bibliographically approved
Alonso-Fernandez, F., Hernandez-Diaz, K., Buades Rubio, J. M., Tiwari, P. & Bigun, J. (2025). Deep network pruning: A comparative study on CNNs in face recognition. Pattern Recognition Letters, 189, 221-228
Open this publication in new window or tab >>Deep network pruning: A comparative study on CNNs in face recognition
Show others...
2025 (English)In: Pattern Recognition Letters, ISSN 0167-8655, E-ISSN 1872-7344, Vol. 189, p. 221-228Article in journal (Refereed) Published
Abstract [en]

The widespread use of mobile devices for all kinds of transactions makes necessary reliable and real-time identity authentication, leading to the adoption of face recognition (FR) via the cameras embedded in such devices. Progress of deep Convolutional Neural Networks (CNNs) has provided substantial advances in FR. Nonetheless, the size of state-of-the-art architectures is unsuitable for mobile deployment, since they often encompass hundreds of megabytes and millions of parameters. We address this by studying methods for deep network compression applied to FR. In particular, we apply network pruning based on Taylor scores, where less important filters are removed iteratively. The method is tested on three networks based on the small SqueezeNet (1.24M parameters) and the popular MobileNetv2 (3.5M) and ResNet50 (23.5M) architectures. These have been selected to showcase the method on CNNs with different complexities and sizes. We observe that a substantial percentage of filters can be removed with minimal performance loss. Also, filters with the highest amount of output channels tend to be removed first, suggesting that high-dimensional spaces within popular CNNs are over-dimensioned. The models of this paper are available at https://github.com/HalmstadUniversityBiometrics/CNN-pruning-for-face-recognition. © 2025.

Place, publisher, year, edition, pages
Amsterdam: Elsevier, 2025
Keywords
Convolutional Neural Networks, Deep learning, Face recognition, Mobile biometrics, Network pruning, Taylor expansion
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:hh:diva-55571 (URN)10.1016/j.patrec.2025.01.023 (DOI)2-s2.0-85217214565 (Scopus ID)
Funder
Vinnova, PID2022-136779OB-C32Swedish Research CouncilEuropean Commission
Note

This work was partly done while F. A.-F. was a visiting researcher at the University of the Balearic Islands . F. A.-F., K. H.-D., and J. B. thank the Swedish Research Council (VR) and the Swedish Innovation Agency (VINNOVA) for funding their research. This work is part of the Project PID2022-136779OB-C32 (PLEISAR) funded by MICIU/ AEI /10.13039/501100011033/ and FEDER, EU.

Available from: 2025-02-28 Created: 2025-02-28 Last updated: 2025-10-01Bibliographically approved
Alonso-Fernandez, F., Hernandez-Diaz, K., Buades, J. M., Raja, K. & Bigun, J. (2025). Exploring Complementarity and Explainability in CNNs for Periocular Verification Across Acquisition Distances. In: Proceedings of the 24th International Conference of the Biometrics Special Interest Group: . Paper presented at 24th International Conference of the Biometrics Special Interest Group, BIOSIG, Darmstadt, Germany, 25-26 September, 2025 (pp. 1-7). IEEE
Open this publication in new window or tab >>Exploring Complementarity and Explainability in CNNs for Periocular Verification Across Acquisition Distances
Show others...
2025 (English)In: Proceedings of the 24th International Conference of the Biometrics Special Interest Group, IEEE, 2025, p. 1-7Conference paper, Published paper (Refereed)
Abstract [en]

We study the complementarity of different CNNs for periocular verification at different distances on the UBIPr database. We train three architectures of increasing complexity (SqueezeNet, MobileNetv2, and ResNet50) on a large set of eye crops from VGGFace2. We analyse performance with cosine and χ2 metrics, compare different network initialisations, and apply score-level fusion via logistic regression. In addition, we use LIME heatmaps and Jensen–Shannon divergence to compare attention patterns of the CNNs. While ResNet50 consistently performs best individually, the fusion provides substantial gains, especially when combining all three networks. Heatmaps show that networks usually focus on distinct regions of a given image, which explains their complementarity. Our method significantly outperforms previous works on UBIPr, achieving a new state-of-the-art.

Place, publisher, year, edition, pages
IEEE, 2025
Series
GI-Edition Lecture Notes in Informatics (LNI), ISSN 1617-5468, E-ISSN 2944-7682
National Category
Signal Processing
Identifiers
urn:nbn:se:hh:diva-57692 (URN)10.1109/BIOSIG65492.2025.11358084 (DOI)
Conference
24th International Conference of the Biometrics Special Interest Group, BIOSIG, Darmstadt, Germany, 25-26 September, 2025
Funder
EU, Horizon Europe, PopEyeSwedish Research Council
Available from: 2025-10-30 Created: 2025-10-30 Last updated: 2026-03-27Bibliographically approved
Alonso-Fernandez, F., Hernandez-Diaz, K., Tiwari, P. & Bigun, J. (2024). Combined CNN and ViT features off-the-shelf: Another astounding baseline for recognition. In: Proceedings - 16th IEEE International Workshop on Information Forensics and Security, WIFS 2024: . Paper presented at 16th IEEE International Workshop on Information Forensics and Security, WIFS 2024, Rome, Italy, December 2-5, 2024 (pp. 1-6). IEEE
Open this publication in new window or tab >>Combined CNN and ViT features off-the-shelf: Another astounding baseline for recognition
2024 (English)In: Proceedings - 16th IEEE International Workshop on Information Forensics and Security, WIFS 2024, IEEE, 2024, p. 1-6Conference paper, Published paper (Refereed)
Abstract [en]

We apply pre-trained architectures, originally developed for the ImageNet Large Scale Visual Recognition Challenge, for periocular recognition. These architectures have demon-strated significant success in various computer vision tasks beyond the ones for which they were designed. This work builds on our previous study using off-the-shelf Convolutional Neural Network (CNN) and extends it to include the more recently proposed Vision Transformers (ViT). Despite being trained for generic object classification, middle-layer features from CNNs and ViTs are a suitable way to recognize individuals based on periocular images. We also demonstrate that CNNs and ViTs are highly complementary since their combination results in boosted accuracy. In addition, we show that a small portion of these pre-trained models can achieve good accuracy, resulting in thinner models with fewer parameters, suitable for resource-limited environments such as mobiles. This efficiency improves if traditional handcrafted features are added as well. ©2024 IEEE.

Place, publisher, year, edition, pages
IEEE, 2024
Series
IEEE International Workshop on Information Forensics and Security, ISSN 2157-4766, E-ISSN 2157-4774
Keywords
Periocular recognition, deep representation, biometrics, transfer learning, one-shot learning, Convolutional Neural Network, Vision Transformers
National Category
Signal Processing
Identifiers
urn:nbn:se:hh:diva-54713 (URN)10.1109/WIFS61860.2024.10810712 (DOI)001422478100039 ()2-s2.0-85215518296 (Scopus ID)979-8-3503-6442-2 (ISBN)979-8-3503-6443-9 (ISBN)
Conference
16th IEEE International Workshop on Information Forensics and Security, WIFS 2024, Rome, Italy, December 2-5, 2024
Funder
Swedish Research CouncilVinnova
Available from: 2024-10-06 Created: 2024-10-06 Last updated: 2025-10-01Bibliographically approved
Hernandez-Diaz, K. (2024). Ocular Recognition in Unconstrained Sensing Environments. (Doctoral dissertation). Halmstad: Halmstad University Press
Open this publication in new window or tab >>Ocular Recognition in Unconstrained Sensing Environments
2024 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

This thesis focuses on the problem of increasing flexibility in the acquisition and application of biometric recognition systems based on the ocular region. While the ocular area is one of the oldest and most widely studied biometric regions thanks to its rich and discriminative elements and characteristics, most modalities such as retina, iris, eye movements, or oculomotor plant have limitations regarding data acquisition. Some require a specific type of illumination like the iris, a limited distance range like eye movements, or specific sensors and user collaboration like the retina. In this context, this thesis focuses on the periocular region, which stands out as the ocular modality with the fewest acquisition constraints. 

The first part focuses on using middle-layers' deep representation of pre-trained CNNs as a one-shot learning method, along with simple distance-based metrics and similarity scores for periocular recognition. This approach tackles the issue of limited data availability and collection for biometric recognition systems by eliminating the need to train the models for the target data. Furthermore, it allows seamless transitions between identification and verification scenarios with a single model, and tackles the problem of the open-world setting and training bias of CNNs. We demonstrate that off-the-shelf features from middle-layers can outperform CNNs trained for the target domain that followed a more extensive training strategy when target data is limited.

The second part of the thesis analyzes traditional methods for biometric systems in the context of periocular recognition. Nowadays, these methods are often overlooked in favor of deep learning solutions. However, we show that they can still outperform heavily trained CNNs in closed-world and open-world settings and can be used in conjunction with CNNs to further improve recognition performance. Moreover, we investigate the use of the complex structure tensor as a handcrafted texture extractor at the input of CNNs. We show that CNNs can benefit from this explicit textural information in terms of performance and convergence, offering the potential for network compression and explainability of the features used. We demonstrate that CNNs may not easily access the orientation information present in the images that are exploited in some more traditional approaches.

The final part of the thesis addresses the analysis of periocular recognition under different light spectra and the cross-spectral scenario. More specifically, we analyze the performance of the proposed methods under different light spectra. We also investigate the cross-spectral scenario for one-shot learning with middle-layers' deep representations and explore the possibility of bridging the domain gap in the cross-spectral scenario by training generative networks. This allows using simpler models and algorithms trained on a single spectrum.

Place, publisher, year, edition, pages
Halmstad: Halmstad University Press, 2024. p. 49
Series
Halmstad University Dissertations ; 114
Keywords
Biometrics, Computer Vision, Pattern Recognition, Periocular Recognition
National Category
Signal Processing Computer graphics and computer vision
Identifiers
urn:nbn:se:hh:diva-53257 (URN)978-91-89587-43-4 (ISBN)978-91-89587-42-7 (ISBN)
Public defence
2024-05-28, S3030, Kristian IV:s väg 3, 08:00 (English)
Opponent
Supervisors
Available from: 2024-04-24 Created: 2024-04-24 Last updated: 2025-10-01Bibliographically approved
Alonso-Fernandez, F., Hernandez-Diaz, K., Buades Rubio, J. M. & Bigun, J. (2024). SqueezerFaceNet: Reducing a Small Face Recognition CNN Even More Via Filter Pruning. In: Hernández Heredia, Y.; Milián Núñez, V.; Ruiz Shulcloper, J. (Ed.), Progress in Artificial Intelligence and Pattern Recognition. IWAIPR 2023.: . Paper presented at VIII International Workshop on Artificial Intelligence and Pattern Recognition, IWAIPR, Varadero, Cuba, September 27-29, 2023 (pp. 349-361). Cham: Springer, 14335
Open this publication in new window or tab >>SqueezerFaceNet: Reducing a Small Face Recognition CNN Even More Via Filter Pruning
2024 (English)In: Progress in Artificial Intelligence and Pattern Recognition. IWAIPR 2023. / [ed] Hernández Heredia, Y.; Milián Núñez, V.; Ruiz Shulcloper, J., Cham: Springer, 2024, Vol. 14335, p. 349-361Conference paper, Published paper (Refereed)
Abstract [en]

The widespread use of mobile devices for various digital services has created a need for reliable and real-time person authentication. In this context, facial recognition technologies have emerged as a dependable method for verifying users due to the prevalence of cameras in mobile devices and their integration into everyday applications. The rapid advancement of deep Convolutional Neural Networks (CNNs) has led to numerous face verification architectures. However, these models are often large and impractical for mobile applications, reaching sizes of hundreds of megabytes with millions of parameters. We address this issue by developing SqueezerFaceNet, a light face recognition network which less than 1M parameters. This is achieved by applying a network pruning method based on Taylor scores, where filters with small importance scores are removed iteratively. Starting from an already small network (of 1.24M) based on SqueezeNet, we show that it can be further reduced (up to 40%) without an appreciable loss in performance. To the best of our knowledge, we are the first to evaluate network pruning methods for the task of face recognition. © 2024, The Author(s), under exclusive license to Springer Nature Switzerland AG.

Place, publisher, year, edition, pages
Cham: Springer, 2024
Series
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 14335
Keywords
Face recognition, Mobile Biometrics, CNN pruning, Taylor scores
National Category
Signal Processing
Identifiers
urn:nbn:se:hh:diva-51299 (URN)10.1007/978-3-031-49552-6_30 (DOI)2-s2.0-85180788350 (Scopus ID)978-3-031-49551-9 (ISBN)978-3-031-49552-6 (ISBN)
Conference
VIII International Workshop on Artificial Intelligence and Pattern Recognition, IWAIPR, Varadero, Cuba, September 27-29, 2023
Funder
Swedish Research CouncilVinnova
Note

Funding: F. A.-F., K. H.-D., and J. B. thank the Swedish Research Council (VR) and the Swedish Innovation Agency (VINNOVA) for funding their research. Author J. M. B. thanks the project EX-PLAINING - "Project EXPLainable Artificial INtelligence systems for health and well-beING", under Spanish national projects funding (PID2019-104829RA-I00/AEI/10.13039/501100011033).

Available from: 2023-07-20 Created: 2023-07-20 Last updated: 2025-10-01Bibliographically approved
Alonso-Fernandez, F., Hernandez-Diaz, K., Buades, J. M., Tiwari, P. & Bigun, J. (2023). An Explainable Model-Agnostic Algorithm for CNN-Based Biometrics Verification. In: 2023 IEEE International Workshop on Information Forensics and Security (WIFS): . Paper presented at 2023 IEEE International Workshop on Information Forensics and Security, WIFS 2023, Nürnberg, Germany, 4-7 December, 2023. Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>An Explainable Model-Agnostic Algorithm for CNN-Based Biometrics Verification
Show others...
2023 (English)In: 2023 IEEE International Workshop on Information Forensics and Security (WIFS), Institute of Electrical and Electronics Engineers (IEEE), 2023Conference paper, Published paper (Refereed)
Abstract [en]

This paper describes an adaptation of the Local Interpretable Model-Agnostic Explanations (LIME) AI method to operate under a biometric verification setting. LIME was initially proposed for networks with the same output classes used for training, and it employs the softmax probability to determine which regions of the image contribute the most to classification. However, in a verification setting, the classes to be recognized have not been seen during training. In addition, instead of using the softmax output, face descriptors are usually obtained from a layer before the classification layer. The model is adapted to achieve explainability via cosine similarity between feature vectors of perturbated versions of the input image. The method is showcased for face biometrics with two CNN models based on MobileNetv2 and ResNet50. © 2023 IEEE.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2023
Keywords
Biometrics, Explainable AI, Face recognition, XAI
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:hh:diva-52721 (URN)10.1109/WIFS58808.2023.10374866 (DOI)2-s2.0-85183463933 (Scopus ID)9798350324914 (ISBN)
Conference
2023 IEEE International Workshop on Information Forensics and Security, WIFS 2023, Nürnberg, Germany, 4-7 December, 2023
Projects
EXPLAINING - ”Project EXPLainable Artificial INtelligence systems for health and well-beING”
Funder
Swedish Research CouncilVinnova
Available from: 2024-02-16 Created: 2024-02-16 Last updated: 2025-10-01Bibliographically approved
Kolf, J. N., Alonso-Fernandez, F., Hernandez-Diaz, K., Bigun, J. & Yang, B. (2023). EFaR 2023: Efficient Face Recognition Competition. In: 2023 IEEE International Joint Conference on Biometrics, IJCB 2023: . Paper presented at IEEE International Joint Conference on Biometrics (IJCB 2023), Ljubljana, Slovenia, 25-28 September 2023. IEEE
Open this publication in new window or tab >>EFaR 2023: Efficient Face Recognition Competition
Show others...
2023 (English)In: 2023 IEEE International Joint Conference on Biometrics, IJCB 2023, IEEE, 2023Conference paper, Published paper (Refereed)
Abstract [en]

This paper presents the summary of the Efficient Face Recognition Competition (EFaR) held at the 2023 International Joint Conference on Biometrics (IJCB 2023). The competition received 17 submissions from 6 different teams. To drive further development of efficient face recognition models, the submitted solutions are ranked based on a weighted score of the achieved verification accuracies on a diverse set of benchmarks, as well as the deployability given by the number of floating-point operations and model size. The evaluation of submissions is extended to bias, cross-quality, and large-scale recognition benchmarks. Overall, the paper gives an overview of the achieved performance values of the submitted solutions as well as a diverse set of baselines. The submitted solutions use small, efficient network architectures to reduce the computational cost, some solutions apply model quantization. An outlook on possible techniques that are underrepresented in current solutions is given as well. © 2023 IEEE.

Place, publisher, year, edition, pages
IEEE, 2023
Series
IEEE International Conference on Biometrics, Theory, Applications and Systems, ISSN 2474-9680, E-ISSN 2474-9699
National Category
Signal Processing
Identifiers
urn:nbn:se:hh:diva-52967 (URN)10.1109/IJCB57857.2023.10448917 (DOI)001180818700054 ()2-s2.0-85171755032 (Scopus ID)
Conference
IEEE International Joint Conference on Biometrics (IJCB 2023), Ljubljana, Slovenia, 25-28 September 2023
Funder
Swedish Research CouncilVinnova
Note

Acknowledgment: This research work has been funded by the German Federal Ministry of Education and Research and the Hessian Ministry of Higher Education, Research, Science and the Arts within their joint support of the National Research Center for Applied Cybersecurity ATHENE. This work has been partially funded by the German Federal Ministry of Education and Research (BMBF) through the Software Campus Project.

Available from: 2024-03-26 Created: 2024-03-26 Last updated: 2025-10-01Bibliographically approved
Zell, O., Påsson, J., Hernandez-Diaz, K., Alonso-Fernandez, F. & Nilsson, F. (2023). Image-Based Fire Detection in Industrial Environments with YOLOv4. In: Maria De Marsico; Gabriella Sanniti di Baja; Ana Fred (Ed.), Proceedings of the 12th International Conference on Pattern Recognition Applications and Methods ICPRAM: . Paper presented at 12th International Conference on Pattern Recognition Applications and Methods, ICPRAM, Lisbon, Portugal, February 22-24, 2023 (pp. 379-386). Setúbal: SciTePress, 1
Open this publication in new window or tab >>Image-Based Fire Detection in Industrial Environments with YOLOv4
Show others...
2023 (English)In: Proceedings of the 12th International Conference on Pattern Recognition Applications and Methods ICPRAM / [ed] Maria De Marsico; Gabriella Sanniti di Baja; Ana Fred, Setúbal: SciTePress, 2023, Vol. 1, p. 379-386Conference paper, Published paper (Refereed)
Abstract [en]

Fires have destructive power when they break out and affect their surroundings on a devastatingly large scale. The best way to minimize their damage is to detect the fire as quickly as possible before it has a chance to grow. Accordingly, this work looks into the potential of AI to detect and recognize fires and reduce detection time using object detection on an image stream. Object detection has made giant leaps in speed and accuracy over the last six years, making real-time detection feasible. To our end, we collected and labeled appropriate data from several public sources, which have been used to train and evaluate several models based on the popular YOLOv4 object detector. Our focus, driven by a collaborating industrial partner, is to implement our system in an industrial warehouse setting, which is characterized by high ceilings. A drawback of traditional smoke detectors in this setup is that the smoke has to rise to a sufficient height. The AI models brought forward in this research managed to outperform these detectors by a significant amount of time, providing precious anticipation that could help to minimize the effects of fires further.

Place, publisher, year, edition, pages
Setúbal: SciTePress, 2023
Series
International Conference on Pattern Recognition Applications and Methods (ICPRAM), ISSN 2184-4313
Keywords
Fire detection, Smoke Detection, Machine learning, Computer Vision, YOLOv4
National Category
Signal Processing
Identifiers
urn:nbn:se:hh:diva-48793 (URN)10.5220/0011689400003411 (DOI)978-989-758-626-2 (ISBN)
Conference
12th International Conference on Pattern Recognition Applications and Methods, ICPRAM, Lisbon, Portugal, February 22-24, 2023
Projects
2021-05038 Vinnova DIFFUSE Disentanglement of Features For Utilization in Systematic Evaluation
Funder
Swedish Research CouncilVinnova
Available from: 2022-12-09 Created: 2022-12-09 Last updated: 2025-10-01Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-9696-7843

Search in DiVA

Show all publications