Machine Learning on the Fly: A Hands-On Tutorial for Streaming Data
2025 (English)In: 2025 IEEE 41st International Conference on Data Engineering, ICDE 2025: Proceedings / [ed] Lisa O’Conner, Piscataway: IEEE Computer Society, 2025, p. 4513-4516Conference paper, Published paper (Refereed)
Abstract [en]
Data stream learning is an emerging machine learning paradigm designed for environments where data arrive continuously and must be processed in real time. Unlike traditional batch learning, which assumes access to a fixed dataset, stream learning addresses the unique challenges of non-stationary distributions, bounded memory, and strict computational constraints. These challenges are increasingly relevant across domains such as IoT, finance, cybersecurity, and environmental monitoring, where timely and adaptive decision-making is essential. This tutorial introduces key concepts and techniques in data stream learning, blending foundational theory with practical demonstrations. It features CapyMOA, an open-source library that provides efficient algorithm implementations through a high-level Python API. We demonstrate the use of this tool through practical examples, with all source code available at https://github.com/adaptive-machine-learning/CapyMOA, and supporting tutorials and installation guides accessible at https://capymoa.org/. © 2025 by The Institute of Electrical and Electronics Engineers, Inc.
Place, publisher, year, edition, pages
Piscataway: IEEE Computer Society, 2025. p. 4513-4516
Series
Data engineering, Proceedings, ISSN 2375-026X
Keywords [en]
data stream learning, concept drift, anomaly detection, supervised learning
National Category
Artificial Intelligence
Identifiers
URN: urn:nbn:se:hh:diva-56259DOI: 10.1109/ICDE65448.2025.00342ISBN: 979-8-3315-3603-9 (print)OAI: oai:DiVA.org:hh-56259DiVA, id: diva2:1965537
Conference
2025 IEEE 41st International Conference on Data Engineering, Hong Kong, China, 19–23 May, 2025
2025-06-092025-06-092025-10-01Bibliographically approved