MARATTO

article · Journal of Information and Intelligence

Automated data processing and feature engineering for deep learning and big data applications: A survey

2024125 citationsOpen accessCape Coast Technical University

In plain language

Modern artificial intelligence aims to design algorithms that learn directly from data, especially within supervised deep learning. Although model training is highly automated, data processing tasks have traditionally required manual collection, cleaning, and augmentation before data can effectively train models. Rising demands to handle large, complex, and heterogeneous datasets for big data applications have spurred the creation of automated techniques. Automated machine learning, known as AutoML, now enables end-to-end systems that convert raw data into usable features by automating intermediate stages. This review covers methods across data preprocessing, including cleaning, labelling, missing value imputation, and categorical encoding. It also examines data augmentation, such as synthetic data generation via generative artificial intelligence, and feature engineering covering extraction, construction, and selection, alongside tools that optimise entire deep learning workflows concurrently.

Key takeaways

  • Conventional deep learning pipelines still frequently depend on manual data collection, preprocessing, and augmentation.
  • Automated machine learning frameworks can transform raw data into useful features for big data applications by automating intermediate stages.
  • Automated data preprocessing methods cover data cleaning, labelling, missing value imputation, and categorical data encoding.
  • Techniques for data augmentation include synthetic data generation using generative artificial intelligence alongside automated feature extraction, construction, and selection.
  • Automated machine learning tools enable the simultaneous optimisation of all stages within deep learning pipelines.

Why it matters

Deep learning algorithms require vast amounts of data, yet preparing that data manually is time-consuming and complex. By automating data cleaning, labelling, and feature engineering, organisations can process massive, heterogeneous datasets much more efficiently. This shift reduces manual effort, speeds up artificial intelligence development, and allows complex big data applications to generate actionable insights more rapidly and reliably.

Commercialisation angle

The review focuses on automated machine learning tools and end-to-end data processing systems for big data and deep learning applications. Potential users include enterprise data science teams and developers seeking to automate repetitive tasks like data labelling, cleaning, and feature selection. As a survey of existing methods, tools, and generative artificial intelligence techniques, it highlights technologies actively applied in machine learning, though the abstract provides no specific deployment metrics or commercially ready product pathways.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Modern approach to artificial intelligence (AI) aims to design algorithms that learn directly from data. This approach has achieved impressive results and has contributed significantly to the progress of AI, particularly in the sphere of supervised deep learning. It has also simplified the design of machine learning systems as the learning process is highly automated. However, not all data processing tasks in conventional deep learning pipelines have been automated. In most cases data has to be manually collected, preprocessed and further extended through data augmentation before they can be effective for training. Recently, special techniques for automating these tasks have emerged. The automation of data processing tasks is driven by the need to utilize large volumes of complex, heterogeneous data for machine learning and big data applications. Today, end-to-end automated data processing systems based on automated machine learning (AutoML) techniques are capable of taking raw data and transforming them into useful features for Big Data tasks by automating all intermediate processing stages. In this work, we present a thorough review of approaches for automating data processing tasks in deep learning pipelines, including automated data preprocessing– e.g., data cleaning, labeling, missing data imputation, and categorical data encoding–as well as data augmentation (including synthetic data generation using generative AI methods) and feature engineering–specifically, automated feature extraction, feature construction and feature selection. In addition to automating specific data processing tasks, we discuss the use of AutoML methods and tools to simultaneously optimize all stages of the machine learning pipeline.

Research topics

  • Machine Learning and Data Classification
  • Machine Learning and Algorithms
  • Advanced Neural Network Applications

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1016/j.jiixd.2024.01.002

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.