MARATTO

review

Record Linkage Approaches in Big Data: A Comprehensive Review

Abstract

Analyzing data and making the right decisions have become crucial objectives in various domains. Record linkage is one of the most important processes for guaranteeing good data quality for analysis. The aim of record linkage is to find records in a dataset that represent the same real-world entity across many different data sources. This process becomes complex in the context of Big Data due to the high volume, variety of sources, and rapid velocity of data. This paper provides a comprehensive review of record linkage processes that can be adaptable to big data, the challenges and issues involved, and the limitations that remain. We give a comparative study of parallel processing approaches for big data that can reduce the execution time of record linkage processes such as Hadoop MapReduce, Apache Spark, and Apache Flink. We find that these last two technologies are more efficient on several characteristics such as data processing, performance, optimization, processing speed, and real-time analysis.

Research topics

  • Data Quality and Management
  • Privacy-Preserving Technologies in Data
  • Cloud Data Security Solutions

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/iscv60512.2024.10620076

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.