MARATTO

article

AI-Enhanced Techniques for Extracting Structured Data from Unstructured Public Procurement Documents

Abstract

This paper presents a methodology for extracting and structuring procurement data from scanned Summary Minutes documents obtained from the Moroccan Public Procurement Portal. Leveraging web scraping techniques with Scrapy-Selenium and Beautiful Soup, scanned PDFs were collected and processed using PaddleOCR for Optical Character Recognition. The treated text files were stored in a MongoDB database, and structured data extraction was performed using the RAG process with the Mistral LLM. Our methodology resulted in the extraction of 439,048 records, offering valuable insights into procurement practices. We discuss the data extraction process, including cleaning and mining techniques, and highlight limitations encountered, particularly regarding PaddleOCR’s performance across different languages and scripts. Despite challenges, our methodology demonstrates the feasibility of extracting structured data from unstructured sources for informed decision-making in procurement analysis.

Research topics

  • Data Quality and Management
  • Imbalanced Data Classification Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/isas64331.2024.10845583

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.