article
This paper presents a methodology for extracting and structuring procurement data from scanned Summary Minutes documents obtained from the Moroccan Public Procurement Portal. Leveraging web scraping techniques with Scrapy-Selenium and Beautiful Soup, scanned PDFs were collected and processed using PaddleOCR for Optical Character Recognition. The treated text files were stored in a MongoDB database, and structured data extraction was performed using the RAG process with the Mistral LLM. Our methodology resulted in the extraction of 439,048 records, offering valuable insights into procurement practices. We discuss the data extraction process, including cleaning and mining techniques, and highlight limitations encountered, particularly regarding PaddleOCR’s performance across different languages and scripts. Despite challenges, our methodology demonstrates the feasibility of extracting structured data from unstructured sources for informed decision-making in procurement analysis.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/isas64331.2024.10845583
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.