MARATTO

article

Multi-Label Biomedical Text Classification of Abstracts According to the Hallmarks of Cancer

Abstract

Automatically identifying the biological processes implicated in cancer development remains one of the fundamental challenges in biomedical research. In this study, we address the task of multi-label classification (MLC) of biomedical abstracts according to the HoC (Hallmarks of Cancer) framework, which organizes key mechanisms involved in tumor progression. To tackle this challenge, we propose a hybrid PubMedBERT-CNN model that combines the contextual understanding of PubMedBERT, a transformer pre-trained on biomedical literature, with convolutional layers tailored for multi-label prediction. Our approach includes rigorous preprocessing steps, such as domain-specific tokenization, label vectorization, and class imbalance mitigation. Evaluated on the HoC dataset, the model achieves an F1-score of 0.89 on the test set, outperforming baseline methods including CNN, SVM, and GHS-NET. Beyond strong overall results, the model also performs consistently well across most individual Hallmarks, confirming its robustness. These findings highlight the potential of transformer-based models, when carefully adapted, to extract structured knowledge from unstructured biomedical texts in oncology research.

Research topics

  • Text and Document Classification Technologies
  • Biomedical Text Mining and Ontologies
  • Topic Modeling

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/sita67914.2025.11273579

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.