article
Automatically identifying the biological processes implicated in cancer development remains one of the fundamental challenges in biomedical research. In this study, we address the task of multi-label classification (MLC) of biomedical abstracts according to the HoC (Hallmarks of Cancer) framework, which organizes key mechanisms involved in tumor progression. To tackle this challenge, we propose a hybrid PubMedBERT-CNN model that combines the contextual understanding of PubMedBERT, a transformer pre-trained on biomedical literature, with convolutional layers tailored for multi-label prediction. Our approach includes rigorous preprocessing steps, such as domain-specific tokenization, label vectorization, and class imbalance mitigation. Evaluated on the HoC dataset, the model achieves an F1-score of 0.89 on the test set, outperforming baseline methods including CNN, SVM, and GHS-NET. Beyond strong overall results, the model also performs consistently well across most individual Hallmarks, confirming its robustness. These findings highlight the potential of transformer-based models, when carefully adapted, to extract structured knowledge from unstructured biomedical texts in oncology research.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/sita67914.2025.11273579
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.