article
Semantic Textual Similarity (STS) is a critical task in biomedical natural language processing, supporting applications such as clinical information retrieval, summarization, and decision support. Although pretrained transformer models are now widely used, the influence of architectural type, model scale, and pretraining domain on STS performance in biomedical contexts remains insufficiently understood. This paper presents a systematic zero-shot evaluation of nine transformer-based models on the BIOSSES benchmark. The models span encoder-only, decoder-only, and encoder-decoder architectures. We compute sentence similarity using both cosine and normalized Euclidean distances, and assess model performance using Pearson and Spearman correlations with expert-annotated gold scores. Our findings show that small, general-domain encoder-only models, such as MiniLM, consistently outperform larger biomedical-specific models like BioBERT. We also observe that the effectiveness of similarity functions varies by model, and that differences between Pearson and Spearman correlations indicate a disconnect between score calibration and ranking precision. These results challenge prevailing assumptions about the superiority of large-scale or domain-specialized models and highlight the importance of architecture-aware model selection. We conclude by outlining directions for future research, including adaptive similarity metrics and evaluation on more diverse biomedical datasets.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/cist65886.2025.11224208
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.