MARATTO

article

Knowledge Distillation Between 2D and 3D Vision Transformers for Point Cloud Quality Assessment

Abstract

Point Cloud Quality Assessment (PCQA) is essential for preserving the perceptual fidelity of 3D data across various applications, such as industrial design, autonomous driving, and immersive gaming. Existing no-reference (NR) approaches face significant challenges due to the limited availability of labeled datasets and the irregular structure of point clouds. Methods that rely on 2D projections benefit from the strengths of well-established 2D networks and achieve competitive performance. However, the pre-processing required for these methods makes them less practical for real-time applications. Conversely, point-based techniques directly process 3D data, eliminating the need for projections, but they often struggle to accurately capture perceptual quality. To address these limitations, we introduce a Cross-Modal Knowledge Distillation framework that transfers quality-relevant features from a deep 2D Vision Transformer (Teacher) to a lightweight 3D Vision Transformer (Student). This approach aligns perceptual features across 2D and 3D modalities during training, enabling the 3D model to achieve enhanced quality assessment performance. At inference, the student model requires only the 3D input, making it highly suitable for real-time applications. Experimental results on benchmark datasets demonstrate that the proposed method achieves competitive accuracy with significantly reduced model complexity and inference time.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/ipta66025.2025.11222053

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.