SignMAE: Segmentation-Driven Self-supervised Learning for Sign Language Recognition

K Xie, Z Cai, K Stefanov - International Conference on Pattern Recognition, 2026 - Springer
K Xie, Z Cai, K Stefanov
International Conference on Pattern Recognition, 2026Springer
Subtle hand differences make sign language recognition challenging, yet many existing
methods rely on encoders pretrained on generic action datasets that poorly capture such
fine-grained cues. We propose a self-supervised pretraining method for sign language
recognition that uses segmentation-based masking to adapt to the presence and motion of
key body parts, rather than treating hand poses as static visual tokens. The resulting mask-
and-reconstruct objective improves fine-grained sign representation learning. On WLASL …
Abstract
Subtle hand differences make sign language recognition challenging, yet many existing methods rely on encoders pretrained on generic action datasets that poorly capture such fine-grained cues. We propose a self-supervised pretraining method for sign language recognition that uses segmentation-based masking to adapt to the presence and motion of key body parts, rather than treating hand poses as static visual tokens. The resulting mask-and-reconstruct objective improves fine-grained sign representation learning. On WLASL, NMFs-CSL, and Slovo, our encoder achieves state-of-the-art performance, improving per-instance and per-class Top-1 accuracy while using fewer input frames and modalities than comparable encoders.
Springer