/Robust Vision Transformers
Abstract

Apparatuses, systems, and techniques to generate a robust representation of an image. In at least one embodiment, input tokens of an input image are received, and an inference about the input image is generated based on a vision transformer (ViT) system comprising at least one self-attention module to perform token mixing and a channel self-attention module to perform channel processing.

Full Text

What is claimed is:

Apparatuses, systems, and techniques to generate a robust representation of an image. In at least one embodiment, input tokens of an input image are received, and an inference about the input image is generated based on a vision transformer (ViT) system comprising at least one self-attention module to perform token mixing and a channel self-attention module to perform channel processing.
Timeline
Filed
05/12/2026
Published
09/10/2026
Granted
Not Available
IPC Codes(8)
G06V 10/82:using neural networks
G06F 21/10:Protecting distributed programs or content, e.g. vending or licensing of copyrighted material (protection in video systems or pay television H04N 7/16)
G06V 10/30:Noise filtering