Read this first. RetinaScope is a course demonstration of Transformer models applied to ophthalmic data. It is not a medical device, it has not been validated on any prospective cohort, and it must never inform the care of a real patient. The classifier was trained on one curated dataset from a small number of sites and knows only four categories, so anything outside those four gets forced into the nearest one. Please do not upload identifiable patient data.
A Vision Transformer reads a single macular B-scan and returns a posterior over four categories. The heat map beside it is an attention rollout, which traces how much each image patch feeds the classification token once the residual stream is accounted for across all twelve blocks.
Example B-scans (Kermany OCT2017 test split, CC BY 4.0)