The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
arXiv · · Significant research
Summary
The paper introduces the Prism Hypothesis, which posits a correspondence between an encoder's feature spectrum and its functional role, with semantic encoders capturing low-frequency components and pixel encoders retaining high-frequency information. Based on this, the authors propose Unified Autoencoding (UAE), a model that harmonizes semantic structure and pixel details using a frequency-band modulator. Experiments on ImageNet and MS-COCO demonstrate that UAE effectively unifies semantic abstraction and pixel-level fidelity, achieving state-of-the-art performance.
Keywords
autoencoding · semantic encoder · pixel encoder · frequency spectrum · representation learning
Get the weekly digest
Top AI stories from the GCC region, every week.