Canonical locks that encode part-whole hierarchies
Organizations: Department of Computer Science, University of Central Florida, United States.
Abstract
One of the challenges in representational learning is how to encode part-whole hierarchies in a neural net. Prior works rely on flattening tree-like structures into string-like sequences and training a sequence-to-sequence model via autoregression. While such a representation works for parse-trees in NLP, it is not entirely clear how to make it work for images. Thus, we propose a geometric primitive called canonical locks. The key idea is that parts/wholes can be modelled as higher-dimensional vectors (), and information can be encoded in their relative phase differences. Inductively, the net consists of positionally-bound bottom-up and top-down neural fields, which drive each other to achieve a state of thermal equilibrium. Additionally, we show the existence of a few symmetrical configurations in the net. The computational iterations taken to break these symmetries depend on the angle between parts/wholes arranged on a disk (or more precisely a ring) in higher dimensions. It also appears to have connections to the psychological phenomenon of mental rotation.
Figures & tables
| Symbol | Description | Symbol | Description | |
|---|---|---|---|---|
| Input image | All bottom-up weight matrices | |||
| Current level in the hierarchy, total levels | All top-down weight matrices | |||
| Spatial grid coordinates at level | Outer-loop weights | |||
| Encoder | Inner-Loop weights | |||
| Bottom-up Net at level | Optimizer for | |||
| Bottom-up Net’s weights | Optimizer for top-down net in level |