We show the formal equivalence of linearised self-attention mechanisms and fast weight controllers from the early '90s, where a ``slow" neural net learns by gradient descent to program the ``fast weights" of another net through sequences of elementary programming instructions which are additive oute...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!