Reformer: The Efficient Transformer

Explore this paper's citation graph

Summary

This work replaces dot-product attention by one that uses locality-sensitive hashing and uses reversible residual layers instead of the standard residuals, which allows storing activations only once in the training process instead of several times, making the model much more memory-efficient and much faster on long sequences.

Type
article
Published
2020-01-13
Cited by
3,064
References
26
Access
Open access

Keywords

Transformer, Computer science, Locality, Residual, Hash function

References

Cited by

Related papers