LAION-5B: An open large-scale dataset for training next generation image-text models

Explore this paper's citation graph

Summary

This work presents LAION-5B - a dataset consisting of 5.85 billion CLIP-filtered image-text pairs, of which 2.32B contain English language, and shows successful replication and fine-tuning of foundational models like CLIP, GLIDE and Stable Diffusion using the dataset, and discusses further experiments enabled with an openly available dataset of this scale.

Type
preprint
Published
2022-10-16
Cited by
5,499
References
109
Access
Open access

Keywords

Computer science, Robustness (evolution), Artificial intelligence, Image (mathematics), Scale (ratio)

References

Cited by

Related papers