CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Explore this paper's citation graph

Summary

CogVideoX is adept at producing coherent, long-duration, different shape videos characterized by significant motions, and develops an effective text-video data processing pipeline that includes various data preprocessing strategies and a video captioning method, greatly contributing to the generation quality and semantic alignment.

Type
preprint
Published
2024-08-12
Cited by
2,312
References
50
Access
Open access

Keywords

Transformer, Computer science, Diffusion, Electrical engineering, Engineering

References

Cited by

Related papers