CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Explore this paper's citation graph
Summary
CogVideoX is adept at producing coherent, long-duration, different shape videos characterized by significant motions, and develops an effective text-video data processing pipeline that includes various data preprocessing strategies and a video captioning method, greatly contributing to the generation quality and semantic alignment.
- Type
- preprint
- Published
- 2024-08-12
- Cited by
- 2,312
- References
- 50
- Access
- Open access
- OpenAlex
- https://openalex.org/W4402387301
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:271855655
Keywords
Transformer, Computer science, Diffusion, Electrical engineering, Engineering
References
- Paper
- MoCoGAN: Decomposing Motion and Content for Video Generation
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Denoising Diffusion Probabilistic Models
- Taming Transformers for High-Resolution Image Synthesis
- Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
- CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
- Imagen Video: High Definition Video Generation with Diffusion Models
- Phenaki: Variable Length Video Generation From Open Domain Textual Description
- LAION-5B: An open large-scale dataset for training next generation image-text models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution
- MAGVIT: Masked Generative Video Transformer
- Effective Long-Context Scaling of Foundation Models
- Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
- CogVLM: Visual Expert for Pretrained Language Models
Cited by
- EDA-DM: Enhanced Distribution Alignment for Post-Training Quantization of Diffusion Models
- VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- Dynamic Typography: Bringing Text to Life via Video Diffusion Prior
- EG4D: Explicit Generation of 4D Object without Score Distillation
- EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
- OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- GenAI Arena: An Open Evaluation Platform for Generative Models
- Aid: Adapting Image2video Diffusion Models for Instruction-Guided Video Prediction
- Training-free Camera Control for Video Generation
- ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning
- ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
- Diffusion Feedback Helps CLIP See Better
- Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
- SkyScript-100M: 1,000,000,000 Pairs of Scripts and Shooting Scripts for Short Drama
- SurGen: Text-Guided Diffusion Model for Surgical Video Generation
- FLUX that Plays Music
- Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
- Pyramidal Flow Matching for Efficient Video Generative Modeling
Related papers
- ИСПОЛЬЗОВAНИЕ ПОТЕНЦИAЛA СОЦИAЛЬНЫХ ПAРТНЕРОВ В ПОДГОТОВКЕ БУДУЩИХ ПЕДAГОГОВ
- Using DataGrid Control to Realize DataBase of Querying in VB6.0
- Study and Two Types of Typical Usage of DataGrid Web Server Control
- STKVS: secure technique for keyframes-based video summarization model
- PACWON: A parallelizing compiler for workstations on a network
- SLA Aware Optimized Task Scheduling Model for Faster Execution of Workloads Among Federated Clouds
- Bidirectional Sort and Choosing a Row to Update or Delete by Click Any Cell in DataGrid
- Edge caching and computing of video chunks in multi-tier wireless networks