MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Explore this paper's citation graph

Summary

MiniGPT-4 is presented, which aligns a frozen visual encoder with a frozen advanced LLM, Vicuna, using one projection layer to uncovers that properly aligning the visual features with an advanced large language model can possess numerous advanced multi-modal abilities demonstrated by G PT-4.

Type
preprint
Published
2023-04-20
Cited by
3,328
References
62
Access
Open access

Keywords

Computer science, Usability, Repetition (rhetorical device), Modal, Code (set theory)

References

Cited by

Related papers