Enhancing Multimodal Large Language Models with Vision Detection Models: An Empirical Study
Enhancing Multimodal Large Language Models with Vision Detection Models: An Empirical Study
Qirui Jiao,Daoyuan Chen,2 Authors,Ying Shen
2024 · DOI: 10.48550/arXiv.2401.17981
arXiv.org · 18 Citations
TLDR
An empirical study on enhancing MLLMs with state-of-the-art SOTA object detection and Optical Character Recognition models to improve fine-grained understanding and reduce hallucination in responses.
