Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens
Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens
Fan Ma,Xiaojie Jin,3 Authors,Yi Yang
2023 · DOI: 10.48550/arXiv.2312.08870
arXiv.org · 33 Citations
TLDR
V ISTA - LL A MA is proposed, a novel framework that maintains the consistent distance between all visual tokens and any language tokens, irrespective of the generated text length, and presents a sequential visual projector that projects the current video frame into tokens of language space with the assistance of the previous frame.
