UPDF AI

Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens

Fan Ma,Xiaojie Jin,3 Authors,Yi Yang

2023 · DOI: 10.48550/arXiv.2312.08870
arXiv.org · 33 Citations

TLDR

V ISTA - LL A MA is proposed, a novel framework that maintains the consistent distance between all visual tokens and any language tokens, irrespective of the generated text length, and presents a sequential visual projector that projects the current video frame into tokens of language space with the assistance of the previous frame.