The generation of realistic-looking videos from prompts does not indicate that a system understands the physical world. The JEPA (Joint Embedding Predictive Architecture) aims to generate abstract representations of video continuations. Joint Embedding architectures produce better representations of visual inputs than generative architectures.

1m read timeFrom twitter.com
Post cover image
9 Impressions