Meta has officially launched its latest image generation model Muse Image and showcased an early preview of Muse Video for the first time. Both solutions were developed by Meta's Superintelligence Labs division and mark significant progress in the field of AI media.
Muse Image: Agentic Approach and Social Context
Muse Image is now available in the Meta AI app, on the meta.ai platform, as well as in Instagram Stories for users in the US and in WhatsApp in a limited number of countries. Meta calls it its most advanced image generation model. The key difference is the agentic architecture: Muse Image uses internet search, tools for writing and executing code, and can refine images during the generation process. The model follows prompts more accurately, can edit images, compose scenes from multiple references, and also takes into account the social context of Instagram.
Integration with Muse Spark and New Capabilities
Muse Image is integrated with Muse Spark. The combination of models leverages code and media generation to create animated GIFs, websites with embedded images, and even interactive visual games. This transforms generative AI from a simple tool for creating images into a full-fledged platform for multimedia creativity.
Muse Video: Preview and Ambitions
In parallel, Meta showed an early preview of Muse Video. According to the company, the model is built on the same pre-training base as Muse Image, delivers high visual quality, and natively supports audio. However, the company acknowledges that it is still refining audio-video synchronization and physically accurate rendering of fast motion. This indicates that the full launch of video generation may take several more months.
Safety and Metrics
Meta has implemented a copyright protection system: all images created with Muse Image in the Meta AI app receive an invisible watermark called Content Seal. According to the company, it persists after cropping, compression, resizing, or screenshotting. In the future, Meta plans to extend this mark to videos as well.
According to data from the internal Arena platform, Muse Image ranks second in the categories of text-to-image, single-image editing, and multi-image editing, trailing only GPT Image 2 from OpenAI. This confirms the high level of competitiveness of the solution.
As an analyst, I note that the agentic approach of Muse Image is an evolutionary step forward. The model's ability to independently use external tools and social network context makes it not just a generator, but a full-fledged creative assistant. However, the issue of audio-video synchronization in Muse Video remains critical — without it, video generation will appear raw. Overall, Meta is confidently strengthening its position in the generative AI race, and I expect that in the coming quarters we will see even deeper integration of these models into the company's ecosystem.