Meta Launches Muse Image, an Agentic AI Image Generator, in Instagram and WhatsApp
Meta Superintelligence Labs shipped Muse Image, a text-to-image model with tool-calling and search grounding, ranking #2 on Arena. A Muse Video preview with native audio is coming next.
Meta Superintelligence Labs shipped Muse Image, its most advanced text-to-image model to date, and it’s already live in the Meta AI app, meta.ai, Instagram Stories in the US, and WhatsApp in select countries.
What makes it different
Muse Image isn’t a standard diffusion model bolted onto a chat interface — it’s built to act agentically. It follows multi-step instructions precisely, handles multi-reference composition for edits (feed it several source images and it composites them coherently), and calls external tools mid-generation: plotting charts, generating QR codes, producing animated GIFs. It also grounds outputs via search, pulling in real-world reference data rather than relying purely on training priors.
Quality scales with test-time compute — spend more inference time, get a better result, the same tradeoff curve that’s become standard in reasoning models is now showing up in image generation.
Benchmarks and provenance
Muse Image ranks #2 on Arena across three separate categories: text-to-image, single-image editing, and multi-image editing. Meta didn’t specify the #1 model in its announcement, but placing second across all three categories against a field that includes Google’s Nano Banana and OpenAI’s image tools is a real result, not a marketing number.
Every output carries an invisible “Content Seal” watermark engineered to survive cropping, recompression, and re-uploading — a direct response to the provenance problem that’s dogged every major image-gen release since Stable Diffusion.
Muse Video is next
Alongside Muse Image, Meta previewed Muse Video, built on the same base model but extended with native audio generation — meaning generated clips come with synchronized sound rather than requiring a separate audio pass. It’s currently ranked #3 on Arena for text-to-video, a preview rather than a full launch.
Why it matters
Meta’s play here is distribution, not just capability. Instagram Stories and WhatsApp put Muse Image in front of billions of users who will never touch a standalone AI tool, which is a different competitive lane than OpenAI or Google chasing developer and enterprise API traffic. If Muse Video ships with the same agentic tool-calling and search-grounding as the image model, Meta will have a native-audio video generator embedded directly in the world’s largest messaging apps — a distribution advantage neither OpenAI’s Sora nor Google’s Veo currently has at that scale.