The multimedia apocalypse already happened

People already speak in clips, voice notes and camera rolls. Some models understand that world out of the box. The text-first ones still need a transcript.

A mate threw a Discord icon at Muse and told it to make a cool GIF. One shit prompt later, it had.

That sent us down the model rabbit hole. Could the just-released GLM-5.3 do that? Does Muse actually understand video and audio? Which of the latest models can?

The more I looked, the less this seemed to be about one GIF or another model comparison. Some model families are already being built for people whose default language is increasingly not text. Others are still waiting for those people to write everything down.

Text is becoming legacy media

That sounds dramatic until you look at what happened to the internet.

Blogs became posts. Posts became photos. Photos became Stories, TikToks, Shorts, livestreams and endless voice notes. People still read and write, obviously, but an enormous amount of culture now arrives through a camera and a microphone.

Writing was always the invented bit. Humans see, hear, speak, point and copy each other long before anybody teaches us an alphabet. Text is an incredible compression format, but it is still a learned interface layered over how we naturally experience the world.

The chat box asks everyone to perform that conversion themselves. See the problem, understand it, turn it into a decent paragraph, then give the paragraph to the clever machine.

For somebody with dyslexia, limited literacy, a language barrier, a disability that makes typing difficult or simply a deep hatred of phone keyboards, the work starts before the model has done anything.

Video and audio do not make people stupid. They remove a translation step.

A Zoomer’s bug report is a screen recording. A kid wearing smart glasses points at the problem. Somebody’s entire brief is a reference video and “do this to mine”.

That is not a future interface. People already communicate like this. Models that need the world converted into prose first already feel broken to them.

Google, Meta, Qwen and MiniMax are building for those people

Google and Meta are not mainly selling coding models to people staring at terminals. Their models are supposed to end up in phones, glasses, search, social feeds and messaging apps used by billions of normal people.

Of course multimodality matters more to them.

Google’s new Gemini 3.7 Flash takes text, images, video, audio and PDFs directly. It only outputs text, and Google still has separate Live, image and video models, because apparently even the AI apocalypse needs a product family nobody can name without a chart.

But I can hand it the original screen recording. I do not have to turn the thing into a flipbook and transcript first.

Muse Spark 1.1 can inspect visual and audio information, keep those details through a longer workflow and operate tools with them. Meta also says Muse Spark and Muse Image can work together to make animated GIFs. That may be what happened with the Discord icon. I do not particularly care which internal model moved which pixel. The capability arrives ready to use.

Then look at where Meta can put it: Instagram, Facebook, WhatsApp and smart glasses. A model that understands the media people already make there is a much more obvious mass-market product than another clever text box.

Whether I want Zuck’s machine looking through all that personal context is a separate and fairly enormous question.

Qwen may have gone further than either. Alibaba’s Qwen3.5-Omni takes any combination of text, images, audio and video in one request. It can handle up to three hours of audio or an hour of video, understand speech and sound effects, use tools and answer with text or voice. The newer general Qwen3.7 models also take images and video directly, although audio still belongs to the Omni branch.

MiniMax sits somewhere between the two approaches. MiniMax M3 combines frontier coding and agent work with image and video input in the same open-weight model. Speech, music, image and video generation still live across MiniMax’s wider model range rather than one model doing the lot, but it is plainly not waiting for somebody to convert a video into text before it can reason about it.

Some models still need to catch up

The current Grok 4.6 API model takes text and images, while voice and video sit elsewhere in the wider Grok product.

OpenAI says its latest general models also take text and images. ChatGPT has voice, transcription, image generation and other product machinery around them, but GPT-5.6 itself is not where I send a video.

GLM-5.3 just launched and looks great for coding and long-running agent work. It still has no multimedia input, so I cannot really use it for a lot of my stuff without bolting other models around it. Z.ai sends image and video work to GLM-5V-Turbo and speech to a separate transcription model.

Then there is Claude. Anthropic’s current model overview lists text and image input across Fable, Opus and Sonnet. No audio. No video. Bring your own frames and transcript.

Claude is still excellent at coding and agents. Anthropic can make a lot of money owning professional, text-heavy work. It is not about to vanish because a teenager cannot send it a voice note.

But come on. On multimedia, it is mighty behind.

The apocalypse already happened to us

We still compare models as if the job is to give the smartest answer to a written question. Prompt goes in, benchmark answer comes out, everybody argues over three points and a price per million tokens.

Meanwhile, people have reorganised huge parts of their lives around TikTok, Shorts, Reels, video calls, voice messages and cameras that are always within reach. Seeing something and showing it to somebody is more natural than stopping to describe it perfectly.

Text is not going away. I live in the stuff. It is precise, searchable and still the best medium for plenty of serious work.

But it is becoming a specialist medium rather than the universal front door.

That changes what mass adoption looks like. The amazing shit I keep looking for might not begin with everyone learning how to prompt. It might begin when they no longer have to.

The multimedia apocalypse is not coming.

It already happened to the humans. The models are still catching up, and the ones waiting for a transcript had better hurry.