Skip to main content

The end of text hegemony: with Gemini 3, GEO becomes multimodal

Until now, Generative Engine Optimization (GEO) focused essentially on text: structuring semantics, refining named entities, and shaping the rhetoric that convinces LLMs.

Antoine HeftlerCo-founder

The end of text hegemony: with Gemini 3, GEO becomes multimodal

Until now, Generative Engine Optimization (GEO) focused essentially on text: structuring semantics, refining named entities, and shaping the rhetoric that convinces LLMs. That era is over. The arrival of natively multimodal models means AI no longer only "reads" the web. It looks at it and listens to it. For brands, that imposes an immediate redefinition of what counts as a "source" in the eyes of answer engines.

1. The image is no longer an illustration. It is data

For a long time we treated visuals (images, infographics, videos) as "reassurance" for humans, or as decoration for the interface. For an AI such as Gemini 3, those elements become raw data to interpret.

The technical path was already drawn. In September 2025, Mistral AI showed with Pixtral-12B how fast computer vision was advancing: thanks to LoRA adaptation, the model's accuracy on satellite imagery jumped from 56% to 91%. That quantitative leap proves that models no longer settle for aggregating metadata (the alt text). They analyze the substance of the image itself, to spot discrepancies and sort the reliable from the approximate.

If your text claims an expertise that your visuals contradict by being generic or inconsistent, the AI will downgrade your authority score.

2. The notion of "source" expands

AI's mantra remains unchanged: "sources, sources, sources." The perimeter of those sources, though, now extends to your visual and video corpus.

A technical demonstration video or an expert interview is no longer merely "rich content." It becomes proof of authority that the engine indexes in order to validate an answer. A brand that does not map the situations in which its products are used — including visually — will be systematically overtaken.

The risk is dilution. If your visual assets (owned and earned) are not designed as reliable anchor points, the AI will sink them into an uncertain narrative. GEO therefore has to include a visual semantics: do your images carry the same brand attributes as your texts?

3. Toward total narrative coherence

This shift to the multimodal forces companies to break the silos between "brand," "content," and "SEO" teams. AI makes no such distinction. It tries to build what is plausible, sometimes at the expense of what is true.

If the AI perceives a dissonance between the text (the promise) and the image (the proof), it will introduce a "blur" into how the brand is rendered. In fuzzy logic as in GEO, the aim is to maximize the probability of being cited as a reference, to move from 0.55 to 0.92 in the association with a given expertise.

The question is no longer only "what have we written?" It is "what are we giving the model to see and to hear that confirms our authority?" Optimization for generative engines becomes an exercise in total coherence.