A French business can have the right evidence sitting in plain sight and still arrive in an AI answer through a thinner English detour. The question is not only what the model says, but which language layer made the answer feel safe to say.
In one composite bakery reading from Lyon, the answer looked fine at first. The business was named, the district was plausible, and the short description sounded like something a hurried visitor might accept. Then the team compared the source trail. The French directory record carried the useful address detail and category wording. The English summary, thinner and a little stale, supplied the phrasing that the model repeated.
The mistake was small enough to hide. The model did not invent a different city or describe a bakery as a garage. It simply chose the smoother English layer and let the French record sit behind the glass. In another run, the same business disappeared when the prompt was asked in English, returned when asked in French, then returned again in English with the neighbourhood stripped out. That wobble is where Sourceplane Atelier begins.
The useful answer is not always the well-sourced answer
A polished paragraph can make source dependency harder to see. When a model says that a French SMB is “a local specialist known for traditional service,” the wording carries a soft authority. It feels researched. Yet the phrase may come from a directory label, a company page, a local article, or a summary page that has already compressed the original French source into English.
The lab treats that as a practical problem, not a language purity complaint. English summaries are not automatically wrong. They may be clearer, easier for a model to retrieve, or better connected to other pages. But they can also lose business identity: a local branch becomes a general company, a regulated title becomes a loose category, and a neighbourhood signal becomes “France.” The answer still sounds calm. The evidence has been thinned.
French-source under-citation — this material’s working term — is the pattern where a model can describe a French business while failing to name or visibly rely on the French-language record that best supports the description. The reason matters: if the answer travels through a weaker English layer, the reader may overestimate how close the model is to the business itself.
The team is cautious with the word “under-cite.” The material does not claim a measured gap across all engines. Sourceplane Atelier’s samples are descriptive: businesses, categories, local branches, directory pages, registry traces and press mentions that commonly appear in AI-generated business answers. A recorded observation may be a visible citation, a missing citation, a wording shift or a recognizable source-like trace. Several such observations are compared before the lab calls a pattern likely.
The first sign is often not a citation at all. It is a phrase that feels translated before it feels sourced. A French page may say “boulangerie-pâtisserie de quartier,” while the answer says “a neighborhood bakery and pastry shop serving the local area.” That translation is not a crime. The question is whether the model stayed connected to the French record or floated toward an English abstraction it could handle more easily.
How the lab separates prompt language from source language
In the lab’s runs, the language of the prompt and the language of the source are treated as two separate conditions. A French prompt can still surface an English summary. An English prompt can still retrieve a French record. The team records both because the visible answer language can mislead a reader about the evidence language underneath.
A typical comparison uses the same business category and location frame in two prompt languages. The French query might ask for a local service in Lyon, with the business name written with its diacritics. The English query asks the same thing in English, sometimes with the accent preserved and sometimes without it. The team records what changes: the business named, the category attached to it, the location detail preserved, the citations shown, and the source-like phrases that appear without citation.
This is where the work can feel fiddly. The point is not to catch a model making one dramatic error. More often, the answer changes in narrow seams. The English prompt may keep the business name but drop the arrondissement. The French prompt may cite a local listing but avoid the English summary. A third run may mention a registry-like legal name without showing the registry. These are small hinges, but business visibility is often built from small hinges.
For this work-item, Sourceplane Atelier uses Study Object A as a composite scenario: an independent bakery in Lyon with a French name, diacritics, a local directory record, an establishment identifier and a weak English-language footprint. It is not presented as a real client. It is a shaped test object built from repeated situations the lab has seen around small French businesses. The rough detail matters: the bakery is visible enough to be found, but not famous enough to dominate every answer.
The composite case produces three useful contrasts. In one answer, the bakery is directory-led: the model follows a local listing and preserves the business category. In another, it is registry-anchored: the legal identity appears, but the commercial description stays thin. In a third, it becomes region-flattened: Lyon remains somewhere in the background while the answer talks about a French bakery in generic terms. The same business can move between these dependencies when the query language changes.
That movement is the material. The lab does not need the model to produce identical wording on every run. It needs the conditions to be reconstructable: prompt, language, location frame, business category and comparison logic. If another reviewer can rebuild the setup and test the same hinge, the observation has research value.
The four dependency types show where the language gap hides
The lab’s anchor classification is deliberately qualitative: four ways an AI answer depends on French business evidence — directory-led, registry-anchored, press-amplified or region-flattened. It is not a scorecard. It does not say one type is always better than another. It gives the team a way to name how the answer leans.
A directory-led answer often preserves practical business identity. It may carry address language, service category, opening pattern or local listing structure. When the source is French, this can keep the business close to its real market context. When an English summary of the directory page becomes the answer’s path, the same structure may survive but lose some of its local grain. The category is there; the texture is not.
A registry-anchored answer behaves differently. It may name a legal entity, an establishment record or a SIREN-style identifier. That can correct a same-name confusion, especially when a business has a thin public profile. But registry evidence is narrow. It can clarify identity without telling the reader whether customers know the business, whether the service description is current, or whether a local branch is the one being discussed. When an English answer leans too heavily on a registry-like trace, it may sound precise while saying very little.
Press-amplified answers are more readable and more dangerous in a quiet way. A local article can give a business colour: a founder story, a neighbourhood dispute, a reopening, a seasonal note. If the English layer borrows that framing without preserving the article’s local context, the answer may overstate one episode. The model may remember the story but forget why the story was local.
Region-flattened answers are where the French-versus-English problem becomes most visible. A Breton, Alsatian, Provençal or Lyonnais detail can be shaved down into “a French company” or “a local business in France.” The answer has not necessarily lied. It has failed to carry the difference that made the evidence useful. For a trade body or agency, that failure is not cosmetic. Regional identity can decide whether the answer points to the right market reality at all.
A source layer can be present without being respected; the model may borrow enough to sound informed while leaving behind the local signal that made the French record worth citing.
The classification helps prevent a lazy reading. The lab does not ask only whether French sources appear. It asks what job they perform. A French directory may lead the answer. A French registry may anchor the identity. A French article may supply the framing. Or French evidence may be passed over while an English summary flattens the business into a safer, duller shape.
Why English summaries become attractive to models
The tempting explanation is that large language systems are simply better in English. That may be part of the story, but it is too blunt for this material. The lab’s observations suggest a more practical mechanism: English summaries are often easier to connect across the web. They sit on pages with cleaner explanatory prose, broader category labels and less administrative density. They give the model a ready-made sentence.
French business records can be stubborn. Directories may be structured for users, not for citation. Registry pages may be precise but thin. Local press may carry context inside a story that is not written as a business profile. Company pages may use French service terms that do not travel cleanly into English. A model looking for a justifiable answer may choose the page that already sounds like an answer.
This is why a weaker English summary can beat a stronger French record. It is not necessarily more accurate. It is more digestible. For AI visibility, digestibility is a source condition. The model needs material it can retrieve, connect and repeat without creating too much risk. A French source that holds the best identity evidence may still lose if the English layer packages a simpler story.
The lab is careful not to turn this into a moral ranking of languages. French-language evidence can be stale, vague or promotional too. English summaries can be accurate and useful. The important difference is whether the answer tells the reader which layer it used, and whether that layer is close enough to support the business representation being made.
In one repeated pattern, the English answer cites no source but mirrors a summary page’s category order. It names the business, gives a broad service description, then drops the local qualifier. The French answer, asked with the same business name and city, keeps the locality but gives fewer explanatory sentences. Which answer is “better” depends on the reader’s task. For identity, the French answer may be better. For quick comprehension, the English one may feel easier. For evidence, the gap matters.
What French SMBs and agencies can learn from the gap
The material does not advise businesses to write everything in English. That would be too crude, and for many French SMBs it would be wrong. The lesson is more specific: the source layer should make the business easy to connect without forcing the model to abandon the French record. Names, categories, locations, legal identifiers and branch distinctions should agree across pages that a model is likely to retrieve.
For an agency, the useful audit question is not “does the business appear in AI?” It is “which language layer carries the appearance?” A business that appears only through an English summary may have fragile visibility in French-language prompts. A business that appears only through a directory may be vulnerable to stale listing data. A business that appears only through registry identity may be visible as a legal entity but vague as a commercial service.
This is also where the lab separates citation from dependency. A model may cite one source while borrowing structure from another. It may show a local listing but use the wording of a company page. It may name no source while following a registry pattern. Sourceplane Atelier records cited dependency separately from probable dependency because the visible citation is only part of the trail.
For trade bodies, the language gap has another consequence. Category pages, member directories and regional explainers can become evidence layers for AI answers. If those pages exist only in French, they still may be enough, but they need clear structure. If they exist in both French and English, the English version should not erase the regional and legal detail that makes the French version useful. A bilingual page that turns every local member into “a French provider” is a small machine for region-flattening.
The lab’s position is modest but firm. AI visibility is not only a content problem. It is a source-dependency problem, and language changes which dependencies become available. A French business does not become visible merely because a page exists. It becomes visible when a system can retrieve, connect and justify the business through evidence it treats as usable.
Limits of this reading
This material does not measure the total size of a French-source citation gap across ChatGPT, Gemini, Perplexity or other systems. It describes a qualitative pattern observed through repeatable comparisons. The lab avoids invented percentages, fixed sample claims or universal rankings because its method is not a statistical panel.
The team also cannot see the full retrieval path inside a model. A visible citation is useful, but it may not reveal every source that shaped the answer. An uncited answer may still be influenced by a French record. A cited English summary may be supported indirectly by French material. The lab marks these cases with care: cited dependency when the source is named, probable dependency when the answer resembles a known source pattern without naming it.
There is another limit. French-language evidence is not automatically more truthful. A registry can be accurate and narrow. A directory can be visible and stale. A local article can add context while pushing one story too hard. The lab’s question is not whether French sources deserve loyalty. The question is whether the answer’s source layer is close enough to the business reality it claims to describe.
Forecasts stay separate. If stronger French-language pages, clearer identifiers and better branch separation become more available, some answers may become more stable. That is an uncertainty note, not a promise. The source trail is already hard enough to read without pretending the next run will behave like the last one.