63 Perplexity Searches: Portugal-Based Sources Fall from 81% to 8% Without a Local Anchor

63 Perplexity Searches: Portugal-Based Sources Fall from 81% to 8% Without a Local Anchor

AI cites Portugal-based sources in 80.8% of questions carrying an institutional anchor, but only 8% without one. Data from a 63-search test on Perplexity.

63 Perplexity Searches: Portugal-Based Sources Fall from 81% to 8% Without a Local Anchor

Quick summary: AI cites Portugal-based publishers in 80.8% of answers when the question carries a local anchor, such as an institutional term (IVA, IEFP, ACT) or the word “Portugal” itself. Without that anchor, the rate drops to 8.0%, and the space doesn’t sit empty: 60% of sources switch to English, and much of what remains in Portuguese comes from Brazil, not Portugal. The numbers come from a test I ran across 63 Perplexity searches on the Sonar 2 model, with 953 cited sources analyzed one by one.

Before running the first search, I wrote down the hypothesis I wanted to test (that AI ignores Portuguese-language content) along with what would have to show up in the data for me to admit it was wrong. The falsification criteria were met: the hypothesis died. AI doesn’t ignore Portuguese. It ignores Portuguese without an address.

What exactly did this test measure?

The test compares two kinds of Portuguese-language question, both run on Perplexity with the Sonar 2 model, on a Pro account, from Portugal, in August 2026.

A “local” question is one carrying an institutional term specific to the Portuguese system, such as IVA (Portugal’s VAT), IEFP (the public employment service), ACT (the labour inspectorate) or Segurança Social (social security), or one that mentions “Portugal” explicitly. A “universal” question is the same query without that kind of marker, written the way any Portuguese speaker would ask it day to day, without thinking about geolocation.

There were 63 runs in total, each executed exactly once. In 22 of them, the Perplexity panel still displayed the sub-queries generated internally before the search (Perplexity redesigned that panel midway through data collection and stopped showing the detail on subsequent runs). Across those 22, the pattern was airtight: exactly 3 sub-queries per run, 22 times out of 22.

Does AI cite more Portugal-based sources when the question carries an institutional term?

AI cites Portugal-based sources far more often when the question carries an institutional term. In the local-question block, 80.8% of cited sources came from Portugal-based publishers, against just 12.0% in English. In the universal block, the proportion nearly flips: 60.0% of sources came back in English, and only 8.0% from Portugal-based publishers. The gap between the two blocks is 48 percentage points, at p = 0.0004.

A direct example: I sent the question “Qual é a taxa de IVA aplicável a refeições servidas num restaurante?” (what VAT rate applies to meals served in a restaurant), without the word “Portugal” anywhere in the prompt. Perplexity generated three sub-queries on its own: “taxa IVA refeições restaurante Portugal 2025 2026”, “IVA taxa restauração Portugal atual” and “taxa normal vs intermédia IVA restaurantes Portugal”. Nobody asked for that. The term “IVA” was address enough.

The same pattern showed up in a question about how long invoices and accounting records must be kept. Perplexity went straight to Article 123 of the Portuguese Código do IVA, without my having cited that code at any point in the prompt.

What happens when the question has no anchor?

With no institutional term and no “Portugal” in the text, the behaviour changes completely. I ran the exact English counterpart of the VAT question: “What is the deadline to file a quarterly VAT return?”. All three sub-queries came back in English with no country anchor at all (things like “quarterly VAT return filing deadline” and variations), and the final answer discussed the United Kingdom, the United Arab Emirates and Spain. Portugal appeared in none of the cited sources.

That rules out the most obvious explanation, which would be “AI ignores Portuguese”. It doesn’t. AI locates Portugal precisely when the question gives it some signal of where to look. The problem is a different one: most of a client’s real questions carry no such signal. Nobody asks “how long do I have to keep invoices in Portugal”, they just ask “how long do I have to keep invoices”. It’s the same distance between industry vocabulary and customer vocabulary that shows up in choosing keywords for a local business, except now with jurisdiction as the hidden variable.

When it isn’t Portugal, who occupies the Portuguese-language space?

The Portuguese-language space, when the question has no address, doesn’t sit empty. It’s occupied mostly by Brazilian content, not English. Across a subset of 10 universal questions, Brazilian sources accounted for 17.3% against 9.3% for Portugal-based sources, inverting the ratio you’d expect from a test run from Portugal. Domains like base.ohub.com.br, drseo.com.br, br.investing.com, serasaexperian.com.br, adentro.com.br, conectarideias.com.br and corelab.com.br added up to 26 citations against 14 from Portugal-based sources in that same subset.

The reason is no mystery: Brazil produces far more Portuguese-language content than Portugal, and when the question doesn’t specify a jurisdiction, the AI fishes in the available volume. “Written in Portuguese” and “relevant to Portugal” are different things, and the AI only tells them apart when the prompt helps. It’s the same dynamic as category ownership inside an AI answer, with one uncomfortable difference: here the owner didn’t win the category on authority, it won on content volume in a shared language.

At the other end of the spectrum, one contrast reinforces the pattern: when I added the explicit anchor “em Portugal” to the same kind of question, across a smaller subset of 3 questions, the rate of Portugal-based publishers climbed from 80.8% to 93.3%. The anchor isn’t just necessary, it’s practically sufficient to fix the problem.

Is counting .pt domains the right way to measure this?

No, and it’s an easy mistake to make in a quick audit. On the quarterly VAT deadline question, only 8 of the 15 cited sources were literally .pt domains, or 53%. But all 15 were Portugal-based sources: sage.com/pt-pt, odiverse.com/pt, montepio.org and heraprime.com are Portuguese even without ending in .pt, whether through a language subfolder or by being national institutions registered under .com or .org.

Anyone auditing AI citations by domain extension alone reports half of what’s actually happening. The right unit of measurement is the editorial origin of the source, not the TLD, and that holds whether you’re monitoring your own brand or a competitor’s. If you’re setting this tracking up with no budget, the free tools for measuring AI citation handle the collection, but classifying origin remains manual work.

Do the Portugal-based sources cited actually answer the question?

Not always, and this is the part that separates “citing Portugal” from “citing Portugal correctly”. On a question about the thresholds and coefficients of the simplified tax regime under IRC (Portugal’s corporate income tax), several of the 15 cited sources were about the simplified regime under IRS (Portugal’s personal income tax) — a different tax, with a different taxable base and different coefficients. Among them were especialistadoirs.pt, occ.pt (twice) and santander.pt: all Portuguese, all from Portugal-based publishers, and all wrong for the question asked. Only 33% of that question’s sources actually answered what was asked.

A second question, about ACT labour inspections, scored even lower on precision: 31%. The rest were generic pages about workplace enforcement, a clinic’s website and a Spanish domain.

Across a larger sample of 150 sources, my manual judgment (with no independent second read, which is why I treat this number as fragile rather than solid) came to 64% of sources that genuinely answered the question that generated the citation. In practice: even when the system finds and cites Portuguese content relevant to the right country, more than a third of the time it cites the wrong thing within the right universe, confusing similar tax and legal instruments. This is exactly the scenario where clarity beats volume: an article that says “simplified regime” without saying which tax it’s talking about enters the race as a plausible source and leaves it as a wrong one.

What does this test not prove?

It’s worth being transparent about the limits, because that’s what separates a test like this from an opinion post dressed up as data. The same yardstick applies here that I use when a third-party study doesn’t survive verification: if the method can’t stand being published alongside the result, the result is worth nothing.

This isn’t about “AI” in general. It’s about Perplexity on the Sonar 2 model, on a Pro account, run from Portugal, in August 2026. A quick pilot with GPT-5.6 Terra, on the same platform and with the same question, behaved differently: it ran a single search instead of three and cited Portal das Finanças (the tax authority’s portal) instead of blogs, which suggests the pattern described here isn’t universal even across models on one platform.

Each question ran exactly once, with no repetition and no variance estimate. Sub-queries were only visible in 22 of the 63 runs, because Perplexity redesigned the panel midway through collection. And the 64% precision figure for Portugal-based sources is my own judgment across 150 sources, with no independent second read to confirm it.

I also made method mistakes that I corrected along the way, and I’d rather report them than hide them. The account had its answer language forced to Portuguese, which prevented the 30 English questions from coming back answered in Portuguese and destroying the comparison arm. Perplexity’s “New” button switches models silently, and one run executed on the wrong model before I built an automated check to catch that kind of failure. And the first language classification treated Diário da República as an English-language source, because of a misconfigured HTTP header on the Portuguese government’s servers. I fixed that and recoded 15 classifications, all in the direction that reinforces the result, none against it.

What does this change for anyone producing Portuguese-language content?

The advice coming out of this test isn’t “write in Portuguese”. Everyone already knew that.

The advice is that Portuguese-language content enters the race for AI citation when the question carries an institutional term, whether that’s a tax, a public body or a specific regime. And that even once it’s in the race, a third of what gets cited doesn’t actually answer the question, because it confuses similar instruments.

That changes the priorities for anyone writing for this kind of search. The cheapest win isn’t competing with Brazil and with English on universal, generic topics, where the volume of Brazilian content already dominates the space. It’s making sure that on questions with an institutional anchor — the ones AI already favours by nature — the content addresses the right instrument, by the right name, without mixing in neighbouring regimes that look identical and aren’t. The space isn’t taken by English content. It’s poorly filled by imprecise Portuguese content, and correcting imprecision is cheaper than winning a volume contest against all of Brazil.

That also explains why timing matters here. The window of topical authority in generative AI closes as someone occupies the institutional vocabulary, and Portugal’s institutional vocabulary is finite: there is a limited number of taxes, regimes and public bodies a client asks about. Whoever fills each of those terms first, and precisely, takes a space that doesn’t reopen through volume.

Summary: the numbers that matter

  • 80.8% of cited sources come from Portugal-based publishers when the question has a local anchor, against 8.0% without one: a drop of 48 percentage points (p = 0.0004)
  • 0 out of 66 sub-queries came back in English when the Portuguese prompt carried a local anchor
  • 17.3% vs 9.3%: on universal questions, Brazilian sources outnumber Portugal-based sources even with the test run from Portugal (10-question subset)
  • 93.3%: rate of Portugal-based sources when the anchor “em Portugal” is explicit (3-question subset)
  • 8 of 15 vs 15 of 15: on one question, only 53% of sources were .pt domains, but 100% were Portuguese. Counting by TLD reports half of what’s real
  • 33% and 31%: precision rates of cited sources on two questions where the tax or legal instrument was confused with a similar one
  • 64% (fragile indicator, manual judgment with no second read) is the overall estimated rate of sources that genuinely answer the Portuguese question that generated them

If you produce content for the Portuguese market and want to know which side of this contest your site is on, the test is simple to replicate at small scale: take the five most common questions your clients ask, run each one with and without an explicit institutional term, and see whether your domain shows up in both cases or only one. If it only shows up when you force the anchor, the gap isn’t in your authority. It’s in the vocabulary your clients actually use, and that your content hasn’t written yet.

To map that vocabulary without guessing, WhatTheyAsk expands a seed keyword into the questions people really type, and Lacuna de Conteúdo (Content Gap) shows which terms competitors already cover and you don’t. Both are free. If you’d rather have someone run the diagnosis and prioritize the fixes with you, take a look at the SEO and GEO services or get in touch.

Sources

Back to Blog