Fireplexity

Field guide

That little [1] is
not a receipt.

AI search engines put a number after almost every sentence, and it feels like proof. Usually it means only that the model was handed a page on the same topic. This guide covers what the studies found, how to check any answer in about a minute, and what it takes to build an answer engine that cites honestly.

The short version

What a citation usually means

“This page was in the pile I was given, and it’s about the same thing.” The link almost always works and the page is usually on topic. Whether the page actually says what the sentence claims is a separate question, and in most systems nothing checks it.

What to do about it

Check the one claim you’d act on, not all of them. Open its source, use Find in page to search for the specific number or name, and confirm the page is the original. That takes about a minute and catches most bad citations.

The numbers, briefly

Three independent studies, run in different years with different methods, keep landing on the same shape of result:

60%+ of answers were wrong when eight AI search engines were asked to find the article a news quote came from. Tow Center, 2025
~1 in 4 citations didn’t fully support the sentence they were attached to. About half of all statements weren’t fully backed by any citation. Stanford, 2023
24–77% of deep-research citations passed a fact check against the source, even though 94%+ of the links worked for frontier models. Cited but Not Verified, 2026

That third number shows the whole problem. The link works and the page is on topic, which is everything you can see at a glance. Whether the page actually supports the claim is the part you can’t see without opening it.

Try it: spot the bad citation

Before the theory, a quick test. Each card below shows a sentence an answer engine might write, along with the part of the cited source it was working from. Decide whether the source backs up the sentence, then reveal the answer. These are made-up examples, but each one is a pattern the studies above keep finding.

Answer says

The bridge reopened to traffic on 14 March.[2]

[2] …officials said the bridge is expected to reopen in mid-March, pending a final safety inspection…
Supported?

NoThe source makes a prediction and the answer reports it as a fact, with a specific date the source never gives. This is one of the most common errors: the answer drops the source’s hedges. “Expected to” becomes “did”, and “mid-March” becomes a precise day.

Answer says

The study followed 1,200 adults over ten years.[1]

[1] …the ten-year study, which tracked 1,200 adults aged 40 to 65, found that…
Supported?

YesBoth numbers are in the source and mean the same thing there. This is what a good citation looks like: you can point to the words. It’s also the easy case. Most of the time, engines do get simple facts like this right.

Answer says

Remote workers are 13% more productive than office workers.[4]

[4] (a productivity blog) …a well-known experiment found a 13% performance increase among call-centre staff who worked from home…
Supported?

PartlyThe number is right but the scope isn’t. The study covered call-centre staff at one company, and the answer applies it to all remote workers. The source is also a blog describing the study, not the study itself. A quick search for “13%” on the page would find a match, which is exactly why this kind of error slips past.

Answer says

The company employs around 4,000 people.[3]

[3] (the company’s About page) …founded in 2009, we build logistics software for mid-sized retailers across Europe and North America…
Supported?

NoThe page is the right one to cite for this company, but it never mentions headcount. The number came from the model’s memory (or nowhere) and the citation was attached because the page is on topic. This is the most common failure in the 2026 study, and the hardest to spot without opening the link.

If you got all four, you already have the instinct. The rest of this guide explains why these errors happen and turns that instinct into a routine.

What the research actually found

Tow Center (Columbia Journalism Review), March 2025

The researchers took 200 news articles from 20 publishers, pasted a direct quote from each into eight AI search tools, and asked a simple question: which article is this from, who published it, when, and what’s the URL? Across 1,600 queries, the tools got it wrong more than 60% of the time.

Perplexity did best, wrong 37% of the time. Grok 3 did worst at 94%, and 154 of its 200 citations led to error pages. The premium tiers were confidently wrong more often than the free ones, because they answered more readily instead of declining. ChatGPT misidentified 134 articles but signalled any uncertainty only 15 times out of 200 and never declined to answer. Many tools also cited syndicated copies on sites like Yahoo News rather than the original publisher.

BBC and European Broadcasting Union, October 2025

Journalists at 22 public broadcasters in 18 countries reviewed more than 3,000 answers from ChatGPT, Copilot, Gemini and Perplexity about the news. 45% had at least one significant issue. Sourcing was the biggest problem at 31%: attributions that were missing, misleading or wrong. And 20% had major accuracy problems, including invented details and out-of-date information. Gemini did worst, with significant issues in 76% of its answers.

Stanford, “Evaluating Verifiability in Generative Search Engines”, 2023

This is the paper that gave us the two measures worth knowing. Citation recall: what share of an answer’s factual statements are fully backed by a cited source? Citation precision: what share of citations actually back the sentence they’re attached to? Across four engines, recall averaged 51.5% and precision 74.5%. Perplexity had the highest recall of the four at 68.7%, with precision of 72.7%. The engines tested are now two generations old, but the measures still hold, and the later studies suggest the gap hasn’t closed.

“Cited but Not Verified”, May 2026

The newest and most revealing of the four. The authors parsed every inline citation in research reports from 14 models and scored each one three ways. Does the link work? 94% or more for frontier models. Is the page on topic? Above 80% for frontier models. Does the page actually support the claim? Between 24.4% and 76.8%, depending on the model.

They also found that more searching made things worse. As agents made more tool calls, from 2 up to 150, fact-check accuracy for GPT-5.4 fell from 79% to 17%, and for Claude Opus 4.6 from 80% to 58%. The link and topic scores stayed above 92% the whole time. More sources meant more citations that looked right and fewer that were.

Reading these numbers fairly

The four studies measure different things. Tow tested one hard task (finding where a quote came from), the BBC study looked only at news, Stanford’s engines are from 2023, and the 2026 paper studied long research reports rather than quick answers. None of them gives “the” accuracy rate of AI search. What they share is the pattern: the citation looks more trustworthy than it is, every time, with every engine.

Why it happens: the anatomy of a [1]

It’s easier to stop trusting the footnote once you’ve seen how one gets made. Fireplexity is a good example because its whole pipeline is open source and short enough to read in one sitting. This is what it does with your question:

SearchFirecrawl returns up to six web results and scrapes each page to markdown.
TrimEach page is cut to about 2,000 characters: its opening, the three paragraphs that best match your keywords, and its ending.
NumberThe excerpts are labelled [1] to [6] and handed to the model along with your question.
WriteThe prompt says to “include citations inline as [1], [2]”. The model writes, adding numbers as it goes.

Three things follow from this, and they apply to most answer engines, not only this one.

  • The citation is a formatting instruction, not a check. The model places [3] the same way it places every other word, by predicting what fits next. Nothing afterwards confirms that source 3 contains the claim. This is the “cited but not verified” gap in practice.
  • What the model read and what you click are different. The model saw about 2,000 characters of the page, chosen by keyword matching. The citation links to the whole page. So a claim can be cited to a page where the supporting paragraph was cut before the model ever saw it, so the claim must have come from the model’s memory, or from nowhere. It can also go the other way: you find the sentence on the page and assume the model read it, when it didn’t.
  • The model is never told it may say “the sources don’t say”. Asked a question its six excerpts don’t answer, it does what the Tow study saw again and again: answers anyway, fluently, with numbers attached.

None of this makes Fireplexity unusually bad. It’s the standard design, and the studies above show Perplexity’s answers failing in exactly the same ways. The difference is that here you can read the code that does it, which also means you can change it. More on that below.

The 60-second check

You don’t need to check every citation. You need to check the one you’re about to rely on. This routine catches every kind of error in the quiz above:

  1. Pick the claim that matters.The number you’ll quote, the date you’ll plan around, the fact you’ll repeat to someone. Skip the background sentences.
  2. Search the page for the specific detail, not the topic.Open the source and use Find in page (Ctrl+F or ⌘F) for the number, the name or the date. Searching for the topic will always find something; searching for “4,000” won’t if it isn’t there.
  3. Read the sentence around it.This catches the subtler errors from the quiz: “expected to” turned into “did”, “call-centre staff” turned into “remote workers”. It also catches old figures presented as current ones.
  4. Check who published it.Is this the original publisher, or a copy, aggregator or blog writing about it? If the claim is important, go one step further back to the original.
  5. Notice what has no citation.Sentences without a number are where the model is speaking from its own memory. They may be right, but treat them as unsourced.
  6. If you can’t find it, ask for the quote.Ask the engine to “quote the exact sentence from source [3] that supports this”. Then search the page for that quote. Models can invent quotes too, but an invented quote is easy to catch with Find in page, and an invented paraphrase isn’t.

Step 6 is the most useful trick in this guide. A quote can be checked by an exact search, but a paraphrase can’t. Whenever a citation matters, get the engine to commit to exact words.

If you’re building one: measure first

If you run an answer engine of your own, whether that’s Fireplexity, Vane or something you built yourself, you can’t improve citation quality until you measure it. The Stanford measures turn into a weekend project:

  • Collect 50 real questions from your users or your domain. Save every answer along with the exact excerpts the model was shown, not just the URLs.
  • Split each answer into individual factual claims and record which source, if any, each one cites.
  • Recall = claims backed by their cited excerpt ÷ all factual claims.
  • Precision = citations whose excerpt backs the claim ÷ all citations.
  • Rerun the same 50 questions after every prompt or retrieval change. Without this, you’re judging a change by reading a few answers and guessing.

You don’t have to do the claim checking by hand. MiniCheck, from an EMNLP 2024 paper, is a family of small models trained for exactly this task: given a claim and a passage, does the passage support it? Its best version has 770M parameters and matches GPT-4 on the authors’ benchmark at roughly 400 times lower cost. Vectara’s HHEM does a similar job. Checker models have blind spots too: on FaithBench, a benchmark of deliberately tricky cases, the open HHEM-2.1 model scores only about 67% balanced accuracy. So use them to flag answers for review rather than to approve them, and read a sample yourself each week.

If you’re building one: four fixes

In order of effort, cheapest first. All four fit into a codebase as small as Fireplexity’s.

1. Let the model say “the sources don’t say”

Add one line to the system prompt: answer only from the numbered sources, and if they don’t cover something, say so instead of filling the gap. It costs nothing and directly addresses the confident-but-wrong pattern the Tow study found. Expect more answers to come back with gaps. That’s the point.

2. Make every citation carry a quote, then check the quote

Ask the model to write citations as "exact words" [n] and check each quote against the excerpt it was shown. It’s a string search, so it’s deterministic, nearly free, and catches invented support without needing a second model:

// Each citation must carry a quote found in
// the excerpt the model saw, not the full page.
const CITE = /"([^"]{12,300})"\s*\[(\d+)\]/g

function checkQuotes(answer: string, excerpts: string[]) {
  const norm = (s: string) => s.toLowerCase().replace(/\s+/g, ' ').trim()
  return [...answer.matchAll(CITE)].map(([, quote, n]) => ({
    source: Number(n),
    quote,
    found: norm(excerpts[Number(n) - 1] ?? '').includes(norm(quote)),
  }))
}

Anything with found: false gets dropped, flagged in the interface, or sent back to the model for another try. Check against the excerpt, not the full page. Matching against the full page reintroduces the gap described above.

3. Link to the sentence, not the page

Every major browser now supports text fragments: a link ending in #:~:text=exact%20words scrolls to that phrase and highlights it. Build citation links from the verified quote and every click lands on the evidence. A reader who sees the highlighted sentence trusts the answer for a good reason. A reader whose link lands on nothing has learned something useful too.

4. Run a checker on what’s left

Pass every sentence and its cited excerpt through MiniCheck or HHEM before the answer streams out, or right after it, and mark the sentences that fail. This adds some latency and one more model to run, but it’s the only one of the four fixes that catches the “right number, wrong scope” errors that a quote check misses.

None of these fixes is possible from outside a closed product. That’s the real argument for an open answer engine, and it’s a better argument than cost, as we found when we compared Fireplexity with Perplexity.

Common questions

Are Perplexity’s citations accurate?

More often than most rivals, but not reliably. It had the lowest error rate in the Tow Center’s 2025 test, at 37%. In Stanford’s 2023 study, 72.7% of its citations fully supported their sentence, so roughly one in four didn’t. Treat a Perplexity citation as a lead worth checking, not as proof.

Why do AI search engines cite sources that don’t say what the answer says?

Because in most systems nothing checks. The model is told to add [1], [2] after its sentences and does so the way it writes everything else, by predicting what fits. The source gets cited because it’s on topic, not because the claim is in it. And the model often sees only a trimmed excerpt, while the link goes to the whole page.

How do I check if an AI answer is grounded?

Open the source behind the one claim you care about and use Find in page to search for the specific number, name or date. Read the sentence around it, check that the page is the original publisher, and treat uncited sentences as unsourced. If you can’t find the claim, ask the engine for the exact quote and search for that.

Which AI search engine has the most accurate citations?

None is accurate enough to skip checking, and rankings change with every model release. Perplexity led the Tow Center’s 2025 test, and Gemini trailed in the 2025 BBC and EBU study with issues in 76% of answers. In 2026 research on deep-research agents, models ranged from 24.4% to 76.8% on a fact check of their own citations.

Can AI citations be checked automatically?

Partly. Checker models like MiniCheck judge whether a passage supports a claim, at GPT-4 level accuracy for far less cost. A simpler check is to require exact quotes and confirm by string search that they appear in the source. Both are worth running, and neither replaces reading a sample yourself.

Sources

  1. AI Search Has a Citation Problem — Jaźwińska and Chandrasekar, Tow Center for Digital Journalism, Columbia Journalism Review, March 2025
  2. News Integrity in AI Assistants — EBU and BBC, October 2025, with a summary at infoDOCKET
  3. Evaluating Verifiability in Generative Search Engines — Liu, Zhang and Liang, Findings of EMNLP 2023
  4. Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — Onweller et al., May 2026
  5. MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents — Tang, Laban and Durrett, EMNLP 2024
  6. Vectara’s Hallucination Leaderboard and Benchmarking LLM Faithfulness in RAG — HHEM and FaithBench results
  7. Fireplexity search route and content selection — the retrieval, trimming and citation prompt described above
  8. Text fragments — MDN reference for #:~:text= links

The quiz examples are illustrative, not quotes from real answers. Study figures are as reported by their authors and measure different tasks, so they shouldn’t be compared directly as engine rankings. This site is an independent community page and is not affiliated with Perplexity AI or Firecrawl.

Read the pipeline.
Then fix the footnotes.

Fireplexity is MIT licensed, a few hundred lines long, and runs locally in a few minutes with two API keys.