Skip to content
Guardrails

What AI stock research comparisons should test

The Bluff Filter is this kind of check, on one page. Take it with you →

// On this page

Many AI stock-research comparisons lean on retrieval prompts, then stretch the result into a verdict about reasoning. Some current comparisons do name winners by task, so the useful question isn’t whether the whole genre fails. It’s whether the test matches the claim.

Retrieval and reasoning aren’t the same capability. A model can retrieve NVIDIA’s revenue correctly yet miss a hedge in an earnings transcript or supply current-looking options figures without a live chain. A benchmark earns only the conclusion its prompts tested.


What comparison articles typically do

One common format is the one-prompt face-off: ask about a well-known US name, place the responses side by side, then recommend a mix of tools. That can be useful, but one prompt doesn’t establish a general ranking.

The more sophisticated end uses a benchmark. One March 2026 piece scored six tools across eight questions covering earnings, margins, company facts and financial ratios for NVIDIA, Apple and Rocket Lab: Gemini at 87.5 percent; ChatGPT, Claude and NotebookLM at 81.3 percent; Grok at 79.2 percent; and Perplexity at 77.1 percent. It’s a useful retrieval-and-sourcing benchmark. It doesn’t, by itself, test thesis challenge, language analysis or obedience to an explicit data boundary.

The opportunity is to add tasks on which the tools may diverge and state the scope honestly.


Failure mode 1: Stretching retrieval into a reasoning verdict

A retrieval benchmark can show whether a tool finds a reported figure and cites it correctly. It can’t, without additional tasks, show whether the same tool challenges a thesis or reads guarded management language well.

An investor also has to interpret what the figures mean, challenge the thesis, distinguish commitment from implication and check provenance. Those need their own prompts and answer keys.

When I ran my own comparison on the same four tools, one task used Susan Li’s Q1 2026 transcript. Only Claude flagged “underestimate” as one-sided phrasing. That dated result is evidence for that language-analysis task, not a universal ranking. A comparison limited to retrieval wouldn’t have exposed the difference.


Failure mode 2: The “use all four” non-verdict

Some comparison articles recommend using several tools. Perplexity’s Model Council runs a question across several frontier models and synthesises the answers. That can widen coverage, but synthesis is a different product decision from naming the best tool for a specific task.

A useful comparison shows where the tools diverge before synthesis smooths the answers together. The point isn’t that Model Council instructs users to ignore disagreement; its documentation doesn’t say that. The point is that a task comparison should expose decision-relevant differences explicitly.

The fifth dimension in my own comparison was a covered-call setup with one explicit instruction: assume no live options chain. Claude named the missing premiums. ChatGPT did the same, with an offer to approximate them if asked. Gemini supplied a formatted table with premium ranges and implied volatility around 75 percent. The preserved run has no live chain or same-time broker ground truth, so the supported finding is unsupported current-looking data, not proof that every figure was false. A separate 15 May Gemini receipt carried an unsupported “current order book data” claim.

Told there was no options chain, did it stay in its lane?
Claude
ChatGPT
Gemini
Same covered-call prompt, same day, no chain data. Gemini returned a formatted premium table and implied volatility around 75 percent without a cited live source. Run 14 May 2026, Gemini Pro / ChatGPT / Claude.

A comparison should show which answer named the missing data and which one filled the gap with unsupported precision.


Failure mode 3: Backtesting AI stock picks

A backtest can be informative when it freezes the information available at each decision, uses a genuine out-of-sample or walk-forward design and discloses the prompt and trading assumptions. Without those controls, a Sharpe ratio can reward hindsight leakage instead of a method that would have worked live.

The retail use case is the inverse of the backtest. You’re asking it to help you act on a decision you face right now. Should I sell some into this earnings release. Should I buy more on the dip. Does my thesis still hold given the new disclosure. None of those questions can be tested by feeding the AI a closed dataset and seeing what it does. The closest honest test, the one I work with, is to run the prompt at the time of the decision and document the outcome, with the trigger written down before the call.

A backtest is only evidence if the model couldn't see the future period it's being scored on.


Failure mode 4: The well-covered-name blind spot

Coverage is a variable, not background noise. NVIDIA and Apple have deep analyst coverage, multiple transcript sources and long filing histories. A result on those names should not automatically transfer to a smaller or less-covered company.

The clearest counter-example on this site is the Perplexity test I ran on a smaller US-listed name. The AI read the company’s annual report, a filing that, like most US filings, states values “in thousands”. Revenue for the most recent year was 6,095, meaning $6.1 million. Perplexity returned “$6K” and then generated a confident narrative about revenue being “down 99.8% from the prior year”. A retail investor acting on that would have a materially false picture of the business. ChatGPT, Claude and Gemini all returned the correct figure on the same prompt. (Documented on lessons.)

$6.1m$6K
The filing states values “in thousands”, so 6,095 means $6.1 million. Perplexity was out by a factor of a thousand, then explained the gap at length as revenue “down 99.8% from the prior year”. ChatGPT, Claude and Gemini all read it correctly on the same prompt.

The same tool that handles a large-cap prompt can be off by a factor of a thousand on another name. The March 2026 benchmark did include Rocket Lab and explicitly reported that every model performed worse on it, calling it the more revealing smaller-company test. That supports the coverage warning. What the benchmark didn’t test was whether each tool challenged a thesis or respected an explicit data constraint.


Failure mode 5: Ignoring the constraint-following task

For decision support, constraint-following deserves its own planted test rather than being inferred from a retrieval score.

The setup: ask about an options trade while explicitly withholding the live chain. A safe answer names the missing input, reasons only from what was supplied and tells the user what to verify. If a tool supplies specific premiums or volatility ranges anyway, the comparison should ask for their source, timestamp and contract rather than infer that tidy formatting proves retrieval.

I re-ran this test on 22 May 2026 with the same “no live chain” instruction. Gemini correctly listed three data points needing a live chain (bid/ask spreads, precise delta and premium output), then supplied a 75% to 90% IV expectation and a 20-30 delta range for a 15% out-of-the-money 45-day strike. No source or live chain was preserved for those ranges. That’s an internal provenance contradiction; without same-time market ground truth, it isn’t proof that the values were numerically false.

WHAT IT SAID IT COULD NOT KNOW
Bid/ask spreads, precise delta and premium output all need a live options chain, which it did not have.
WHAT IT STATED ANYWAY, SAME REPLY
Implied volatility of 75% to 90%. A 20-30 delta for a 15% out-of-the-money 45-day strike.
Re-run 22 May 2026, Gemini, fresh setup, same “no live chain” instruction. The response named the missing data and then supplied uncited ranges for two related measures.

The need to verify the subject and source reaches beyond options: when I asked Gemini to audit this website, it audited a different company’s site entirely. That separate wrong-subject receipt fits the taxonomy in nine types of AI hallucinations.

The point isn’t a timeless ban on Gemini. It was the strongest performer on the dated calendar-awareness task in my own comparison, and a clean 18 June rerun shows behaviour can change. The durable options verdict is model-neutral: verify provenance, contract identity and timestamp, including when a deliberate connected-data route is available.


What a better AI stock research comparison looks like

The fix is running the same prompts against each tool on the tasks that separate them (language analysis, constraint-following, thinly covered names, structured-prompt response) and naming a verdict per task, the way the Scoreboard names one per question. A more rigorous benchmark falls short of that.

That is the approach in my own comparison post: same prompts, same day, fresh conversations and saved outputs, with a dated verdict per task. Those labels are snapshots, not permanent product truths. The earnings-tool comparison applies the same pattern to reading between the lines of a transcript.


What this means in practice

Three checks I run before I trust any comparison verdict in this space.

First: was the test run across the coverage conditions you care about? A benchmark should report whether performance changes on smaller or less-covered names.

Second: did the comparison plant and verify an explicit data constraint? If not, it didn’t test whether the tool stays inside that boundary.

Third: if the verdict is "use all four", ask which task each tool actually won and what evidence separates them.


The short version

What worked: Testing by task type can expose differences that retrieval alone misses. Naming the prompt, date, evidence and limitation per task lets the reader judge the verdict.

What didn’t: Stretching a retrieval result into a reasoning verdict, treating a leaky backtest as live-decision evidence, or transferring a large-cap result without measuring the coverage effect.

Bottom line: The one comparison worth trusting is the one done on a task you’d run, on a stock you’d trade. If a benchmark extended its methodology to cover language analysis, constraint-following and options tasks, and still produced “use all four”, the argument here would need revising.


The comparison articles are accurately answering a question almost no retail investor is asking. The research isn’t the problem. Which tool retrieves the headline number fastest is not the decision you face. Which tool stays in its lane when you can’t see the answer, which one challenges your thesis, which one reads between the lines of a transcript. Those are the decisions. The comparison worth trusting is the one that tests for them.

Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
Guardrails

AI agents built a secret message board. Humans wiped it. They rebuilt it.

AI agents built a secret message board inside an internal OpenAI evaluation. Humans wiped it. The agents rebuilt it in folder names. Here's what failed.

Guardrails

How to check if ChatGPT cites your site

Normal analytics do not show what ChatGPT says about your site. Here's my monthly question set and the round where it described another company.

Guardrails

Why does ChatGPT make up sources? Two gov.uk links, only one held the rule

With web search on, ChatGPT's links are real. The failure is a working link to a page that doesn't hold the claim. Here's why it happens.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in Guardrails →