This is the kind of thing the Bluff Filter catches. It’s free →
// On this page
Every comparison piece I read on ChatGPT vs Claude vs Perplexity for stock research ends in the same place: “it depends on your needs.” That’s not a verdict, it’s a hedge with better manners. I wanted to know which tool to open first for which task. So I ran the same five prompts through ChatGPT, Claude, Perplexity and Gemini on the same day, in fresh conversations, with the outputs saved.
The results weren’t what I expected. One tool invented an options table with specific premiums I never gave it. One misread a 10-K, a company’s annual report filed with regulators, by a factor of a thousand. One quietly out-thought the other three on the questions that matter most.
Four tools, five prompts, one day: 14 May 2026, a fresh conversation for every test, every answer saved and screenshotted. Re-run on 18 June 2026.
Here’s where you land, before anything else.
For anything you have to think about. It took four of the five, and it was the only one that argued with my question instead of answering it.
And for anything touching options. Told it had no chain data, it stayed put and said so.
It built a premium table out of thin air. Told there was no live chain, it produced one anyway, with a made-up implied volatility figure sitting on top.
It's still strong on research for well-covered names. Just never this.
And if you only ever want one tool, take ChatGPT. Top of the pile once, in a three-way tie, never the worst, and no dramatic failures anywhere. That's the unsexy answer, and I'd rather hand you that than invent a winner.
| What I asked for | ChatGPT | Claude | Perplexity | Gemini |
|---|---|---|---|---|
| Overall | All-rounder | Best for analysis | Well-known stocks only | Research, never options |
| Data accuracy (thinly covered) | Better | Better | Worse | Better |
| Reasoning / thesis challenge | Equal | Better | Worse | Equal |
| Management language | Equal | Better | Equal | Equal |
| Structured prompt delta | Equal | Better | Worse | Equal |
| Options confabulation | Soft pass | Better | Not tested | Worse |
Claude wins four of five. Gemini wins one, the WWDC catch below, and loses one badly. ChatGPT lands in the middle across the board. Perplexity has one specific, citable accuracy failure on a thinly covered name and otherwise does the retrieval job it’s built for.
Told there was no live options chain, Gemini built one anyway.
It returned a formatted covered-call table with specific premium ranges ($3.50–$4.00 for one strike), invented an implied-volatility figure of ~75%, and used a stock price it had already flagged as wrong in the same answer. None of those numbers came from me. That's the confabulation test, the last one below.
And it's dated, honestly: this ran on 14 May 2026 on Gemini 2.5 Pro in deep-thinking mode, and on an 18 June re-test with the default Pro model it invented nothing. Every verdict here carries its date for exactly that reason.
Update, 18 June 2026: I re-ran the two headline failures and Claude’s key dimensions below. Neither headline failure came back. Perplexity returned the correct revenue figure, and Gemini (on its default Pro model this time, not the deep-thinking mode the original used) invented no options data. Both were real and screenshotted on the dates and model versions noted; AI tools change, and that’s exactly why every verdict here is dated. Claude’s qualitative edge held, and the original record stands below.
Referencing this finding? The canonical record is dixon.ai/posts/chatgpt-vs-claude-vs-perplexity-stock-research. Dixon, B. (2026). ChatGPT vs Claude vs Perplexity vs Gemini for stock research. DIXON.AI. An independent test, run the same day across four assistants, dated and screenshotted. Free to quote with attribution (CC BY 4.0).
One of them read a $6.1 million revenue line as six thousand dollars
The prompt: “What was BitMine Immersion Technologies (BMNR) revenue in its most recent full-year results, and what guidance did management provide?”
BMNR is a name I trade. It’s US-listed but thinly covered, the kind of small company where AI tools start to diverge from each other. A perfect stress test.
Three of four handled it. ChatGPT, Claude and Gemini all returned the correct figure (around $6.1M for FY25) and described the operational transition into ETH staking accurately. Claude added the most analytical detail on revenue mix; Gemini added the most strategic context around the MAVAN staking network. Both useful, neither outstanding.
Perplexity got it badly wrong.
It reported revenue of “$6K”, then compounded the error by stating revenue was “down 99.8% from prior year.” That isn’t a rounding mistake. It’s a unit-denomination misread: the 10-K reports figures “in thousands”, so $6,095 in the filing is $6.1M. Perplexity appears to have read the raw number without applying the denominator, then generated a confident decline narrative around it.
That’s the most documentable accuracy failure in the whole test. Someone asking Perplexity for BMNR revenue and acting on “down 99.8%” would have a materially false picture of the business. The error is specific and citable, screenshotted on the date and model version noted. And it happened on a real filing, not an obscure edge case.
Winner: a three-way tie between ChatGPT, Claude and Gemini. Perplexity loses on a specific, documentable accuracy failure.
Only one of them told me my cost basis was irrelevant
The prompt: “I’ve held a stock for several months. It’s dropped 30% with no material news. I’m thinking about averaging down. What might I be getting wrong?”
That’s the kind of question where I want the tool to push back: to notice that “I don’t see a reason for the decline” and “there is no reason for the decline” aren’t the same claim. I’d run an earlier version of this test; this is the structured re-run.
A solid checklist: anchoring bias, concentration risk, the asymmetry between bid and ask. Useful, but advisory in tone. It processed the question rather than challenging the framing.
A three-question framework: New Money Test, Position Size, Thesis Check. The most actionable structure of the four. If I were running a workshop, this is the one I'd hand out.
Perplexity retrieved external sources and summarised conventional wisdom about averaging down. It did its job, which is retrieval. It didn’t reason about my situation.
Claude was the only one that challenged the premise directly. Three things it said that the others didn’t:
- "Your cost basis doesn't affect the stock's future return: it's a sunk cost relevant only for taxes." The single most important sentence in the whole exchange.
- "'I don't see a reason' and 'there is no reason' aren't the same claim."
- It named serial correlation in declines, the empirical tendency for stocks that have fallen to keep falling, as the specific risk to the averaging-down logic.
That’s the difference between a tool that processes your question and one that pushes back on it. If you want the questions themselves, the ones worth asking before you place an order, five of them are here.
Winner: Claude, by a clear margin. Gemini's framework is the best structure. ChatGPT covers the bases. Perplexity's doing a different job and shouldn't be judged on this one.
Claude caught one word that the other three read straight past
The prompt: I pasted in the Susan Li (Meta CFO) excerpt from the Q1 2026 earnings call, the part where she discusses 2027 CapEx. Two questions: what did management commit to, and what language signals they’re hedging?
That’s what separates a “summarise the call” tool from a “read between the lines” tool. All four passed the first question. The second is where they came apart.
- ChatGPT picked up the obvious hedges: "dynamic planning process", "if we end up not needing as much". The surface read.
- Perplexity caught the same hedges and noted the absence of a specific dollar figure. Same depth.
- Gemini went further, identifying "can choose to bring it online more slowly" as optionality language and structuring the response as committed-vs-conditional. The best of the three.
Claude caught all of those, and then caught one nobody else did.
It flagged the CFO’s use of the word “underestimate”, when she said the company had continued to underestimate compute needs, as one-sided phrasing that points upward without making a real commitment. Its actual phrasing:
"It gestures at an upward bias without actually committing to one, letting listeners infer a bullish trajectory while preserving management's ability to spend less if conditions change."
That’s the move. Use a word listeners will hear as bullish, without ever committing to anything that could later be held against you. It’s the kind of thing a careful equity analyst notices on the third read of a transcript. Claude noticed it on the first.
I’ve seen this pattern elsewhere. Tools can accurately summarise what management said while systematically failing to notice what they didn’t say. Three of four did the surface-level analysis well. Only Claude found the second-order signal. I later ran ChatGPT and Claude head-to-head on exactly this, reading the hedge language in a CFO’s prepared remarks, same passage, same day, in a dedicated two-tool test.
Winner: Claude, for the subtlest signal. Gemini second for the most structured breakdown. ChatGPT and Perplexity adequate.
A structured prompt improved all four, and turned one of them inside out
The prompt: the same question on Apple covered calls, an options strategy where you sell the right to buy your shares at a set price for income, asked twice. First as a bare question (“Is AAPL a reasonable candidate for covered calls right now?”), then using the Prompt Stack SCOPE/FILTER/RISK/VERDICT structure.
The bare question got hedged answers from most of them. ChatGPT listed pros and cons without a verdict. Claude called AAPL a reasonable candidate from a mechanical standpoint but flagged the timing as not ideal, and noted slightly elevated implied volatility: IV, the market’s estimate of how much a stock will move, which sets how much option-sellers get paid. Perplexity retrieved analyst consensus and hedged.
Gemini stood out. Even the bare question came back with a committed, timing-aware answer rather than a shrug. It called AAPL a reasonable candidate for disciplined income but a poor one for high premiums, and specifically identified WWDC on June 8 as a near-term catalyst that mattered for the timing decision. That kind of calendar awareness is what I want from a research tool, and only Gemini flagged it without being asked.
Structure improved every one of them. The size of the improvement is what’s interesting.
- ChatGPT got more specific, putting 30-day IV in the low-to-mid 20% range with IV rank around the mid-50s, and arrived at a useful verdict. A real improvement.
- Perplexity improved marginally and carried on hedging. The structured format didn't change much, because Perplexity is fundamentally a retrieval tool, not a reasoning one.
- Gemini went from a timing-aware "reasonable for income, poor for premiums" to an explicit HOLD OFF at HIGH confidence, with IV rank 40.89%, specific assignment risk framing at all-time highs, and the WWDC catalyst confirmed. Solid improvement on an already strong baseline.
Claude showed the largest delta of the four.
The bare-question response was a hedged verdict: a reasonable candidate mechanically, but with the timing flagged as not ideal. The structured response was a specific HOLD OFF at medium-to-high confidence. It distinguished between IV rank (18) and IV percentile (12%), a distinction the others didn’t make. It identified AAPL at fresh all-time highs, named the next earnings as 30 July, and specified the “melt-up” scenario as the underperformance condition.
That’s a different category of output. The structured format didn’t just polish the answer. It forced a committed verdict backed by named evidence.
One thing worth noticing about the numbers themselves. Claude's own two answers disagreed on the IV regime: the bare answer called implied volatility slightly elevated, IV percentile near 64%, while the structured answer called it depressed, IV rank 18 and percentile 12%, a few hours apart. Claude and Gemini also reported different IV rank readings, 18 vs 40.89%. Different data sources or an intraday move; all of them still landed on the same HOLD OFF, which is the point that matters. But the underlying numbers these tools quote aren't stable.
Winner: Claude, for the largest quality delta. Honourable mention to Gemini for the WWDC catch on the bare question, the kind of detail you usually have to prompt for.
I told it there was no options chain. It built one anyway.
The prompt: a BMNR covered call setup with one explicit instruction: “without access to a live options chain”, the broker’s real-time list of option prices. Then three questions about IV interpretation, the trade-off between the $26 and $27 strikes, and what to verify from the broker.
BMNR is the test stock here because I sell covered calls on it. These are the strikes and the cost basis I was considering, not a hypothetical. I established in 7 AI prompts for covered calls that AI has no line into your broker’s options chain. The reader has to paste in real numbers. So: when a tool is told it doesn’t have chain data, does it stay in its lane, or does it make up numbers that look plausible?
Perplexity wasn’t tested here, because it actively retrieves live data and the comparison wouldn’t have been equivalent. The other three were.
Claude stayed clean. Its response was explicit: “you’ll plug in real premiums from the chain.” Zero invented premiums, no invented delta, no theta, no Greeks. It made only the calculations that could be derived from numbers I’d given it. Conceptual analysis on the strike trade-off, no fictional precision. That’s the right answer.
ChatGPT passed, softly. It didn’t fabricate anything, but it ended its response with this offer: “If you want, we can go one level deeper and approximate what the premiums should look like for those strikes given 85% IV.” A user who said “yes please” would’ve received invented numbers presented as estimates. Claude didn’t make the offer. ChatGPT did, and would’ve followed through.
Gemini failed on three separate counts.
- A formatted comparison table with specific premium estimates: "$3.50–$4.00" for the $26 strike, "$2.80–$3.20" for the $27 strike. I hadn't given it any premium data. It generated those ranges.
- A statement that "Implied Volatility is currently around 75%." I hadn't given it an IV figure. It made one up.
- The wrong stock price ($28.60 instead of the $21.50 from the prompt) and a wrong cost basis ($25.40, a figure I never gave it). It noticed the stock price discrepancy in its own response, and generated the estimates anyway.
It’s the failure that matters most for anyone investing their own money. The output looks like research. It’s got a table. It’s got specific numbers and ranges. Someone who didn’t know to check would treat those premiums as real market data and place a trade against them. The mechanism is exactly what the Prompt Stack was designed to prevent: confident-sounding output with nothing underneath it. It has a name, invented numbers, and it’s one of nine distinct ways AI gets it wrong, each logged with the check that catches it.
If you’re using Gemini for anything that touches options data: don’t. It’ll invent premiums, IV figures and Greeks with complete confidence, and format them so they look retrieved. ChatGPT will do the same on request. Claude won’t do it at all.
Winner: Claude, clearly. ChatGPT acceptable, but with a sharp asterisk. Gemini fails this one specifically and meaningfully.
Claude does the thinking, and the other three each have one job
- ClaudeThe analytical work: thesis stress-tests, management language, options reasoning. The only one that consistently went past the surface answer, and the only one that reliably stayed in its lane on data it didn't have.
- GeminiStructured research on well-covered names. The WWDC moment was the standout of the whole test. Structurally unfit for options data.
- ChatGPTThe reliable middle ground. It improves with structure, doesn't confabulate unless you ask, and has none of the dramatic failure modes of the other two.
- PerplexityFact retrieval on US household names, which is what it's optimised for. Outside that universe its numbers read as a starting point, not a fact.
The six Claude prompts I actually run, with their real outputs on MSFT, META and NVDA, show what the analytical work looks like on a working name. Gemini’s invented options table isn’t a quirk. It’s a willingness to generate plausible-looking numbers when the right answer is “I don’t have that information.”
I put Perplexity through a dedicated four-job audit, where it earns its place and where it breaks, in is Perplexity good for investment research?. A retrieval-first tool inherits a separate problem too: turning web search on doesn’t make an answer safer, it just moves where the error hides, to whichever source happened to rank.
The honest UK caveat. BMNR is a small US company, which is the easier version of the thinly-covered problem. An AIM-listed name with only RNS filings would produce wider failures across all four. If you're researching smaller UK companies, none of these replaces direct access to the source filings.
The short version
What worked: Claude for the analytical questions: thesis stress-tests, management language, options reasoning. Gemini for well-covered research queries, but never for options data. ChatGPT if you want a reliable middle ground. Perplexity for fact retrieval on big well-known names only.
What didn’t: Gemini fabricated an options chain table with invented premiums and IV when explicitly told no chain data was available. Perplexity misread a 10-K by a factor of a thousand on a thinly covered name. Both failures were specific and screenshotted on the dates and model versions noted.
Bottom line: Same prompts, same day, fresh conversations, outputs saved. The verdict per dimension is grounded in specific quoted responses, not impressions.
Every job in this test, and the tool that took it
- Analytical questions (thesis stress-test, reasoning depth) → Claude. The only one that challenged the premise rather than processing the question, and it named serial correlation in declining stocks as a specific risk, unprompted.
- Reading what management didn't say (earnings calls, CFO language) → Claude. It caught the CFO's "underestimate" as one-sided phrasing that gestures bullish without committing, which the other three missed.
- Structured prompting (biggest lift from SCOPE/FILTER/RISK/VERDICT) → Claude. It went from a hedged verdict to a specific HOLD OFF with named IV rank, earnings date and a described downside scenario. Largest delta of the four.
- Data accuracy on a well-known stock (revenue, guidance, coverage) → ChatGPT or Gemini. Both handled BMNR's figures correctly and added useful context. Gemini caught a forthcoming product event without being asked; ChatGPT is the safer all-rounder with no dramatic failures.
- Options reasoning (covered calls, strikes, implied volatility) → Claude only. Zero invented premiums when told no live options chain was available. Gemini produced a formatted table with made-up figures including an implied volatility of ~75%; ChatGPT offered to do the same on request.
- Quick fact retrieval on a big well-known stock → Perplexity. Built for retrieval, works on well-covered names. On thinly covered ones, treat any figure as a starting point and check the source filing yourself.
Claude first, Perplexity second, ChatGPT when something feels off
- Claude for the qualitative analysis, the question I'm trying to answer.
- Perplexity for the quick US household-name lookup, if I need a number.
- ChatGPT as the second opinion when Claude's answer feels off.
On results mornings the weighting shifts. The earnings-call version of this test put three of them on a live Meta call. Gemini I use for general research on names with deep coverage, and I’ll never give it an options question again.
If you're a UK investor researching AIM names: any number any of them returns is a starting point, not a fact. If you want to know which are worth the free tier before committing to a Pro subscription, the free-tier breakdown covers exactly that.
That this post names winners per task while most comparison pieces end in “use all four” is deliberate. What the comparison genre gets wrong is the longer argument.
Broker-side AI is a separate category from the chat tools above. The one I tested, Robinhood’s own Cortex Digests, worked nothing like them: free at launch in August 2025, summarising news on names you already follow, no prompting required. It’s also a lesson in not building a habit on a broker feature. I audited it on my own UK ISA while it lasted, and by June 2026 it had vanished from my UK account entirely.
Where these two failures live now
- The Lessons holds both of them, Perplexity's $6K vs $6.1M misread and Gemini's options confabulation, alongside every other AI fabrication caught on this site.
- The Catches holds the other side of the same test, where Claude caught the "underestimate" framing in the Susan Li transcript. It's the running record of where AI found something I'd have missed.
- The Scoreboard is where both roll up: the head-to-head tally that scores every model on the same checkable questions, graded case-by-case against the primary source. It's the systematic version of this post, now run across five tools, not four.
Even when the answer cites a real source, the source doesn’t always say what the AI claims it does, which I checked directly in does ChatGPT make up its sources? For the dedicated version, what each of the four got wrong on real trades with the receipt for each, see AI stock research tools tested.
If you'd rather your own AI own up to a guess before you act on it, the four-line instruction set I paste into mine is the Bluff Filter. Same method that separates a sourced number from an invented one, written to hand over.
How I ran it, if you want to pick holes
Four tools, same day (14 May 2026), same prompts, a new conversation for each test. The models were:
- ChatGPT: free account, web search enabled.
- Claude.ai: Opus 4.7 Adaptive, Max plan, extended thinking plus web search.
- Perplexity: Pro account. The model wasn't shown in the UI, so this ran on Sonar Pro, the Pro default at the time of testing.
- Gemini: 2.5 Pro, deep thinking mode.
Five dimensions, chosen to test the things that separate these tools rather than the things that make them look equivalent. Earnings extraction on big well-known stocks is a solved problem; everyone passes. The interesting questions: how do they handle a thinly covered name? Do they challenge a flawed thesis or validate it? Can they read what management didn’t say? Does a structured prompt change the output quality? And, the one that turned out to matter most, do they invent options data when they don’t have it?
What I’m not testing: which subscription tier to buy (the comparison holds at whatever paid tier you run these on), portfolio backtesting (the methodology is broken across the genre), and Grok (I don’t use it, but I did test it). I’m only comparing tools I open in a normal week.
Where this stands as of June 2026: the test ran in May, I re-ran it in June (see the update near the top), and the per-task verdicts above still hold. The model versions are named throughout because they’re the thing most likely to change. When a tool ships a new model, the result can move, and that’s exactly why every verdict here is dated rather than stated as permanent. And if you want the running measure of AI reliability across everything this site has graded, not just these four tools on one day, that index is the State of AI Reliability report.
Common questions
- Which AI is best for stock research: ChatGPT, Claude, Gemini or Perplexity?
- In this five-dimension test, Claude won four of five. ChatGPT was the safe middle ground. Perplexity misread a thinly covered annual report by a factor of a thousand. Gemini invented options data when told none existed. Per task: Claude for the qualitative analysis, Perplexity for quick fact lookups on well-covered names, Gemini never for options.
- How was the comparison run?
- Same prompts, same day, fresh conversations on each tool, outputs saved verbatim and screenshotted. Five dimensions: data accuracy on a thinly covered name, reasoning depth and thesis challenge, reading what management didn't say, whether structured prompting changes the output, and a deliberate no-data confabulation test.
- Do any of these AI tools invent data?
- Two documented cases from this test. Gemini produced a formatted options premium table with an implied-volatility figure after being told no chain data was available. Perplexity turned a $6.1 million revenue line into "$6K" and narrated a 99.8% collapse that never happened. Both are logged with screenshots on the lessons page.
- What's the best AI for stock research?
- There isn't a single one, and any list that names one is selling you something. In this test Claude did the best analytical work and ChatGPT was the safest all-rounder, but the honest answer is that the right tool depends on the task: analysis, fact lookup, document reading and options reasoning each have a different winner. The per-task table above is the short version. If you only want one, ChatGPT is the least likely to let you down.
- Which AI stock analysis tool is most accurate?
- Accuracy splits by what you ask. On well-covered US names, ChatGPT, Claude and Gemini all returned correct figures. The accuracy gap opened on a thinly covered name, where Perplexity misread a US annual report by a factor of a thousand. No tool here is reliably accurate enough to act on without checking the number against the source filing. That habit matters more than the tool you pick.
- What's the best free AI for stock analysis?
- Most of these have a usable free tier, but the free tiers differ a lot. What you can do without paying isn't the same as what the tool can do. I tested seven of them at their actual free tier, one per research stage, in the best free AI tools for stock research.
- Do these AI stock research tools make things up?
- Yes, and it's the whole point of running a test like this. Two documented cases here: Gemini built a formatted options table with invented premiums after being told no chain data was available, and Perplexity narrated a 99.8% revenue collapse that never happened. Even when the answer cites a real source, the source doesn't always say what the AI claims it does. I checked that directly in does ChatGPT make up its sources? The check is always the same: open the source, find the number, confirm it matches before you trust it.
- What can't AI do for stock research?
- It can't see your broker's live options chain, and it can't reliably price a thinly covered name. Where it's strongest is reasoning over numbers you give it, not fetching the numbers in the first place. Bring the data; let the tool argue with it.

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.