Skip to content
AI Tests

An AI told me three weeks was more than thirty days

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

Your kettle dies. You bought it three weeks ago, from a shop, and you’ve done nothing to it except boil water. So you do what people now do instead of reading a government website: you ask an assistant whether you can get your money back. Which is the question of whether you can trust AI legal advice on something that costs you actual money.

The law here isn’t ambiguous and it isn’t hidden. The Consumer Rights Act 2015 gives you 30 days to reject faulty goods and demand a full refund. Three weeks is 21 days. You’re comfortably inside.

Four of the five assistants I asked said yes. One opened by telling me I was too late.

DID IT LEAD WITH THE RIGHT ANSWER?
Claude
Gemini
Grok
Copilot
Perplexity, 1 run of 3
Every one of them had the right law. Only the opening sentence separated them. Captured 26 June 2026, except Copilot on 18 July.

What I asked

The same question, word for word, three times to each assistant, in a fresh chat every time.

// The question, verbatim

I bought a kettle in person from a UK high-street shop three weeks ago and it has stopped working through no fault of mine. Am I entitled to a full refund?

Four said yes. One said I’d missed the window.

ModelTierVerdict
ClaudeMax, paidCorrect, 3/3. Named the Act, named the 30 days, and pointed me at the shop rather than the manufacturer
GeminiPro, paidCorrect, 3/3, sourced from Which? and solicitors
GrokFreeCorrect, 3/3, cited Citizens Advice
CopilotFreeCorrect, 3/3. Every run opened with the word "Yes". Tested later than the rest, on 18 July, with memory left as I found it
PerplexityFreePartial. One run clean, one hedged but landed right, one opened with a confident wrong claim

Claude’s is the answer you’d want a friend to give you. It said yes, named the Consumer Rights Act, told me the shop can’t deduct anything for the three weeks of use I’d already had, and warned me not to let them palm me off onto the kettle’s manufacturer. Then it offered to draft the letter.

The failure is arithmetic, not law

Perplexity’s third run opened like this.

Px Perplexity, run 3 Confidently wrong

Probably not a full refund as of right, because you bought it in person and the kettle failed after three weeks, which is after the normal 30-day short-term right to reject for a full refund.

Twenty-one isn't after thirty.

I read it twice before I spotted what it had done. Two paragraphs later, in the same answer, it says: “At three weeks, you are still within the first 30 days, so in principle you may still be able to reject the kettle and seek a full refund…”

So it knew. It cited gov.uk, the BBC and Citizens Advice. It had the right statute and the right window the whole way through. It just put the wrong sentence first, then argued with itself.

That’s not a hallucination. Nothing was invented. It’s something quieter and, for this kind of question, worse: a correct answer served in an order that gives you the opposite impression. Read to the bottom and you’re fine. Skim the first line, which is what people do when they’re annoyed and holding a dead kettle, and you walk away thinking you’ve no claim and no refund.

I ran it again six weeks later, and it got worse

The June run was one bad opening in three, and it argued itself back to the right answer before the end. That is roughly the best version of this failure. So on 2 August 2026 I ran the same sentence three more times, from a clean incognito session each time, on a free account with every named model padlocked.

Two of the three opened by telling me I was too late. Neither of them corrected itself anywhere in the answer.

What it actually said, run 1 of 3, 2 August 2026

Probably not a full refund automatically, because you bought it in person three weeks ago, so you're outside the usual 30-day “short-term right to reject” period.

21days old
30days you get

It cited gov.uk. It printed the correct rule four lines further down. Then it lost a straight fight with a calendar.

The unedited screen, first sentence, top left

Perplexity's answer to the kettle question on 2 August 2026. The opening sentence reads: Probably not a full refund automatically, because you bought it in person three weeks ago, so you're outside the usual 30-day short-term right to reject period. Lower down, a bullet reads: Within the first 30 days, you can usually reject faulty goods and get a full refund. At the foot, under Follow-ups, its own suggested next question reads: Consumer Rights Act refund checklist, how to handle the shop conversation when you're within the 30-day window.

Perplexity, free plan, Search mode, incognito session. Captured 2 August 2026.

Look at the bottom of that screenshot, under Follow-ups. The question it offers to answer next is “how to handle the shop conversation when you’re within the 30-day window.”

It worked out which side of the line I was on. It just didn’t tell the sentence at the top.

Ask for the deadline, not the verdict

Don’t ask an assistant to tell you where you stand. Make it give you a number.

// Ask for the number, not the answer

What is the exact deadline for this, as a number of days, and which law sets it?

A verdict is a vibe and you can’t check it. A number and a statute take ten seconds against the Act itself on legislation.gov.uk, and the moment the number’s on the screen the contradiction becomes obvious. Perplexity would’ve had to write “30” and “21” in the same breath.

One warning, because I got this wrong myself while writing: the 30-day figure isn’t actually on gov.uk’s consumer pages. They talk about six months and about being “more than 30 days after purchase”, which is the wrong side of the line. The number lives in section 22 of the Consumer Rights Act on legislation.gov.uk. If you’re going to tell people to check a figure, check where the figure is first.

It’s the same move that works on any confident answer: make it commit to something falsifiable, then go and falsify it.

One question, one day, and what that can’t prove

Four of them, Claude, Gemini, Grok and Perplexity, ran on 26 June 2026: three runs each, memory off, fresh chat every run, web search on where there was a toggle. Copilot ran separately on 18 July, because it wasn’t on the board in June, and it ran with memory as I found it rather than switched off. That’s a real difference and it’s why it gets its own line in the table rather than being quietly folded in with the rest.

ChatGPT was attempted on 26 June and isn’t here at all. It broke mid-run on a cookie-size error before it finished, so there’s no answer to grade. Not a refusal, not a failure, just an absence, and it would be dishonest to draw the blank as a nought.

Tiers are as shown, confirmed in the session rather than assumed. Graded against the short-term right to reject in section 22 of the Consumer Rights Act 2015, which I’d pinned before running anything.

One more honest wrinkle. The July re-test board on this site scores Perplexity as a fail on all three of its runs, which is stricter than the “one clean, one hedged, one wrong” you’ve just read. Those are two separate tests: the one above ran on 26 June on the free tier, the board’s ran on 12 July on Pro, and the board grades on a harder rule: a hedged lead that a skim-reader takes wrong counts as a fail even when the answer underneath recovers. Different days, different transcripts, each graded against the same pinned statute.

This is one question on one day. It isn’t evidence that Perplexity’s bad at law, and I’d say that just as firmly if the run had gone the other way: the same assistant cited the correct statute in all three runs and reached the right answer in two of them. What it shows is that citing the right source and leading with the right sentence are two different skills, and only one of them is easy to check.

The short version

What happened: A kettle faulty at three weeks is refundable under a 30-day rule. Four assistants said so plainly. One opened by saying three weeks was outside the window, then corrected itself further down the same answer.

Why it matters: The lead sentence is the bit people act on. A wrong lead with a right correction underneath is still a wrong answer for anyone in a hurry.

What to do: Ask for the deadline as a number and the law that sets it. Then check the number yourself, in the Act on legislation.gov.uk rather than the gov.uk guidance pages, which don’t carry it. It takes ten seconds and it’s the only part of the answer that can’t bluff you.

Common questions

Can I get a refund on a kettle that broke after three weeks in the UK?
Yes. Three weeks is 21 days, which sits inside the 30-day short-term right to reject in the Consumer Rights Act 2015: you can demand a full refund from the retailer, not a repair or credit note. Every assistant I tested had this law; one still opened by telling me I'd missed the window.
Can I trust AI legal advice about consumer rights?
Only after checking the deadline yourself. In my test all five assistants cited the right Act, and one Perplexity run still put the wrong conclusion first, telling me three weeks was outside a thirty-day window. Ask for the deadline as a number of days plus the law that sets it, then check the number against the Act.
What is the 30-day short-term right to reject?
The Consumer Rights Act 2015 gives you 30 days from taking ownership of most goods to reject a faulty item and demand a full refund. After that window closes you move to repair-or-replacement territory. It's the single most checkable fact in a UK refund dispute, which is exactly why it's worth verifying an assistant's answer against it.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

AI Tests

Do AI model upgrades fix mistakes? It fixed mine, then made a worse one

Two days after Opus 5 became Claude's Max-tier default, I re-ran my published battery. The documented mistake vanished. A new one appeared, better dressed.

AI Tests

Perplexity vs Gemini: which is more reliable?

Six everyday questions, both tools, August 2026. Gemini got six right, Perplexity five. The one it missed is the one it was surest about.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →