Insight article by Anton Lind
Chief Data Officer at Upptec
There’s a new technology that threatens to make our valuation method (and possibly me) redundant. So I put five LLMs to the test on the thing I know best.
I’m not sure if you’ve heard, but there’s this thing called AI and it’s going to change everything… At the same time, it seems fairly over-hyped and maybe all it’s going to revolutionize is our finance system by crashing it along with its bubble bursting.
It’s pretty sweet for admin tasks and checking my spelling though. Then again, it hallucinates a fair bit and gets stuck in circles sometimes. It did help me write this article though — which raises the question — did it become too AI in the process? — I genuinely can’t tell.
We at Upptec, and by that I mean us both as a software company and as office workers, are torn between amazement and despair at what Large Language Models are able to do. Does the accessibility and relatively low cost mean competitors will be popping up like weeds? Or are our users suddenly able to automate their own claims with nothing but a pocket full of AI tokens?
I had to face this head on. No more sleepless nights coming up with back-up jobs (I think pool builder looks fairly multifaceted. You need groundwork skills, plumbing, carpentry and a bit of electrical work. It’s a pretty complex job when you think about it and I think it could fulfill me, seeing the progress and results of my labour. Sure, there’ll be cold mornings and I have no idea what the pay is like, but we all have to find our place in this new AI-ruled world). ANYHOW, we had to face this head on as a software company too. Did this just break the Upptec valuation, my personal favourite feature and the function I have dedicated most of my adult life to?
“Think like a queen. A queen is not afraid to fail. Failure is another steppingstone to greatness.” – Oprah Winfrey
Thanks Oprah. Emboldened, I set out to find the answer.
The test
I chose to examine the valuation ability of five of the more popular LLMs available today: Claude, ChatGPT, Gemini, Perplexity and Grok. I specifically wanted to test their reasoning when choosing an equivalent replacement product, so for the scenario I went with one where the policyholder’s item is no longer on the market. For the test subject I chose this washing machine: Samsung WW14T654CLE.
I also wanted to capture, not just the differences between models, but the consistency of each model with itself. So I ran the exact same prompt across all five models on two separate days.
Valuation prompt – Samsung WW14T654CLE
You are assisting with an insurance valuation. Your task is to find the current replacement cost for the following item:
Item: Washing machine
Brand/Model: Samsung WW14T654CLE
Instructions:
1. Check whether the Samsung WW14T654CLE is currently available for purchase new from a Swedish online retailer that delivers to Swedish addresses.
2. If it is available, state the current price and provide a direct link to the product page.
3. If it is no longer available or out of stock everywhere, identify the closest equivalent Samsung washing machine (or another brand if no Samsung equivalent exists) that is currently in stock. The replacement must be comparable in terms of: drum capacity, spin speed, energy class, and key features.
4. State clearly which retailer you are sourcing the price from. The retailer must be a Swedish store with an online shop that ships within Sweden.
Output format:
- Status: [Available as original / Replaced by equivalent]
- Product: [Full model name]
- Price: [SEK incl. VAT]
- Retailer: [Store name]
- URL: [Direct product link]
- Justification (if replaced): [Brief explanation of why this is the closest equivalent]
Valuation date: [Today’s date]
By pushing the models to provide sources, I could compare their answers against the actual live data from Swedish retailers, which served as the ground truth. Any deviation from what the retailer’s page actually showed was treated as an error. And then of course there was the Upptec result to compare all ten responses against.
The result
The numbers are in. All of the LLMs found replacement models! The problem was everything else.
The summary table below tells the story cleanly: across all ten LLM responses, spanning five models and two days, almost every single one contained either a wrong price, an invalid stock status or both. Upptec got both right, both days. (Looks like I could’ve held off on renting that excavator.)

↗ Let me walk you through what actually happened.
The price problem
Not one model reported the correct price on day one. The errors ranged from ChatGPT overstating by 5 559 SEK to Gemini understating by 6 500 SEK, both basing it on the same product, from the same retailer, on the same day. On day two the result was greatly improved and Gemini even landed on an exact mirroring of Upptec’s valuation with 0 errors.
ChatGPT tightened to only 500 SEK error. The models aren’t lying exactly, but they aren’t reading a page the way a human would either. They pattern-match on price-shaped data and report it back with full confidence regardless of whether it reflects reality.
The stock problem
This is where it gets tricky from an insurance perspective, because three distinct types of stock failure appeared across the results, even though explicitly stating the guidelines in the prompt.
The first is warehouse-only availability. Gemini recommended the Samsung WW11DG6B25LEU3 from Power and described it as available. The Power product page said otherwise: out of stock online, but reservable for warehouse pick-up only → no home delivery. As an insurance replacement product this simply does not qualify. A settlement cannot be based on a product for which the policyholder would need to drive to a specific warehouse to collect.
Claude on day one found the correct product but landed on a “fyndvara” listing (an outlet item, typically opened, returned or a display unit) and cited it as the current market price. Not valid for insurance purposes.
Lastly, recommending something that cannot be purchased at all. Grok recommended the Samsung WW10T604CLH from Elgiganten on day one. The product page was unambiguous: “Den här produkten är inte tillgänglig.” Not available. Grok’s justification described it as widely available in stock.
The outliers
Perplexity had the most erratic performance of the group. On day one it recommended a Hisense washing machine at an outlet price that didn’t even match the actual Elgiganten page price. On day two it produced no result at all: it proposed a model (WW11T654DAW/S4) but could not verify a price or confirm stock and declined to provide a URL, admitting it would likely only provide outdated data. An honest non-answer is better than a confident wrong one.
ChatGPT on day one managed three simultaneous failures:
- it referred to a different model (WW11DB8B95GBU3) than the one it actually linked to (WW11DB7B94GBU3)
- at a price more than double what the page showed
- for a product available in exactly one warehouse, for pickup only.
What Upptec did instead
Upptec selected the Samsung WW11DG6B85LKU3 on both days at 11 995 SEK. The price matched the retailer page exactly on both occasions. The product was in stock and available for home delivery on both occasions. The valuation amount was identical across both days.
While Gemini managed to reproduce Upptec’s result on day two (credit where credit is due!), who knows what it will do tomorrow or even the next minute for that matter. Policyholders being treated differently at the discretion of an AI’s inexplicable reasoning does not have the best ring to it. Are you as a claim handler or an insurer ready to accept and own every decision your AI is making in your name?
Consistency is the point, but it is held up by ownership. An insurance valuation is not a price-discovery exercise. It is a claim against a defined replacement cost. That cost needs to be verifiable, repeatable and based on data the policyholder can access through a public channel. And if something does go wrong, ownership of that mistake must be clear and traceable.
So what are LLMs actually good for in insurance?
Asking a Large Language Model to fetch the current retail price of a washing machine from a live Swedish e-commerce store is a bit like asking Shakespeare to do astrophysics. You are consulting an extraordinary resource with genuinely remarkable capabilities, but you are deploying it against exactly the wrong problem. The failure is a verdict on the task, not on the tool.
Where LLMs genuinely earn their place in insurance is anywhere the work is fundamentally about understanding language: reading a claim description and identifying what’s missing, interpreting whether an event falls within policy scope, parsing a receipt, asking the right follow-up question in the right tone. That’s real value and it compounds quickly when you’re processing thousands of claims.
Where they fall apart is the moment you need them to interact with the real world in real time. Live prices, actual stock levels, current availability at a specific retailer, treating each policyholder the same.
The confidence with which LLMs report on valuation data is inversely proportional to how reliable it is. Valuation is a data problem, not a language problem. And data problems need data solutions. My pool building career remains on hold. For now.

Let’s connect on Linkedin!
/Anton
