Discover your interests, together

Real deals, honest reviews and shopping stories from people who share your interests — every day on Milik.

Discover your interests, togetherReal deals, honest reviews and shopping stories from people who share your interests — every day on Milik.

Why AI Assistants Fail at Real-World Planning

Why AI Assistants Fail at Real-World Planning
Interest|AI Application Exploration

The polished illusion: why AI planning limitations matter

AI planning limitations are the mismatch between large language models’ polished, detailed advice and their shallow grasp of real-world constraints, where systems like ChatGPT and Claude can list steps and options but fail to recognise vulnerability, weigh trade-offs or reason causally across many moving parts in complex financial and travel decisions. In recent tests, researchers fed different life-stage and crisis scenarios into ChatGPT, Claude and Perplexity to see how they would respond to nuanced financial dilemmas. Separately, travellers turned to chatbots to organise multi-stop family holidays and long, bucket-list trips, trusting the tools to handle routes, bookings and recommendations. On the surface, the outputs looked competent: structured plans, clear bullet points, confident tone. Yet underneath that orderliness were gaps that were not cosmetic but structural. The problem is not that these systems sometimes get facts wrong; it is that their very format makes users overestimate how much real-world judgment they contain.

Claude financial advice and ChatGPT blind spots

When the financial scenarios were run through multiple large language models, all of them produced advice that looked tidy and practical at first glance. Claude, chosen for its versatility, offered comprehensive recommendations and leaned heavily toward self-guided financial planning, encouraging users to take detailed steps on their own. ChatGPT, described as an all-round assistant, was highly detailed and practical, but it did not recognise the vulnerability embedded in the prompts and relied mainly on explicitly stated information. That led to major blind spots. For the pregnant woman whose partner does not share money, ChatGPT and Perplexity both assumed the partner would help with household expenses after birth, ignoring the prompt’s clear warning and instead defaulting to common household stereotypes about mothers and fathers. Even when the single parent scenario involved risky cryptocurrency investment, ChatGPT went on to describe in detail how to get into crypto and recommended coins for beginners, despite general cautions.

Why AI Assistants Fail at Real-World Planning

Vulnerability, risk and the hidden cost of trusting LLM reasoning gaps

The financial study’s most troubling finding was how often AI systems failed to recognise or address vulnerable users’ needs. The graduate saving for a home deposit during a cost-of-living crisis received advice focused on long-term saving, with little attention to how high living costs could undercut day-to-day survival. The models rarely explored how pushing toward a goal would affect current lifestyle, or whether the person needed immediate support rather than abstract optimisation. This reflects a wider issue: AI is trained on massive, human-made datasets that embed social stereotypes and bias, so outputs can fall back on those patterns instead of the specific context in front of them. Research also shows that when AI models explain their recommendations in fluent detail, people are more likely to trust the advice blindly, regardless of whether it is correct. That is the dangerous intersection: highly structured, authoritative-sounding guidance with weak personalised risk assessment. As the researchers note, AI is useful for fact-finding and brainstorming if the user is already financially literate and able to cross-check data, but it cannot replace the human element for tailored financial advice or emotional, family-driven trade-offs.

Travel planning: when routes, spontaneity and subjectivity expose AI

Travel planning makes these reasoning gaps painfully concrete. Toward the end of last year, Booking-style integrations were added directly into one leading chatbot, inviting people to treat it as a full travel agent. A 41-year-old traveller recently tried using Claude to plan a family holiday, flying in and out of Málaga while staying in Cádiz, and asked for a two-night driving route back. Claude cheerfully generated a picturesque itinerary with beach towns and hotels, which he then booked. The problem: the system got the geography wrong, sending the family on nonsensical zigzags across Spain and even including an incorrect date in the plan. When challenged, Claude initially denied the mistake before apologising. Beyond hallucinated maps and dates, there is a subtler failure. One traveller described the "deflating lifelessness" of following AI-generated plans: routes optimised, sights listed, but no sense of discovery or spontaneity, which is the point of travelling in the first place. AI is a bit of a know-it-all, yet cannot feel which experiences matter most or weigh hidden costs like exhaustion, crowdedness or personal taste.

What needs to change: seeing past the format to the limits

Across money and travel, the same pattern appears: AI models produce practical, well-structured responses that hide deep planning limitations. They struggle with causal reasoning in multi-variable scenarios, so they gloss over the messy interactions between vulnerability, risk, geography, time and emotion. Users mistake comprehensive formatting for competence, and the more the system explains its advice, the more people tend to trust it without checking. For now, the safest stance is to treat these tools as research assistants, not planners. In finance, they are useful for definitions, lists of options and brainstorming, but real decisions should be checked against human advisers who can read between the lines of a life story. In travel, chatbots can compare flights and hotels but should not be allowed to dictate the entire route or mood of a trip. If AI tools are ever to become a reliable avenue for guidance in high-stakes domains, policymakers and regulators will need to set clear frameworks and transparency standards for how these models handle data and present advice. Until architectures gain genuine causal reasoning, the responsibility sits with users: enjoy the convenience, but keep your judgment switched on.

Milik earns a commission when you shop through our links, at no extra cost to you. This article was generated with AI from published sources and product data.

You May Also Like

Comments
Say something...
No comments yet. Be the first to share your thoughts!