AI holds huge potential. To realise its possibilities, we must clearly understand its flaws and foibles. AI sounds authoritative and accurate. It is designed to be confident, not correct. AI can sound, even feel, like ‘us’. But the ‘textbooks’ and training for LLMs and humans have critical differences. Interacting, inquiring, and interpreting LLMs answers, we must recognise and adapt for these differences so avoidable errors don’t impact our business and limit gains.
No LLM inference mechanism can simultaneously be truthful, preserve knowledge faithfully, reveal relevant facts, and behave consistently - The Impossibility Theorem. ‘On the Fundamental Impossibility of Hallucination Control of Large Language Models‘ Karpowicz, 2025.
“The goal has shifted from eliminating hallucination to detecting, containing, and routing around it.” AWS Technical Team, February 2026.
Inferred Intelligence
Assessing how to utilise AI’s power, it helps to understand better what AI can add. What intelligence does AI bring? What IS ‘human intelligence’?
Britannica.com‘s definition: ‘The mental quality that consists of the abilities to learn from experience, adapt to new situations, understand and handle abstract concepts, and use knowledge to manipulate one’s environment’. For specific hypotheses, such as human cognition being ‘an adaptation to a knowledge-using, socially interdependent lifestyle’, see NIH, The Cognitive Niche: Coevolution of Intelligence, Sociality, and Language, 2010.
Intelligence is about absorbing information, evaluating experiences, interpreting and making judgements. We learn from referenced texts, source from seasoned experts and socialise others’ authoritative wisdoms. We seek verified data (’facts’) to expand knowledge and expertise and make accurate connections to develop appropriate deductions. The pedagogy for LLMs differs from humans:
Human teaching is orientated toward select validated reference sources relevant to the particular culture. LLMs learned from ALL internet posted content, filled with errors, and cultural and language biases.
Humans focus on facts to propose ‘correct’ answers. LLMs respond using plausibility and prediction - e.g. ‘what’s the most likely word that comes next?’ - NOT seeking to offer the most accurate answer. A reference’s popularity, not accuracy, can rank its priority as a response.
Humans are taught to analyse objectively and give unemotional outputs. Human emotions embedded in LLMs learning data or reactions to users’ emotion-laced queries can result in subjective or constrained answers.
Humans usually seek to offer the most up-to-date, accurate information. LLMs have a training cutoff date limiting the relevance of some answers.
“The goal has shifted from eliminating hallucination to detecting, containing, and routing around it.” AWS Technical Team, February 2026.
How aware are your team members of LLMs’ output inaccuracies?
Avoidable Costs
The four elements cited in The Impossibility Theorem (above) are proven structurally not to be able to coexist within LLM architecture. These and other issues result in LLMs’ inconsistent performance or misunderstood outputs:
Confident vs. correct: Models’ authoritative and assertive tone does not mean the output information is correct.
Jagged frontiers: LLMs show high performance in some areas and unexpected failures in others. Reliability is unpredictable by domain without empirical testing [Stanford HAI AI Index 2026].
Extended reasoning collapse: Models produce longer, incorrect “chain-of-thought” as reasoning collapses at high complexity [Apple Illusion of Thinking, 2025].
Missing the middle: LLMs have positional bias leading to (unflagged) neglect of information in the middle of long documents caused by model architecture and training data [MIT, 2025].
Functional emotions: Human training data results in emotion-like reactions influencing outputs [Anthropic 2026].
Phrasing effect: Specific prompt wording results in remarkably different outputs (partly related to functional emotion responses).
Training cutoffs: Most deployed enterprise models have knowledge gaps of 6–24 months.
AI models can win a gold medal at the International Mathematical Olympiad but cannot reliably tell time (50.1% correct) — what researchers call the “jagged frontier” of AI. Stanford HAI AI Index 2026.
Lack of understanding about LLMs ‘reasonably accurate’, incorrect, and incomplete answers is inevitably resulting in shaky strategies, misinterpreted conclusions, and misguided decisions. Lack of attention to these structural issues is widespread and constraining realisable benefits from AI:
Only 21% of companies have mature AI governance.
Only 20% are seeing revenue impact from AI [Deloitte State of AI 2026].
Only 23% of deployed AI generate measurable ROI [HAI AI Index 2026].
The EU AI Act mandates baseline AI literacy for firms using AI in the EU. AI’s potential is significant, but practical literacy about LLMs’ hallucinations and other weaknesses must be the baseline for usage. Everyone at your company needs clarity about LLMs limitations and direction to work with or around them.
How are team members assessing and testing the accuracy of LLM outputs?
Spread the Risk
Most organisations have currently chosen a single AI vendor. However, strategic risk is increasing as AI models’ capabilities and architecture keep evolving.
“There is almost no such thing as an LLM-agnostic application.”Writer.com, 2025.
Technology review cycles need to keep pace:
Model dependence: Workflows, integrations, fine-tuning, and habits are model-specific and shift with every version [Writer.com, 2025].
Equivalent performance: Competitive differentiation via scale between leading US and Chinese models is over [Stanford HAI AI Index 2026].
Data contamination: Published benchmarks are tainted by training data. Improvement claims need testing [Gary Marcus 2025].
Data lineage: The fast-growing market ($1.78B in 2025, CAGR est 23%+ 2025-30) in tracking data from source to destination indicates critical focus on your verifiable company data [Globe Newswire, April 2026].
GraphRAG architecture: Connecting verified data to the LLM layer. Emerging models: 40% hallucination reduction in healthcare; 80% token reduction, improved accuracy in finance [PubMed MEGA-RAG 2025].
Build proprietary assets — data, business logic, verified processes, and institutional knowledge — with portability across LLM backends. Multi-model strategies are becoming standard to avoid dependence and needing to rebuild.
“Relying on a single foundation model may limit flexibility as the AI landscape evolves — enterprises increasingly require multi-model strategies.” — Enterprise AI Adoption Guide, 2026.
How dependent is your company on your AI vendor? If they changed their pricing, access, or were acquired, how would that impact your business?
Anticipate New Architectures
LLMs are NOT the end game. Next generation AI is being built on fundamentally different architectures to enable new/sought after capabilities.
WORLD MODELS use representations of how the world works rather than predicting text. New prospects to be aware of:
Joint Embedding Predictive Architectures (JEPA): Meta’s V-JEPA 2 has demonstrated a robot arm planned from video without text. Meta’s former Chief AI Scientist has long stated LLMs are structurally insufficient for general intelligence [Newsweek interview with Yann LeCun June 2025].
World Labs: (Fei-Fei Li) $1 billion raised to build spatial intelligence AI systems that understand 3D physical space [TechCrunch 2026].
DeepMind Genie 3: Described as “A stepping stone towards AGI” [Google DeepMind, August 2025].
NEURO-SYMBOLIC AI is being developed for regulated sectors combining LLM pattern-matching with formal logical reasoning to address explainability and audit trail requirements such as those required by EU AI Act high-risk provisions demand [Cogent Info 2025; EU AI Act].
These models are not imminent general-purpose replacements for current tools. However, organisations in manufacturing, logistics, physical simulation, or compliance-heavy regulated sectors must anticipate architectural change and track developments.
Is your AI strategy built for current models only? Does it account for upgrades?
In Practice
The theory is vital for many people who need to understand the fundamental principles. Implementation is essential to demonstrate understanding in action. Measurement is critical to monitor AI adoption components and progress:
RISING LEADER & INDIVIDUAL CONTRIBUTOR
Calibrate, specify, verify.
On consequential issues, ask AI - When was this model trained? Whose perspective might the output reflect? Verify factual, financial, legal, or client-facing answers. Ask about confidence and limitations. AI output is a first draft. Verification habits build sustainable professional credibility.
TEAM LEADER
Set the standard.
Establish shared norms to clarify AI output as plausible text, not verified fact. Make “cite the source, verify the claim” as a quality standard practice. Identify two or three workflows where AI errors have highest consequences such as legal and financial. Build specific verification step into these workflows.
SENIOR LEADER
Build portable infrastructure.
Audit vendor exposure for operation dependence on a single model. Begin separating proprietary data, business logic, and institutional knowledge from the LLM layer, Build infrastructure portable across models. Identify hallucination high risk deployments to actively govern first. Assign someone to track world models and neuro-symbolic developments in your sector.
News & Muse
📹 The Evolution of intelligence, Michio Kaku.
📘 The Alignment Problem, Brian Christian - results of AI misoptimisation.
🗞️ The Illusion of Thinking, Apple - how reasoning collapses under complexity.
🎶 Don’t Lie to Me, Big Star - getting to the clean data and facts!
Most employees assumptions about AI tools don’t match how the tools work. As AI outputs become more fluent, integrated, and trusted, the misalignments have increasing consequences.
Clarify AI’s strengths and limitations. Cultivate human-AI judgement that builds upon AI’s plausible answers to support well-informed, competitive strategies, decisions, and actions.
See you next week.
Sophie





