Bench Talk for Design Engineers: The Race for AGI Has No Finish Line
Artificial intelligence (AI) leaders treat artificial general intelligence (AGI) as a mythical sword lodged in a stone, with only the most worthy tech billionaire fated to wield it. And while they imagine that worthiness is determined by lofty goals of using AGI to benefit all of humanity, the race for AGI is a war of attrition that will be won by whoever amasses the most cash. Assuming, of course, that we can even agree on what the finish line of this race is.
In March 2023, Microsoft released a provocative paper titled “Sparks of Artificial General Intelligence: Early experiments with GPT-4”.[1] The responses were mixed, with critics noting that Microsoft had a version of GPT-4 that was not publicly available, meaning their results could not be independently verified.[2]
Three years and several versions of foundation models later, AI leaders are still trying to stoke the sparks of AGI into something bigger, but they don’t agree on exactly what they’re trying to reach (Figure 1).
Figure 1: Technology leaders continue the race toward AGI, even as questions remain about where the finish line lies. (Source: Miftakhul/stock.adobe.com; generated with AI)
In a January 2025 blog post, Sam Altman wrote, “we are now confident we know how to build AGI as we have traditionally understood it”.[3] But then in December 2025, he told the Big Technology Podcast that “we got it wrong with AGI, we never defined that…so my proposal is we agree that, you know, AGI kind of went whooshing by”.[4]
Sir Demis Hassabis, Co-Founder and former CEO of Google DeepMind, disagrees with Altman.
Personally, I’m not sure that I’d consider some of the most creative and groundbreaking minds in their respective fields mere “general” intelligence. But Hassabis believes that a true AGI should be able to push the boundaries in both creative and technical pursuits.
This blog examines why defining and measuring AGI remains a challenge, explores competing views on what AGI should be, and evaluates whether large language models (LLMs) are capable of achieving that vision.
How Will We Know When We Get to AGI?
The short answer is we won’t. Not until we agree on what AGI is and establish a comprehensive evaluation framework. We’ve already covered how some of the leaders in the AI race don’t agree on what they’re chasing, so now let’s look at the difficulty of declaring if something truly is AGI.
Think about the smartest person you know. How do you know they’re smart? They likely have a strong grasp of abstract concepts, a strong memory, and deep curiosity. Could you come up with a single benchmark test that quantifies their intelligence? Probably not. Human intelligence has too many facets to be accurately measured by Scholastic Assessment Test (SAT)-style standardized tests.
In a 2026 interview about benchmarking AI, cognitive scientist Gary Marcus said, “Anybody who’s worked in the field of psychometrics […] knows it’s very difficult to make a test actually measure what you want it to do.”[5] The challenge of benchmarking intelligence is amplified when the intelligence is artificial, because hidden model weights and secret training data make it unclear whether models are truly reasoning their way to a solution or simply drawing on a bank of practice exams in their training data.
Benchmarks Are for Bragging Rights, Not Quantifying Intelligence
A joint effort by the Center for AI Safety and Scale AI has produced the most difficult benchmark ever conceived: Humanity’s Last Exam. Comprising 2,500 questions curated by a team of 1,000 researchers, the exam undercuts AI’s advantage of having ingested the entire internet by posing questions so obscure that the answers do not readily appear online. Questions include the number of paired tendons supported by a particular bone in a hummingbird’s body and the translation of ancient Palmyrene script (Figure 2).[6]
Figure 2: Screenshot of a Humanity’s Last Exam question about a hummingbird sesamoid bone and paired tendons. (Source: lastexam.ai)
While some tech evangelists believe that a model that passes this intentionally difficult exam will unquestionably be considered AGI, test performance doesn’t always generalize to true intelligence. Standardized tests are sandboxes with rules, patterns, and clear answers. These constraints play to AI’s strengths in pattern recognition, but real-world complexity and nuance expose AI’s lack of common sense. LLMs can pass the bar exam, but they hallucinate citations when attorneys use them for research.[7]
AI doesn’t think like a human, so we need to stop testing it like one. To develop better benchmarks, the Stanford Institute for Human-Centered AI hosted a workshop where AI researchers sought to completely reimagine how we measure AI. During the workshop, as participants indicated whether they agreed or disagreed with statements about AI by standing in particular places in a room, they found that “there is almost no current consensus on how to define concepts such as AI ‘reasoning’”.[8] Much like there is no clear definition of AGI, there is also no reliable way to measure how intelligent a model actually is.
Does AGI Even Matter?
Assuming we can agree on a definition of AGI and create a framework to benchmark it, can we actually build one? Not with an LLM. The entire purpose of an ultra-smart AGI would be to trust it with complex, high-stakes decisions. But LLMs aren’t deterministic, so the same critical tasks we don’t trust an LLM to handle today would still be tasks we wouldn’t trust an LLM-based AGI to perform.
Hallucinations are inherent to LLMs, and while they can be mitigated, an LLM can never achieve 100 percent accuracy.[9] Most institutions would be reluctant to rely on a researcher whose conclusions occasionally contain inaccuracies presented with a high degree of confidence.
Returning to Hassabis’ vision of an AGI that’s on par with the likes of Picasso, Mozart, and Einstein reveals another critical flaw in the LLM-based AGI race: It’s paradoxical to expect an LLM to expand human knowledge when its outputs are limited to preexisting patterns. LLMs ingest vast amounts of text and break it into tokens, which can be single words, fragments of words, or single characters. The model learns statistical relationships between the tokens and uses them to generate outputs. In simpler terms, they generate outputs based on what they’ve seen before.
This means that while LLMs excel at tasks with limited scope and predictable results, they can only produce variations on what they’ve already seen rather than truly novel creations. Picasso, Mozart, and Einstein are household names because they combined their skill and creativity with existing patterns to create entirely new works. Einstein’s work may have started from a foundation of established classical physics, such as Maxwell’s equations, but he established an entirely new paradigm of modern physics that has guided innovation and perplexed engineering undergrads for the last century.
Brilliant human breakthroughs build on existing knowledge yet require a level of creativity and courage that LLMs have not yet demonstrated. The race for AGI might at least produce better models for us, but it will ultimately fall short of expectations unless AI leaders find a foundation beyond LLMs. Until then, we’re merely in the modern version of the infinite monkey theorem: waiting for infinite graphics processing units (GPUs) to come up with the theory of relativity.
Sources
[1] https://arxiv.org/pdf/2303.12712
[2] https://www.nytimes.com/2023/05/16/technology/microsoft-ai-human-reasoning.html
[3] https://blog.samaltman.com/reflections
[4] https://www.youtube.com/watch?v=2P27Ef-LLuQ&t=3350s
[5] https://www.youtube.com/watch?v=iFYF_e1GSGI&t=1508s
[6] https://lastexam.ai/
[7] https://www.reuters.com/legal/litigation/us-appeals-court-rebukes-lawyer-over-fake-hallucinated-case-citations-2026-07-10/
[8] https://hai.stanford.edu/news/smart-enough-to-do-math-dumb-enough-to-fail-the-hunt-for-a-better-ai-test
[9] https://openai.com/index/why-language-models-hallucinate/
About the Author
Matt Campbell
Matt Campbell is a technical storyteller at Mouser Electronics. While earning his degree in electrical engineering, Matt realized he was better with words than with calculus, so he has spent his career exploring the stories behind cutting-edge technology. Outside the office he enjoys concerts, getting off the grid, collecting old things, and photographing sunsets.

