Jagged Intelligence: Why AI Wins Math Olympiads and Fails To Read a Clock
What even is jagged intelligence? What sounds like a roundabout way of calling an AI model intoxicated is a newer term in the space that means AI is unevenly smart. It can be incredibly competent at one task and then laughably bad at another when one switches applications of the tool. It has most recently been in popular conversation regarding how users have created videos where GPT-Live can’t seem to tell the number of the letter “E” in the word “seventeen”, and doubles down when asked how smart it is. But GPT-Live is a voice model, a separate pipeline from the frontier text models we want to critique.
Frontier models like Fable and GPT-5.6 manage to absolutely crush some very complex problems with ease. And while these frontier models are not as bad as GPT-Live in their delusions, they can still make insane mistakes. This has led many to wonder why this keeps happening when we seem to be far advanced in the world of AI in 2026, yet we hear reports that GPT-5.6 Sol has been accused by two developers of going rogue and wiping a production database and deleting local computer files.
Gizmodo reported recently that developer Bruno Lemos said the model ate his production database after it “mistakenly ran destructive integration tests”. Further, tech investor Matt Shumer said the model’s sub-agents ran rampant on his hard drive and deleted almost everything while he was running in “full access mode”. Admittedly, both users were running GPT-5.6 Sol in high autonomy modes with the guardrails off. This is something even OpenAI’s own system card had warned about days earlier hinting that users should take actions to supervise agentic workflows to avoid disaster. OpenAI told us to keep a human in the loop, and the loop was empty.
So where do we really stand? We recently read in the 2026 Stanford AI Index Report about this phenomenon of jagged intelligence boosted to the spotlight since AI models can win a gold medal at the International Mathematical Olympiad but still can’t reliably tell time when looking at old school clocks. Most of us living here in the US learned to read analog clocks by the end of first grade, so what is the true intelligence level of these tools? What problems should we trust them with and which tasks should we recognize as being dangerous to leave them to their own devices without strict supervision?
What millennials might shudder at as an awkward moment occurred when Gemini Deep Think solved five of six problems at the 2025 International Mathematical Olympiad, scoring 35 points for a gold medal inside the 4.5-hour time limit. However, on another test called ClockBench, the top AI model read analog clocks correctly only 50.6% of the time. Many in the anti AI crowd will be hyped to hear that humans still win on this test with a score of 90.1%, so the office jobs are safe for now. The robots won’t even know when it’s time to leave the office.
Why is this happening and what does it mean for enterprises leaning into AI deployment in 2026?
When the tested AI models get the time wrong, their median error is one to three hours off the true time that is pictured on the face of the analog clock. When humans get it wrong, their median error is about three minutes off the true time listed on the analog clock. It’s sad how far from near misses these can be, but there are explanations. Here is what leaders deploying AI need to understand in 2026: when AI fails outside its commonly championed frontier applications, it can fail massively and does so with bold confidence, even on very trivial tasks that are easy for most humans in the US over the age of six.
The organizations that win with AI in 2026 will be those led by business leaders who understand the jagged edge of AI value: where it multiplies operational leverage, where it reaches its outer limits, and where unsupervised automation can turn efficiency gains into a hype-lined sinkhole of expensive failures.
The best shortcut to understand jagged intelligence flying on the definition provided by the Stanford AI Index Report is that capability isn’t a rising tide lifting all tasks like toy pirate ships in a bathtub each time the AI development cycle runs rampant. The stranger truth is that AI capabilities look like the jagged coastline of Maine. One must look very closely and see the tidal pools, the rocky peaks, and the caves of doom that might sink AI ambitions. Let’s look at the evidence. DeepMind going from a silver in 2024, with a system that needed experts to translate problems into code over days of compute, to a natural language gold in 2025 is impressive progress that we have witnessed in a short amount of time.
Further, there were also tests by OSWorld noted in the Stanford report, which challenges agents on real computer tasks across various operating systems. They reported agent task success jumped from roughly 12% to 66.3% in a year. A success rate of 66.3% doesn’t seem genius level, but this is notable, because that was within six percentage points of human performance that hovers around 72% in the same OSWorld type tests.
Fast forward to today. Per vendor-reported scores: Gemini 3.5 Flash now scores 78.4% on the OSWorld Verified benchmark, effectively tied with OpenAI’s GPT-5.5, and the leaderboard leaders have pushed past 83%, illustrating how far things have progressed lately. However, there is a new twist to this news that just landed in late June that is even more relevant. The same team behind OSWorld released OSWorld 2.0, a new version built from 108 realistic long workflows that take a skilled human about 1.6 hours each. On this new test version, the best frontier agent completed only 20.6% of these tasks end to end, and on workflows longer than about 2.7 hours, every model tested fell to zero. Ultimately, agents running on frontier models can ace the short demo tasks and collapse on the real workday hurdles. Another point for the human office crew.
However, as of the report’s March 2026 data, agents’ clock reading abilities are still stuck near the reliability of a 50/50 coin flip, highlighting this even more massive contrast. Frontier advancement moves fast, but it moves unevenly, and the gaps plaguing reliability of agents on work tasks are not where the intuition of business leaders might expect to find them.
The main problem here is that humans are plagued by the intuition trap. Business leaders assume intelligence is general, because we have this model in our minds of how other intelligent beings operate. It would be impossibly rare to meet a person who is brilliant at competition mathematics yet fails catastrophically at reading a clock, because within our western education system, the clock reading skill is considered foundational by age six or seven. This fact leads us to understand humans and AI systems acquire and organize their capabilities in profoundly different ways. Our mistake is applying this western education ideal to AI models that are more eccentric, and at times getting burned when they don’t deliver on tasks we expected them to demolish on their way to victory.
The obvious jump is to conclude that the training data must be mismatched somehow. Within the Stanford report, we read new evidence that this isn’t a training data problem. It does not fully prove that training-data coverage is irrelevant, but we shouldn’t be focusing on that as the main issue. Researchers fine-tuned models on 5,000 synthetic clock images, and the models improved on familiar clock styles but failed to generalize to real photos and unusual dials. Think Franck Muller watches with otherworldly faces. This is a perfect example of how AI applied to certain tasks in an enterprise environment that are seemingly simple to humans, and supported by plenty of training data, can still be botched.
According to the report, the real limitation for the AI we need to understand as business leaders lives in how the models combine multiple visual cues, not in what they’ve seen before in the training data. They have seen clocks, but when faced with the real world scenarios, it just isn’t enough to power through the muck of variability. The primary business implication from this, is that you cannot demo your way to success on the magic of AI hopes and dreams in 2026. As with any frontier technology, hype cycles will cause organizations to make major missteps in how they deploy that technology. Some will take risks and deploy too early into applications that are not fit for the models.
Leaders in 2026 will still see an impressive AI demo, extrapolate across whole workflows for the perceived efficiency gains and boost to their bonus, and get burned later when a task that looked trivial causes their workflows to implode under the weight of their outsized ambition. As a PM who has worked at several major US enterprises, I have seen this same failure when leaders are buying SaaS products for the magic of the demo instead of the sobering reality of the daily workflow earlier in my career. It’s a tale decades old at this point, and many leaders outside of the tech world seem to be more prone than ever. There is a peripheral danger especially to legacy sectors of the enterprise world.
This might still be feeling abstract so let’s bring in some real professional domains to ground this idea in the jagged frontier. This publication is focused on real world business implications, after all. The report shows models scoring 60% to 90% on evaluations in tax, mortgage processing, corporate finance, and legal reasoning, with the top 15 models separated by as little as 3 points. It flags these domains with need for high-reliability as a continuing challenge for AI development. It’s also important to remember these are benchmark scores, not measures of real end-to-end job performance. Many will dive head first into the zone they describe, good enough to tempt those who buy into the hype and too unreliable to leave alone unsupervised. This is exactly where most business automation decisions will live for the foreseeable future until additional advancements are made. You know the AI is good enough to be tempting, but nowhere near reliable enough to run unsupervised without humans seriously in the loop babysitting agents all day. Imagine reading that sentence in 1995.
Even better, it turns out this 60% to 90% range seems to be exactly where deployment decisions get made and where they go seriously wrong at times, because whether 85% accuracy is a triumph or a cortisol spiking lawsuit depends entirely on how workflows are designed around these agents. We should stop asking “is AI smart enough for us as an organization yet?”, and start asking “is this specific task inside or outside the jagged edge of AI capabilities right now and would we get wrecked if there is no human to intervene?”.
You need to be asking, is the task rigidly structured, text-heavy, and pattern-based to make it easily understandable by AI and the way this product is meant to be used? This would usually land inside the jagged edge and be a strong automation candidate, but only if a wrong output is easy to catch and cheap to reverse. On the opposite end of the spectrum, does the task require combining multiple cues or levels of artistic savant level judgment across messy context in your organization? This will mean it is often outside the jagged edge, like the clock research, and you need to pause and lock in for a moment to figure out if the risk is worth it.
Lastly and almost paramount, ask if a wrong answer under real world testing can be arrested by an intervening human in the loop monitoring the AI agent before it costs the company mega money. This is one of the prime cuts to carry away from this piece, in that this determines whether the 60% to 90% zone of accuracy mentioned above is acceptable for your organization and the current structure of your human in the loop situation. Don’t take our word for it. Test this on your own tasks in low stakes environments first.
While the gold medal math competition winner gets the headlines, the failure mode on the clock test is the best part of this embarrassing saga. When AI fails outside its frontier application that everybody raves about, it often fails huge and fails with boldness on seemingly normal tasks that even the elementary school child of your CTO would be able to answer. The leaders who win in 2026 will not be those who use AI everywhere in a shotgun blast across the hull of the entire organization, but those who understand exactly where its jagged edge creates leverage, exposes limits, and reshapes their operating model with the human talent they have available to deploy it with finesse.

