If you feel like everything in AI is moving faster than ever, you are probably right.
The pace of progress in frontier AI models has accelerated dramatically. Leading AI labs are releasing increasingly capable systems at a speed that would have seemed impossible only a few years ago.
But the acceleration is not just about release schedules.
The more important story is that AI capability itself appears to be improving faster than many expected.
Of course, the frontier remains uneven. AI systems are still surprisingly weak in certain areas, and impressive benchmark results do not always translate into flawless real-world performance.
But when we look at what AI systems can actually accomplish in professional environments, the trend becomes much clearer.
AI is not just answering questions better.
It is starting to complete meaningful pieces of work.
Measuring the Rise of AI Capability
One of the biggest challenges in understanding AI progress is finding meaningful ways to measure it.
Traditional benchmarks often capture narrow abilities — answering questions, solving problems, or performing specific tasks.
But the real question is different:
How much human work can an AI system actually perform?
Several organizations have started building evaluations designed around this question.
Two of the most widely discussed examples are assessments from:
- METR
- The UK AI Security Institute
These evaluations attempt to estimate how many hours of professional human work an AI system can complete from a single prompt.
Another important evaluation, GDPval, compares AI performance against human experts across a wide range of professional fields, using expert judges to evaluate results.
Across these evaluations, one pattern is becoming increasingly clear:
AI systems are improving at an accelerating pace.
Their ability to complete longer, more complex tasks is increasing rapidly.
From Hours of Work to Days of Work
Another organization studying AI progress, Epoch, recently tested how much independent work advanced AI systems could complete.
In one experiment, Opus 4.7 was able to work autonomously for approximately 14 hours and produce a software package that would typically require between 2 and 17 weeks of human engineering effort.
The computational cost was relatively small compared with the amount of work completed — around $251 worth of tokens.
Of course, these systems are not perfect.
They still fail tasks. They still require supervision. And running powerful models at scale is not always inexpensive.
But the direction is unmistakable:
AI systems are becoming capable of performing longer, more independent, and more valuable tasks.
In my own experiments, I found that Fable was able to work autonomously for around nine hours, completing complex software projects that would previously have required a team of engineers working for more than a week.
The important shift is not simply that AI is getting smarter.
It is that AI is gaining endurance.
Two AI Frontiers: Closed Models and Open Models
So far, I have focused mainly on frontier models — the systems with the highest levels of capability.
Today, the leading frontier models largely come from three American AI companies:
- Anthropic
- OpenAI
These companies continue to push the limits of proprietary AI systems.
But there is another important category of models developing alongside them:
Open-weight models.
Unlike closed frontier models, open-weight systems allow researchers, developers, and companies to download, modify, and deploy them independently.
Most of the strongest open-weight models currently come from China, and they typically trail the leading closed models by several months.
However, they are improving rapidly as well.
They are following their own steep improvement curve.
My evaluations using tests such as AA-Briefcase — which simulates a complex consulting engagement requiring multiple forms of analysis — show that open-weight models are advancing quickly, even while remaining behind the strongest proprietary systems.
The gap still exists.
But the distance is changing.
Benchmarks Are Not Enough
However, abstract performance charts only tell part of the story.
They can hide one of the most important characteristics of modern AI:
The frontier remains extremely uneven.
A model can achieve excellent benchmark scores while still struggling with practical judgment, creativity, reliability, or long-term execution.
This is why real-world experimentation matters.
The best way to understand AI is not only to look at scores.
It is to actually use these systems for meaningful tasks and evaluate how well they perform in situations that matter.
For example, I created a test where AI systems were asked to build an interactive simulation showing the evolution of a harbor over thousands of years.
The results revealed differences that traditional benchmarks often miss.
The models differed not only in technical ability, but also in:
- design decisions;
- creative judgment;
- understanding of user experience;
- ability to maintain coherence over long tasks.
As AI systems continue taking on longer assignments, these harder-to-measure qualities will become increasingly important.
The Way We Use AI Is Changing
The biggest transformation happening in AI may not be the models themselves.
It may be the way humans interact with them.
Until recently, the dominant AI workflow was what I call co-intelligence.
A human would:
- Ask the AI to complete a task;
- Review the output;
- Provide feedback;
- Guide the AI through the next step.
The human remained the project manager, while the AI acted as a powerful assistant.
This approach remains extremely useful.
But increasingly, it is no longer the most advanced way to use AI.
From Chatbots to Agents
The next phase of AI is built around autonomous agents.
Agents differ from traditional chatbots because they are connected to:
- tools;
- software environments;
- files;
- workflows;
- external systems.
They do not simply generate responses.
They take actions.
This requires a new technological layer: the AI harness.
A good harness gives AI access to the environment where work actually happens.
Examples include:
- Claude Code;
- Claude Cowork;
- OpenAI Codex.
These systems allow AI to operate more like a digital employee than a conversational assistant.
The result is a fundamental change in how work gets organized.
Instead of collaborating with chatbots, people increasingly delegate tasks to agents.
The Rise of the AI Manager
A joint study from OpenAI and academic researchers provides an early look at how this transition is happening inside AI companies themselves.
The interesting finding is that AI agents are not only being used by programmers.
Legal teams, HR departments, and other non-technical groups are adopting agents at similar rates.
OpenAI may represent an early preview of what many workplaces will eventually look like.
Increasingly, work inside these organizations resembles AI management.
Many employees are no longer simply completing tasks themselves.
They are managing multiple AI systems working alongside them.
A significant portion of OpenAI employees regularly operate multiple agents simultaneously.
As coding becomes increasingly automated through specialized AI environments, many professionals are becoming something closer to “AI operators” — people who understand their domain while directing intelligent systems to execute work.
Expertise Matters More Than Technical Background
Interestingly, research on Claude Code users reveals something important:
Success with AI agents is not determined primarily by whether someone is a programmer.
The biggest factor is domain expertise.
People who deeply understand their field are better at guiding AI systems within that field.
They know:
- what questions to ask;
- what mistakes to watch for;
- what results actually matter.
Even more interestingly, users with stronger expertise tend to extract more value from each interaction.
The future may not belong to people who know the most about AI.
It may belong to people who know their own field deeply and can effectively manage AI systems within it.
Living Through an Exponential Moment
The defining characteristic of exponential growth is that every period of improvement is larger than the previous one.
If your organization created an AI strategy before the winter of 2025, you were probably planning around systems capable of completing a few hours of work with significant limitations.
Only months later, some systems can complete sixteen hours or more of work from a single prompt.
This is why AI progress feels so dramatic.
A smooth curve on a graph becomes a series of sudden shocks in real life.
Humans are not naturally good at understanding exponential change.
And right now, we are living inside one.
Why AI Feels So Disruptive
I think this explains much of the turbulence surrounding AI better than simple explanations about hype.
AI does not gradually become a cybersecurity concern.
It appears harmless — until suddenly it reaches a capability threshold where the risks become obvious.
Markets do not slowly adjust to AI disruption.
They move sharply when investors realize that AI may fundamentally change a business model.
These sudden shifts are often interpreted as signs that AI is an immature technology that will eventually stabilize.
I am not convinced.
The instability is not necessarily a temporary phase.
It may simply be the result of institutions moving at human speed trying to keep pace with a technology improving on an exponential curve.
As long as AI capability continues accelerating, the gap between technological progress and human adaptation will continue to grow.
The AI revolution is not slowing down.
We are still learning how to live inside it.

