4 min read by Mahtab Chelani aka Milton

AI & AGI: Real Progress, Real Hype

A plain-English look at where AI actually stands in 2026, what is genuinely impressive, what is still hype, and why the smartest move is to learn how to work with it.

Artificial IntelligenceAGIAI AgentsTechnologySoftware

AI discourse right now seems to have two camps:

"AGI is here. Jobs are finished."

and

"It's all hype. Just autocomplete."

I don't think either is quite right.

The progress itself is very real.

AI systems can now write serious code, use computers, conduct research, work with tools, and complete increasingly complex multi-step tasks.

But that does not mean we suddenly have a flawless digital human.

What is actually real?

According to Stanford's 2026 AI Index, AI agents reached 66.3% accuracy on OSWorld, a benchmark that tests whether agents can perform real computer tasks across operating systems.

That is a huge jump from roughly 12%.

The same report also notes that Gemini Deep Think achieved gold-medal-level performance at the 2025 International Mathematical Olympiad.

So yes, AI has become extremely capable.

But here is where things get interesting.

The same Stanford report says the strongest model on ClockBench could read analogue clocks correctly only 50.6% of the time, compared with 90.1% for humans.

In other words:

AI can solve competition-level mathematics and still struggle with a clock.

Researchers often describe this as jagged intelligence.

It is incredibly strong in some directions and surprisingly weak in others.

Source: Stanford AI Index 2026 - Technical Performance

So, do we have AGI?

Not in any broadly agreed scientific sense.

A useful reality check comes from ARC-AGI-3, a benchmark designed around unfamiliar interactive environments.

The agent is not given the rules or even told what "winning" means. It has to explore, learn, adapt, and figure the environment out for itself.

At launch in March 2026:

  • Humans scored 100%
  • Frontier AI scored 0.51%

That is a pretty large gap.

Source: ARC Prize - Announcing ARC-AGI-3

This does not mean today's models are weak.

It means intelligence is more complicated than getting a high score on coding, maths, or knowledge benchmarks.

The bigger change is AI agents

The most important development may not be "smarter chatbots."

It is the shift from:

question -> answer

to something closer to:

goal -> inspect -> use tools -> act -> test -> retry -> deliver

That is what makes modern AI agents interesting.

They can increasingly operate software, write and test code, search through information, and complete longer chains of work.

METR has been tracking how long and difficult a task frontier agents can complete reliably.

Their results show rapidly increasing capability in areas such as software engineering, machine learning, and cybersecurity.

But METR also makes an important warning:

An AI completing an "8-hour task" in a benchmark does not mean it can replace eight hours of a professional's actual job.

Real work involves prior context, people, unclear requirements, judgment, office politics, changing goals, and success criteria that are often impossible to score neatly.

Source: METR - Task-Completion Time Horizons of Frontier AI Models

Is AI replacing people or helping them?

Right now, the picture is more mixed than the headlines suggest.

Anthropic's January 2026 Economic Index found that on Claude.ai:

  • 52% of conversations were classified as augmentation
  • 45% were classified as automation

"Augmentation" means people were working with AI, iterating, learning, validating, or using it collaboratively.

That does not mean automation is unimportant. In fact, Anthropic notes that automation has increased over time.

But today's usage still looks a lot like:

Human + AI

rather than:

AI alone

Source: Anthropic Economic Index - January 2026 Report

One development worth watching

OpenAI recently said it has reached what it calls an "automated research intern".

By that, OpenAI means a system capable of carrying out well-defined research tasks under human direction, including tasks that could take a skilled researcher a few days.

People still set priorities, decide what matters, judge results, and determine what gets deployed.

So this is not autonomous AGI.

But it is an important signal.

If AI increasingly helps researchers improve AI itself, the feedback loop becomes much more interesting.

Source: OpenAI - Research Acceleration: The View Inside OpenAI

So, is the AGI hype real?

My takeaway is simple:

AI replacing humanity tomorrow? Probably hype.

AI fundamentally changing software, research, business, and knowledge work? That part is already happening.

The mistake is treating this as a binary question.

We do not need to have "AGI" for AI to dramatically change how work gets done.

And we do not need to pretend current systems are human-level at everything just because they are extraordinary at some things.

The most useful question probably is not:

"Will AI replace me?"

It is:

"How much more capable can I become by learning how to work with it?"

That feels like the more practical way to look at what is happening right now.

Sources