Gemini 3 Shows How Far AI Has Come: From Chatbots to Digital Coworkers
I have been testing Google’s latest Gemini 3 model, and my conclusion is simple:
It is very impressive.
But rather than showing you another list of benchmark scores, I want to demonstrate what has actually changed by asking the AI itself to show us.
It has been less than three years since ChatGPT launched and introduced generative AI to the mainstream. Just days before that release, I published my first article about OpenAI’s earlier GPT-3 model.
When ChatGPT arrived, I wrote:
“I am usually pretty hesitant to make technology predictions, but I think this is going to change our world much sooner than we expect, and much more drastically. Rather than automating jobs that are repetitive and dangerous, there is now the prospect that the first jobs disrupted by AI will be more analytical, creative, and involve more writing and communication.”
Looking back, that prediction appears to have been largely correct.
The most surprising part is not simply that AI became better at generating text.
The real transformation is that AI systems are beginning to perform meaningful intellectual work.
Three Years Ago, AI Could Write. Today, It Can Build.
I could explain the difference between the original ChatGPT and Google’s Gemini 3.
But I don’t need to.
Instead, I gave Gemini 3 a screenshot of my original 2022 article and asked:
“Show how far AI has come since this post by doing things.”
Gemini understood the challenge immediately.
Its response was not a paragraph explaining AI progress.
It built something.
It created a fully interactive, playable simulation game:
A Candy-Powered FTL Starship Simulator.
The original 2022 AI could describe a fictional spaceship powered by candy and escaping from otters.
The 2026 AI could:
- Design the concept
- Write the code
- Build the interface
- Create the game mechanics
- Generate the content
- Let me actually play it
The difference is not that AI became better at writing.
The difference is that AI became capable of making things.
And that distinction changes everything.
Coding Tools Are Becoming General-Purpose AI Workspaces
Alongside Gemini 3, Google introduced Antigravity, a new AI development environment.
For programmers, Antigravity will look familiar. It belongs to the same emerging category as tools like Claude Code and OpenAI Codex: AI systems that can access a computer, write software, execute tasks, and work autonomously with guidance.
But if you are not a programmer, you might assume these tools are irrelevant.
That would be a mistake.
The reason is simple:
Coding is becoming the universal language of computers.
Almost everything we do digitally is ultimately controlled by software.
If AI can understand and manipulate code, it can increasingly perform almost any computer-based task:
- Build websites
- Create dashboards
- Analyze files
- Generate presentations
- Automate workflows
- Manage digital tools
The future importance of coding agents is not that everyone will become a programmer.
It is that programming is becoming a gateway for AI to operate the digital world.
From Prompting AI to Managing AI
One of the most interesting parts of Antigravity is that you do not interact with AI agents by writing code.
You simply describe what you want.
You communicate in natural language.
The AI decides how to accomplish the task.
For example, I gave Gemini 3 access to a folder containing years of my newsletter posts and asked:
“Create an attractive website showing my AI predictions, then research online to determine which predictions were correct and which were wrong.”
The AI did not immediately start generating text.
Instead, it:
- Read through hundreds of files.
- Identified relevant predictions.
- Created a plan.
- Asked for approval.
- Conducted research.
- Built the website.
- Tested the result.
- Prepared the final output.
The important part was not that everything was perfect.
It was not.
There were still judgment errors and places where I needed to provide feedback.
But the experience felt fundamentally different from using a chatbot.
I was not asking questions and correcting mistakes.
I was managing a capable collaborator.
Is Gemini 3 PhD-Level Intelligence?
The phrase “PhD-level intelligence” has become common in AI discussions.
But what does it actually mean?
I decided to test it.
I gave Gemini 3 access to a collection of old research files from my academic work on crowdfunding. The files were messy:
- Old spreadsheets
- Statistical files
- Unorganized datasets
- Outdated formats
I asked the AI to understand the data, clean it, and prepare it for new analysis.
It successfully reconstructed the structure of the dataset and prepared it for research.
Then I gave it a much harder assignment:
Write an original academic paper using this data. Conduct deep research, develop an important theoretical question, perform sophisticated analysis, and write it as if for an academic journal.
I provided almost no additional guidance.
The AI:
- Studied the dataset
- Generated research questions
- Developed hypotheses
- Performed statistical analysis
- Created new measurements
- Reviewed academic literature
- Produced a formatted research paper
The result was a 14-page academic-style paper.
The AI Still Has Human-Like Weaknesses
Does this mean Gemini 3 is equivalent to a PhD researcher?
The answer is complicated.
In some ways, yes.
If we define PhD-level intelligence as the ability to perform competent graduate-level research work, Gemini 3 is approaching that capability.
But it also displayed weaknesses similar to human researchers.
The paper had:
- Interesting ideas
- Strong technical execution
- Creative approaches
But it also contained:
- Statistical methods that could be improved
- Conclusions that went beyond the evidence
- Theoretical arguments that needed refinement
These were no longer basic AI failures.
They were closer to the mistakes made by capable but inexperienced researchers.
And when I provided feedback, the AI improved significantly.
The lesson is important:
The future is not an AI that works perfectly without humans.
The future is humans guiding AI systems that are already highly capable.
The End of the Chatbot Era
Gemini 3 represents something much larger than another model release.
It reflects several major shifts happening simultaneously:
- AI models continue improving rapidly.
- AI agents are becoming more capable.
- Computer-use systems are expanding.
- AI is moving from answering questions to completing tasks.
- Human interaction with AI is changing.
Three years ago, we were amazed that AI could write a poem about an otter.
Today, we are discussing research methodology with AI systems that can build their own analysis environment.
The transformation is enormous.
But the biggest change may not be the intelligence of AI itself.
It is the role humans play.
The old model was:
Human asks → AI answers → Human fixes mistakes.
The emerging model is:
Human defines goals → AI works → Human directs and evaluates.
The “human in the loop” is evolving.
We are moving away from humans acting as AI error correctors.
We are moving toward humans becoming AI managers.
And that may be the most important change since ChatGPT first arrived.

