Each day, I dedicate time to staying up-to-date with the latest developments in AI, whether it’s through newly published papers, launch announcements, or tracking the ongoing discussions across various topics. This week, however, marked a major milestone in the field for different reasons.

OpenAI introduced an agent that highlights a shift from generative AI to agentic AI when they unveiled their first AI agent, Operator, while China based DeepSeek-R1 launched a high performing open-source large language model. DeepSeek uses model distillation techniques to create smaller, more efficient models that replicate the behavior of larger, more complex systems and can be a key unlock edge based AI. Both developments represent important strides in the evolution of AI technology.

Recently, the shift from generative AI to Agentic AI is the main topic I am often asked to speak about. For those not familiar with the phrase, agentic AI refers to intelligent systems capable of autonomously performing tasks, making decisions, and interacting with their environments in much the same way humans do.

By day, I work with large CPG organizations. We discuss all aspects of AI-ready data, scalability through simplicity, and how to map AI agents to connect complex ecosystems across business functions to automate tasks by leveraging LLMs, systemic prompting, and various types of organizational data. This shift paves the way for AI that can actively drive operations rather than just assist with insights generation—essentially shifting from insight to action.

One subset of Agentic AI are specialized agents. Specialized agents can take on many forms such as workflow, voice and browser agents. Browser agents, are designed to handle specific tasks like web scraping, information retrieval, and automating online interactions. Open AI’s Operator builds on this concept by combining these specialized abilities into a more versatile, multimodal agent, capable of performing a broader range of tasks across various environments. This evolution marks a significant leap in the adaptability and functionality of agentic AI.

The Operator platform is powered by Computer-Using Agent (CUA), Operator can autonomously navigate digital spaces and perform browser specific tasks on behalf of users. The model combines GPT-4o’s vision capabilities with advanced reasoning through reinforcement learning. This allows the agent to interact with GUIs (graphical user interfaces)—everything from buttons to text fields—just as a human would, without the need for specialized APIs.

In this example, Operator takes an image of a dish and the prompt instructs to purchase the ingredients while consideringt the best price.

This combination of multimodal understanding and problem-solving marks a revolutionary step forward in AI development, as CUA can now break tasks into multi-step actions and adapt dynamically to changing online environments. It can still get stuck and require human assistance, but the initial results have been fascinating to review.

According to OpenAI, CUA processes raw pixel data from screens, using a virtual mouse and keyboard to complete tasks. It operates through an iterative loop that integrates three crucial components:

  1. Perception: CUA analyzes screenshots from the computer to understand the current state of the system.
  2. Reasoning: CUA employs chain-of-thought reasoning, considering past and present observations to determine the next logical steps.
  3. Action: It performs actions—clicking, typing, scrolling—based on its reasoning, while ensuring user confirmation for sensitive tasks, such as entering login details or responding to CAPTCHAs.

The beauty of this process lies in its flexibility and ability to perform tasks without relying on specific OS or web APIs. This opens the door for a new range of applications, such as automatically filling out forms, navigating websites, and even handling unexpected changes in tasks. This alone is an incredibly exciting development.

On the other side of the world, China’s DeepSeek recently launched Deepseek-R1. DeepSeek-R1 is a powerful language model that offers several key advantages. Its strong reasoning capabilities allow it to handle complex problems and provide detailed explanations, with a significantly more cost-effective model that initially benchmarks in the range of large closed models.

What really caught my attention is that DeepSeek-R1 boasts open-source availability. DeepSeek’s release highlighted the accelerating gap between open-source and closed AI. Just six months ago, open-source models seemed far behind their commercial counterparts, but DeepSeek-R1 is causing some closed AI organizations to evaluate and review their approach to model distillation and how they are generating performance benchmarks similar to much larger models.

DeepSeek originally released V3 in December 2024. This is a large model that R1 is based on. It is comprised of 671 billion parameters, but it uses a mixture of experts model. It doesn’t use all parameters at the same time, it only uses 37 billion activations per token. This is key to the efficiency gained.

DeepSeek V3 used 5% of the training time of GPT4, an efficiency of 95%. With V3 as a foundation, R1 uses it as its underlying model. R1 went through a new fine tuning method via unsupervised reinforcement learning. R1 uses chain of thought prompting at time of inference. When you put a prompt in, you will see it going through a reasoning process prior to a final answer.

Deepseek's combination of open source availability and performance have caught a lot of attention

Some have already begun experimenting with DeepSeek + Browser Use, a popular GitHub project that creates an easy way to connect AI Agents with a browser. In this case, connecting DeepSeek-R1 to act in a manner similar to OpenAI’s Operator.

The release of DeepSeek-R1 and smaller distilled models could signify a new chapter in how AI evolves. We are seeing the rise of highly intelligent AI that can now run locally, even on personal devices like smartphones.

Distilled models are becoming more powerful, and AI is no longer confined to centralized servers—it’s moving to the edge, enabling real-time, personalized, and local AI experiences. The future is not just about cloud-based AI, but about edge-based AI that is capable, efficient, and seamlessly integrated into our daily lives.

UPDATE 1.28.25

I was curious what impact this news would have when I originally wrote this post. Since publishing there have been movements in the markets. What I would say is similar to what Satya Nadella, CEO of Microsoft also mentioned and that is Jevon’s Paradox applies here.

The Jevons Paradox describes the situation where technological advancements that increase the efficiency of a resource (like GPU usage for model training) can paradoxically lead to increased consumption of that resource (for inference).

With the decreased emphasis on GPU’s for training, there will be an increased demand for GPUs for inference that will actually create an increase in demand. More efficient models… new applications… increased usage…



Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Trending

Discover more from Tom Edwards AI Keynote Speaker - EY AI Leader - BlackFin360 Blog Thought Leadership

Subscribe now to keep reading and get access to the full archive.

Continue reading