Crypto news

23.07.2026
16:12

Grok 4.5 crushed everyone: The next-generation AI assistant takes over the browser

Developers are creating AI assistants capable of independently visiting websites and performing actions on behalf of a person: opening pages, clicking buttons, filling out forms, and gathering information. You can ask such an assistant to book a ticket or collect data from a website — and it will do it on its own.

The tool based on the Grok model handles this task better than all others. This conclusion was reached by Gregor Zunic, founder of the Browser Use service.

Analysts compared five tools based on one key criterion — how often the AI completes a task on a website without errors. The higher the column on the chart, the more reliably the tool works.

The best result is from Grok — 76.42 percent of tasks completed correctly. It is followed by the assistant from OpenAI (72.70 percent) and two assistants from Anthropic (70.75 and 67.92 percent). In last place is GPT via the OpenClaw plugin with a result of 60.38 percent.

Each assistant runs on its own AI model. Grok uses the Grok 4.5 model, OpenAI assistants use the GPT 5.5 model, and Anthropic assistants use the Opus 4.8 model. The model is the actual artificial intelligence that thinks and makes decisions. The assistant adds the ability to control the browser.

From the comparison, it is clear that under identical conditions, the choice of AI model plays a major role, and here Grok 4.5 outperforms its competitors. The method of connecting to the browser is also important — for the same assistant from Anthropic, the result differed by almost three points depending on how it was connected to the website.

What to Consider

The market for agentic AI is not just a race of models, but a competition of ecosystems. Grok 4.5 shows that integration with the browser at the level of a native tool can give an edge even to the most powerful competitor models. While OpenAI and Anthropic rely on third-party plugins, xAI has built a vertical solution — and it pays off. However, it is worth remembering that the tests were conducted in a controlled environment. In real-world scenarios with dynamic pages, captchas, and JavaScript load, the gap may narrow.