Jev Ultrafast

A screenshot-free browser agent for developers who need web automation to finish in seconds

Fastest and cheapest web agent

Category
AI agents & assistants
Audience
Developers
Language
Python
Licence
MIT

Updated

Jev Ultrafast is a Python browser agent that drives web pages through a dynamic, indexed action space instead of screenshots. It comes from the Browser Use project, built around TypeSafe's Jev model, and it is aimed at developers who already write browser automation and who care about how long each step takes and how many model and browser calls it burns. If you have watched an agent screenshot a page, think for a second, then click once, this is the same job arranged differently.

What it does

You give it one goal in natural language. On every observation it builds a table of the elements currently on the page, then decides one operation and one target from that table. The operation vocabulary is fixed and small: CLICK, TYPE_TEXT, SELECT, SCROLL_UP, SCROLL_DOWN, WAIT, DONE and BLOCKED. Free text is generated only when the chosen operation is TYPE_TEXT, and the README says a small LLM handles that part.

The demo in the repository is a real Google Flights search, Zurich to London, done in 7.1 seconds from a single natural-language goal — with the city names actually generated and page loading waits included, not edited out. The project ships the demo as both a GIF and an MP4, a measurements document, and the agent loop itself as readable source.

It is MIT licensed and, at the time of writing, has gathered over twenty-one thousand stars and roughly fifteen hundred forks within about two weeks of its first commit.

How it works

The element table is the whole idea. Each observation produces numbered rows with a role, a label and the current value, something like:

[1] button    Change ticket type · Round trip
[2] combobox  Where from?        · San Francisco
[3] combobox  Where to?          · empty
[4] textbox   Departure          · empty

One request to the model comes back with the operation plus the candidate targets — a click target, a type-text target, and a select target when a select is present — and the runtime uses whichever target matches the operation it picked. So CLICK [7] and TYPE_TEXT [3] are both resolved from the same single round trip, rather than one call to decide what to do and another to decide where.

Two consequences follow. First, the agent never has to describe a picture of the page to itself, which is where most of the latency and token cost in screenshot-driven agents lives. Second, the action space is rebuilt from the live page every step, and only supported operations and targets are offered, so a click always names an element that genuinely exists on the page right now instead of a guessed coordinate.

The numbers the channel quotes come from the project's own measurements: across six runs the median time dropped by twenty-five percent, and browser calls fell from over a thousand to about a hundred. The loop is in jev_ultrafast/agent.py and the measurements in docs/performance.md, so both claims are checkable rather than taken on faith.

When to use it / when not

This fits tasks that live in ordinary page structure: searches, forms, booking flows, filters, anything where the thing you need to touch appears as a button, combobox or textbox with a readable label. The faster the loop, the more attempts you can afford, and a hundred browser calls instead of a thousand matters if you are paying per session.

It fits less well where the information you need exists only in rendered pixels — a chart, a canvas drawing, an image with text baked into it — because there is no screenshot in the loop to look at. Pages that hide their real controls behind custom widgets without usable roles or labels are the same problem in a different shape.

The other thing to weigh is age. The repository was created in mid-September 2026 and last pushed days later. The star count is enormous for that window, but it tells you about attention, not about how the agent behaves on the twentieth site you point it at. The README also advertises a waitlist for a hosted cloud version, which means the managed path is not open yet — what you can run today is the open-source loop.

Alternatives

The honest comparison is the mainstream browser agent that takes a screenshot each step and reasons over the image; that is the baseline this project measures itself against, and it remains the more general approach when visual understanding is genuinely required. The wider Browser Use project, from the same authors, is the fuller-featured sibling. Jev Ultrafast is the narrow, fast variant: fewer ways to act, chosen faster.

Take this one seriously if you run browser automation at any volume and your bill or your wall-clock time has become the limiting factor. The design is a single clear trade — give up the screenshot, gain a structured action space and one model call per step — and the repository is small and readable enough that you can judge that trade yourself in an afternoon. Treat the published timings as a starting point to reproduce on your own sites rather than a guarantee, and treat the cloud offering as not yet here.

More in AI agents & assistants

All of AI agents & assistants →