Blog Detail Top Background
Kuan Zou

Humans Search with Google. Machines Search with Octen.

For twenty-five years, the user of search has been a person. That premise is changing. The next user of search works at machine speed, takes in information at machine scale, and consumes search on a machine cost structure.

Kuan ZouKuan Zou
LinkedInX
/Founder & CEO of Octen AI/Sep 21, 2026
AIAgentSearch

When we started Octen ten months ago, we had one simple conviction: search infrastructure built for humans cannot carry the AI era.

For twenty-five years, the user of search has been a person. Someone opens a search box, types a question, scans a page of results, clicks a few links, reads, and then decides what to do next. Nearly every major search engine today is built on assumptions drawn from that interaction.

How many searches a user will make, how fast results need to come back, how content is ranked and displayed, how much compute each request deserves, even what counts as "a good search result." Every one of those answers rests on the same premise: search is for people.

But the user of search is changing. It is no longer a person. It is a model, an agent, a device, and soon a robot. And machines search the world in a fundamentally different way.

AI doesn't search like a human

To complete a task, a person might search a few times, one query each. For an AI system, the same task can generate hundreds or thousands of searches, and they can all happen at once.

That changes both the engineering model and the economics of search.

Latency compounds.

When search is one step in a reasoning chain, every extra 100 milliseconds stacks across multiple queries, multiple agents, and multiple rounds of iteration. Latency stops being a UX metric. It directly determines how fast the entire AI system can think and act.

Concurrency goes from edge case to default.

People search serially. Machines have no such constraint. A single agent can explore dozens of hypotheses at once, and many agents can search simultaneously. The better AI gets at thinking in parallel, the higher the concurrency it demands.

Cost goes from an API pricing question to an infrastructure economics question.

A person searches a limited number of times per day, so a slightly expensive retrieval doesn't matter much. But if an AI needs hundreds of searches to finish one task, and a product generates millions or billions of searches a day, the same cost structure simply doesn't hold. Every marginal cent gets amplified at machine scale until it can no longer be ignored.

The definition of "search quality" changes too.

Faced with ten blue links, a person can judge which results are bad, filter out ads and visual noise, parse the page structure, and find what actually matters. A machine needs information it can consume directly: relevant, fresh, clean, well-structured, reliably grounded, and complete enough in context to support the next step of reasoning.

So Machine Search is not an API wrapped around a human search engine. It is an entirely different workload.

You can't build Machine Search with Human Search thinking

After generative AI took off, the most natural first step was to plug large models into existing search systems. As a starting point, that made sense. But from the beginning, we believed it wouldn't be the end state, because the underlying architecture of most search engines today was built on the assumptions of the human internet: a bounded number of queries, largely serial search behavior, result pages and ranking logic designed for human reading.

AI has changed almost every one of those assumptions.

So instead of asking "how do we make existing search work better for AI," we asked: if search were designed for machines from day one, what would it look like?

That question took us down a completely different path. We rebuilt search around what machines actually care about: extremely low latency, massive concurrency, real-time freshness, information machines can consume directly, retrieval quality, and a cost structure that still holds under enormous volumes of machine requests.

It also means we have never treated speed, cost, and quality as three separate product metrics. To us, they have been a single systems engineering problem from day one.

Speed is not a feature

When people talk about search latency, they usually think of "faster" as a better experience. For AI, speed doesn't determine the experience. It determines what kind of AI products can be built at all.

An agent doesn't need to ask questions one at a time the way a person does. It can fan out across a dozen directions at once, and each of those directions can branch into new searches. If retrieval is fast enough, the agent can explore the information space aggressively and still finish the task in seconds.

If retrieval is slow, developers are forced to hold it back: fewer searches, parallel work pushed back into sequence, heavy reliance on caching, models reasoning on incomplete information, research depth traded for response time.

At that point, search infrastructure is effectively setting the agent's intelligence budget.

That is why we have been obsessive about latency from day one. Octen's internal server-side median latency is currently around 30ms, and the system is designed for 1M+ concurrent queries.

But the goal was never to be a few dozen milliseconds faster on a benchmark. It was to make retrieval stop being a scarce resource. Every call used to cost latency and money, so every agent carried an invisible search budget.

Once that budget is lifted, the optimal behavior of AI changes. The question used to be "do I really need to search?" The question becomes "why wouldn't I search again?"

Speed, cost, and quality were never a trade-off

People have long understood search through trade-offs: go faster and quality may drop; want higher quality and you need more compute, so cost goes up; push cost down and performance suffers.

We never believed these trade-offs were laws of physics. Most so-called trade-offs are just the consequences of architectural choices.

If the underlying retrieval system is designed for machine workloads from the start, then every improvement at the system level can move multiple dimensions at once. Better indexing and ranking find more relevant information with less work. Removing unnecessary computation saves both time and money. True parallelism lets machines explore more without linearly increasing end-to-end latency. And infrastructure designed around machine scale has fundamentally different unit economics from human queries.

So being the best on a single dimension isn't enough. Machine Search has to solve speed, cost, and quality at the same time.

Ten months later: SOTA that actually means something in production

Today, Artificial Analysis released its independent Search API Benchmark.

Among all major search providers, Octen ranks first in speed, second in cost efficiency, and top three in search quality. It is the only product in the top three on all three core dimensions, and the only provider in the "Most Attractive Quadrant" when quality and cost are considered together.

I want to call out the part of this result that is easiest to miss: it comes from a single product, in its default configuration.

Leaderboards can be gamed by splitting. A provider can submit multiple APIs, multiple tiers, multiple parameter sets, and let one chase speed, another chase quality, a third chase cost, leaving a nice-looking name in every column.

Occupying three slots on a leaderboard and building one product that holds up on all three dimensions are two completely different things.

The first is a submission strategy. The second is an engineering problem.

Because a product can't be used in pieces.

In production, a developer integrates one endpoint. That endpoint is either fast enough, cheap enough, and accurate enough at the same time, or it doesn't work. Three first-place finishes spread across three configurations are worth nothing to a running agent. It can't use all three in the same call.

So from the beginning, we have only done one thing: make the same product, in the same default configuration, hold up on all three dimensions at once. It is a clearly harder goal, and the only one aligned with the reality developers actually face.

What matters to us isn't any individual ranking. It's that these results are starting to validate the conviction we had ten months ago:

Speed doesn't have to come at the expense of quality.

High quality doesn't mean infrastructure has to be expensive.

Lowering cost doesn't mean sacrificing the search experience.

The frontier of speed, cost, and quality can move forward together. And only when all three hold in the same product does SOTA actually mean anything in production.

The next bottleneck is the search paradigm itself

This benchmark measures single searches, and multiple searches within a task. In both cases, the searches are orchestrated by the agent itself: it decides how many queries to send, in what order, when to wait, and how to stitch the results back together.

In other words, what's being measured is always "one API called many times." And those calls often have no real dependency on one another.

When an agent researches a company, it wants the whole picture: products, customers, competitors, funding history, latest news, open roles. These dimensions are independent of each other and should be fetched simultaneously. Only when they come back together can the agent reason over all of them in a single step, instead of fetching one dimension, reasoning one step, and going back for the next.

So the default behavior should be: one request, multiple search tasks fanning out in parallel, and the richer and more accurate the returned information, the better. For a machine, that is what efficient search actually means. Not asking one question faster, but getting an entire matter answered as fully as possible within the same window of time.

Search should become abundant, not just in the sense that you can afford to search more often, but that a single request brings back enough.

And that orchestration logic shouldn't live in the caller's code in the first place. Which searches can be sent together, which indexes can be reused, which results are actually duplicates, which can be returned early: the retrieval system knows all of this far better than the application above it. Pushing it all onto developers is the wrong division of labor.

Measured by this standard, many of the names that live for leaderboard rankings start to fall apart. Take a browser agent that claims "zero cost" and takes thirty-plus seconds per run: a thousand searches is more than eight hours. And that cost advantage is on paper only. It doesn't count the developer's time, it doesn't count the searches that have to be redone when quality falls short, and it simply bills the cost of search under a different name. That's wordplay, not a technical cost advantage.

Real engineering capability comes from redesigning the shape of the request itself. Over the past few months, the vast majority of our engineering investment has gone here.

More on that soon.

Search is becoming the invisible infrastructure of AI

The first generation of internet search was a destination. You opened Google, typed a question, and clicked Search.

Machine Search is entirely different. It will gradually disappear behind every AI product. A coding agent pulls documentation while it writes code. A financial agent checks markets, filings, and company data before it analyzes. A shopping agent continuously verifies inventory and reviews while comparing prices. AI devices query the world in real time based on what they see and hear. Robots need information beyond their own models before deciding what to do next.

In these scenarios, a person may never see a search box. But out of sight, the volume of search will grow by orders of magnitude.

This is the paradox of search in the AI era: it becomes increasingly invisible to humans, and increasingly essential to machines.

And as models get stronger, they won't search less. They'll search more. Stronger reasoning generates more questions that need verifying. More agents generate more parallel information needs. Continuously running systems need to keep refreshing their understanding, because the world doesn't stop changing when training ends, and it doesn't stop changing when the first retrieval completes.

Whenever an AI needs to know what's happening beyond its weights and its context window, it needs a way to perceive the outside world.

Retrieval is the perception layer of AI.

For twenty-five years, the search industry was built around humans, because humans were the users of search.

That premise is changing.

The next generation of search infrastructure will serve a completely different set of users. They work at machine speed, acquire information at machine scale, and consume search at machine cost structures.

They are models, agents, devices, and the Physical AI of the future.

That is the world we are building Octen for.

In ten months, we've proven that speed, cost, and quality can hold at the same time. What comes next is putting that capability in the hands of the developers who have been holding their AI back, only because search couldn't keep up.

For twenty-five years, humans learned how to search the internet.

Now, machines are starting to search the world.

Humans search with Google.

Machines search with Octen.

Kuan (Colin) Zou
Kuan (Colin) ZouLinkedInX

Founder & CEO of Octen AI