Google Custom Search API Is Shutting Down: Your Migration Options
Google Custom Search JSON API is shutting down at the beginning of 2027. Discover your migration options and how to integrate Octen for AI-native web retrieval.

If the Google Custom Search API is in your production stack, it has a fixed deadline. Google closed the Custom Search JSON API to new customers and gave existing ones until January 1, 2027, to move to something else, according to its documentation.
Two changes are landing at once, which is where most of the confusion comes from. The January 2026 announcement retired the "Search the entire web" option: since January 20, 2026, every new engine has to specify a site list capped at 50 domains, and existing whole-web engines keep working only until January 1, 2027. The JSON API is discontinued on that same date regardless of how your engine is scoped. So if you run a site-restricted engine, the free Search Element still serves your site list after the deadline (a whole-web engine must first be reconfigured to 50 or fewer domains); your programmatic access to them does not.
Most of the migration work is mechanical. The risk sits in what you assume carries over, because the replacement you choose changes your result schema, your pagination model, your ranking behavior, where your domain restrictions live, and what your downstream code is allowed to assume. Teams that treat this as a one-line change in November 2026 will find out the difference in production.
This guide covers what is actually being shut down, what you need to inventory before you migrate, what your realistic options are, and how to pick between them based on what your integration actually does.
What Google Is Shutting Down, and What It Isn't
Five Google products share the "Custom Search" or "Programmable Search" naming. The shutdown applies to one of them. Identify which surface your integration depends on before scoping any migration work.
| Product | Surface | Status | Action |
|---|---|---|---|
| Custom Search JSON API | GET https://www.googleapis.com/customsearch/v1 | Closed to new customers; discontinued January 1, 2027, whatever the engine is scoped to | Migrate |
| Custom Search Site Restricted JSON API | GET .../customsearch/v1/siterestrict | Ceased serving traffic January 8, 2025 | Already migrated, or already broken |
| Programmable Search Engine element | Client-side JavaScript widget (cse.js) | Continues free, capped at 50 domains for new engines (and for existing ones starting Jan 1, 2027) | Confirm your site list fits |
| Search Console API | searchconsole.googleapis.com, first-party analytics | Unaffected, and out of scope | None |
| Vertex AI Search | discoveryengine.googleapis.com, GCP-managed index | Active; Google's recommended path | Evaluate as a new build, not a swap |
Only the first row needs migration work, and the 50-domain cap does not exempt you from it: a cx scoped to five sites loses its API access on the same date as one set to search the whole web. One caution if you go looking: switching "Search the entire web" off is one-way, so leave that toggle alone on a production engine. An engine ID in your console proves nothing on its own either, since the same cx can back both the widget and the API, so grep for the URL, for customsearch/v1 client-library calls, and for wrappers like LangChain's GoogleSearchAPIWrapper that hold the credential for you. Vertex AI Search, the path Google recommends, searches a corpus you supply rather than the public web, and queries run cheaper than CSE's $5 per 1,000 while you take on the indexing and storage bill instead, so treat it as a new build rather than a swap. And since Microsoft retired the Bing Search APIs in 2025, assume you will have to do this again: pick a replacement you can replace.
What Changes for Existing Google Custom Search Users
Today your integration almost certainly looks like this:
HTTP GET → customsearch/v1 → JSON → your parser → your app
The parser is where the migration cost lives. Before you evaluate anything, inventory the following. Most of it takes an afternoon with grep and your billing dashboard, and it is the difference between a controlled cutover and a surprise.
Before you migrate: Inventory checklist
-
Every call site. Application code, cron jobs, notebooks, internal tools, that one Zapier step nobody owns, and any vendored library that calls the API on your behalf.
-
Which
cxengines exist, what each is scoped to, and whether the scoping is load-bearing. Domain restrictions configured in the Programmable Search control panel do not travel with you; they become request parameters or your own filtering logic. -
Query parameters in use.
siteSearch,dateRestrict,lr,gl,cr,fileType,safe,searchType=image,exactTerms. Each one needs an equivalent, a workaround, or a deliberate decision to drop it. -
Result fields you actually read. Usually
title,link,snippet, anddisplayLink. Sometimespagemapfor structured data, which is the hardest field to replace because it comes from Google's own page annotations. -
Pagination behavior. Google returns 10 results per call, and the
startparameter cannot exceed 91, so 100 results is the hard ceiling. If you have a paging loop, it is built around that limit. -
Anything reading
searchInformation.totalResults. That number is an estimate, and most replacement APIs do not provide it. If you display it or paginate off it, decide now what it becomes. -
Quotas and spend. Google's pricing is 100 queries a day free, $5 per 1,000 after that, and a hard 10,000/day ceiling. Your current usage curve drives every pricing comparison you run.
-
Failure handling. Whatever backoff you built around Google's 429s must be retuned for a provider with a different rate-limit model, often QPS rather than daily quota.
Expect at least four incompatibilities in any migration: a different response schema, a different ranking, a different pagination model, and a different billing unit. Plan for all four rather than discovering them one at a time.
Your Migration Options
The useful way to group replacements is not "best to worst" but by what you need the search layer to do. Most teams fall cleanly into one of four buckets.
1. Traditional SERP APIs
Choose this when your application depends on search-engine-style output: rankings, ads, knowledge panels, People Also Ask blocks, image or news verticals, or position tracking. Rank trackers, SEO tools, and competitive monitoring live here.
Top providers include SerpApi, Serper, DataForSEO, and Bright Data, alongside the Brave Search API, which serves its own independent index rather than reselling Google's.
These get you closest to what you have today in terms of result composition. The trade-offs are cost at volume, latency measured in seconds rather than milliseconds because most fetch and parse a live results page, and, for the Google-scraping providers specifically, a legal picture worth watching. Google filed suit against SerpApi in December 2025. The service is still operating, but if the outcome matters to your business, build a swap path from day one.
2. AI-native search APIs
Choose this when search output feeds an LLM, an agent, or a Retrieval-Augmented Generation (RAG) pipeline. These APIs are built for a machine reader, so the unit of output is a ranked, deduplicated passage carrying source metadata and a timestamp, which is what a model can reason over directly.
The main providers here are Octen, Tavily and Exa.
This is a different product from a SERP API, not a cheaper version of one. A page of results is built for a person scanning it, and getting from there to something a model can use means writing and maintaining a scrape-and-clean layer; this category removes that layer. Most providers also let you pull more content per result when a passage is not enough context.
3. Search plus extraction
Choose this when discovery is only the first half of the job, and you need the actual content of the pages you find, not the snippet describing them.
Firecrawl covers both halves, or you can compose your own by pairing Jina Reader with any search API.
This pattern powers research agents, monitoring tools, and anything that summarizes source material. Composing your own is fine and often cheaper, though it means two rate limits, two failure modes, and two bills. Before you build the pipeline, check whether you need one: some AI-native search APIs return page content from the search call itself, Octen among them, through its full_content option, which collapses discovery and retrieval into a single request and removes the second hop entirely.
4. Model-native grounding
Choose this when search is tightly coupled to one model provider and integration simplicity beats retrieval control.
Providers: OpenAI's built-in web search tool, Gemini grounding with Google Search, Anthropic's web search tool.
This is the least code. It is also the least control: you cannot easily inspect, filter, cache, or re-rank what the model retrieved, and you cannot reuse the retrieval layer if you switch models. Per-search pricing here tends to sit at the top of the market, and it is billed alongside tokens. Good for a feature. Awkward as infrastructure.
If your workload is agent retrieval at high concurrency, Octen sits in category two,and its full content option covers what most team reach for category three to get.
How the Alternatives Compare to Google Custom Search
The table below compares categories rather than individual vendors, because within a category the differences are mostly pricing and index quality, and both change faster than any published comparison stays accurate. Verify current numbers against each provider's own pricing page before you commit.
| Google CSE (current) | SERP APIs | AI-native search | Search + extraction | Model-native grounding | |
|---|---|---|---|---|---|
| Index | Google, via a cx engine | Google (scraped) or independent | Provider's own index | Provider's own index | Provider's own |
| Freshness | Google crawl schedule; no control | Live SERP | Provider-dependent; minute-level in some cases | Live fetch at request time | Provider-dependent |
| Result format | items[] with snippet | Parsed SERP blocks | Ranked passages, LLM-ready | Passages plus full page text | Prose answer with citations |
| SERP fidelity | Partial (CSE ≠ google.com) | High | None | None | None |
| Pagination | 10 per call; start ≤ 91, so 100 results max | Page-based | Usually single call, up to 100 | Single call | N/A |
| Total match count | Estimated, via searchInformation.totalResults | Usually | Not returned | Not returned | Not returned |
| Concurrency | No exposed; no QPS control | Plan-dependent | Plan-dependent QPS | Plan-dependent | Model rate limits |
| Full content | No; snippet only | No; snippet only | Returned inline on request | Yes, via a second fetch | Consumed by the model, not exposed |
| Agent/RAG fit | Poor | Poor without cleanup | Good | Good | Good but opaque |
| Pricing unit | $5/1k, 10k/day cap | Per query or monthly plan | Per search, sometimes plus a per-result charge for content | Per query plus per page | Per query, highest tier |
A compatible schema is not a compatible service. Several vendors advertise a drop-in replacement that returns Google-shaped JSON so your parser survives untouched. That saves a day of work, but it tells you nothing about ranking quality, coverage, freshness, filter behavior, or how the service fails under load. Keep the compatible schema if it saves time; spend the time saved testing everything the schema cannot tell you.
Treat published benchmarks as directional. Every provider in this space publishes numbers showing itself in front, usually on a benchmark it selected and a configuration it tuned. Latency figures in particular depend heavily on region, payload size, and whether full content was requested. Your own query distribution is the only benchmark that predicts your outcome.
Choosing the Right Replacement
A short decision rule, based on what your current integration actually does:
-
You need rankings, snippets, ads, or search verticals (rank tracking, SEO tooling, SERP monitoring) → Choose a SERP API. Brave if you want an independent index and predictable terms; Serper or DataForSEO if per-query cost dominates; SerpApi if you want the broadest engine and vertical coverage.
-
Search feeds an agent, chatbot, or RAG pipeline → Choose an AI-native search API. You get structured, LLM-ready results and no HTML-cleaning code to maintain.
-
You need full-page content after discovery → Choose search plus extraction. Either one vendor covering both, or a search API composed with a reader.
-
You are committed to one model provider and want the smallest possible integration → Choose model-native grounding, accepting that you give up control of the retrieval layer.
-
You need fresh web retrieval at high concurrency, controlled independently of your model → Evaluate Octen, which is built around QPS-based limits and near-real-time indexing rather than daily query caps.
If you are in more than one bucket, split the traffic. A rank tracker and an agent do not need to share a provider, and forcing them to usually means one of them gets the wrong tool.
Migrating from Google Custom Search to Octen
If your integration falls into the agent, RAG, or content-retrieval buckets, here is what the Octen path actually involves. Everything below maps to the Octen API reference.
What changes
| Google CSE | Octen | |
|---|---|---|
| Method | GET with query string | POST with JSON body |
| Auth | key query parameter | x-api-key header (or bearer token) |
| Engine config | cx engine ID, configured in a control panel | No engine ID; scoping is per request |
| Results per call | 10 per call, start ≤ 91, 100 results max | count, 1–100 in a single call |
| Domain scoping | Configured in the Programmable Search UI | include_domains / exclude_domains per request |
| Date filtering | dateRestrict=d7 | time_range, or start_time/end_time with time_basis |
| Snippet | snippet | highlight, query-relevant, token-budgeted |
| Full content | Not available | full_content.enable |
| Limits | 10,000 queries/day | QPS-based, 20 QPS on the base tier |
The two changes that ripple furthest into your code are pagination and domain restrictions. Pagination disappears for most use cases because you ask for up to 100 results in one call instead of looping. Domain restrictions move into the request, which means the configuration that lived in Google's control panel now lives in your repository. That second one is an improvement, but it is a migration step people forget until results come back unscoped.
Before and after
Existing Google request:
import requests
resp = requests.get(
"https://www.googleapis.com/customsearch/v1",
params={
"key": GOOGLE_API_KEY,
"cx": SEARCH_ENGINE_ID,
"q": "retrieval augmented generation benchmarks",
"siteSearch": "arxiv.org",
"siteSearchFilter": "i",
"num": 10,
"start": 1,
"dateRestrict": "m1",
},
timeout=10,
)
data = resp.json()
for item in data.get("items", []):
print(item["title"], item["link"], item["snippet"])
Equivalent Octen request:
import requests
resp = requests.post(
"https://api.octen.ai/search",
headers={
"x-api-key": OCTEN_API_KEY,
"Content-Type": "application/json",
},
json={
"query": "retrieval augmented generation benchmarks",
"include_domains": ["arxiv.org"],
"count": 10,
"time_basis": "published",
"time_range": "month",
"highlight": {"enable": True, "max_tokens": 512},
},
timeout=10,
)
body = resp.json()
if body["code"] != 0:
raise RuntimeError(f"{body['code']} {body['msg']} (request_id={body['request_id']})")
for r in body["data"]["results"]:
print(r["title"], r["url"], r.get("highlight", ""))
Note the error handling. Octen returns a business status code in the body alongside the HTTP status, and every response carries a request_id. Log it. It is the first thing support will ask for. Documented failure codes are 401 for an invalid key, 403 for insufficient balance, 429 for rate limiting, and 500 for server errors; the error codes reference covers the full list.
Keeping your downstream code
If your parser, template, or caching layer is widely depended on, do not rewrite it in the same change. Normalize at the boundary instead:
def to_cse_item(result: dict) -> dict:
"""Map an Octen result onto the CSE item shape existing code expects."""
return {
"title": result.get("title", ""),
"link": result.get("url", ""),
"snippet": result.get("highlight", ""),
"displayLink": urlparse(result.get("url", "")).netloc,
}
That gets you to production with one changed module. Once traffic is stable, unwind the adapter and start using the fields Google never gave you: time_published, time_last_crawled, and full_content when you need the page itself rather than a description of it.
searchInformation.totalResults has no equivalent, and no honest one exists. If you paginate off it, switch to a "load more" pattern. If you display it, drop it.
Rate limits and cost
Octen's limits are expressed as QPS rather than a daily cap: 10 QPS on the free tier, 20 once you add credits, and higher on paid plans. That fits bursty agent traffic better than a daily quota, where a single task may fan out into dozens of concurrent queries, and it means your retry logic should back off on the second rather than wait for a midnight reset.
Current per-call rates are on the pricing page. Full content is billed per result rather than per token, at $0.50 per 1,000 results, and the first 10 full-content results of every search come free. At the result counts most integrations use, that means you can leave full_content on by default instead of gating it behind a relevance threshold, and only start counting when you pull deep result sets.
Validate before you cut over
Do not compare feature lists. Run both APIs against the same queries:
-
Pull 500 to 1,000 real production queries from your logs, stripped of personal data.
-
Run each through Google CSE and through each candidate. Store the raw responses.
-
Measure overlap (how many of Google's top 10 URLs appear in the candidate's), freshness (publication dates on time-sensitive queries), latency at p50 and p95 under your real concurrency, and cost at your actual monthly volume.
-
Have someone who knows the domain grade relevance on 50 queries by hand. Automated overlap tells you how similar a provider is to Google, not how good it is.
-
Re-run the same harness on the day you cut over, and keep it. It becomes your regression test for the next provider change.
Try it against your own queries: get an API key and run your existing query set through /search before you change a line of production code.
Migration Checklist
-
Inventory every Google Custom Search call site, including scripts and third-party tools.
-
Export 500–1,000 representative production queries from your logs.
-
Document the result fields your downstream code reads, and every query parameter in use.
-
Record current volume, spend, and p95 latency as your baseline.
-
Shortlist two or three candidates from the category that matches your workload.
-
Run the same query set through each and compare overlap, freshness, latency, and cost.
-
Hand-grade relevance on a sample; do not rely on overlap alone.
-
Write the normalization adapter and unit-test it against stored responses.
-
Re-tune retry and backoff for the new limit model (QPS, not daily quota).
-
Deploy behind a feature flag and canary 5–10% of traffic.
-
Monitor result quality, error rates, and spend for two weeks.
-
Complete cutover well before January 1, 2027, and leave the old path removable, not removed.
-
Delete the Google API key and close the billing account once traffic is at zero.
Conclusion
The right replacement for Google Custom Search depends entirely on what your integration actually does. A rank tracker and a RAG pipeline both call customsearch/v1 today, and they should not call the same thing tomorrow. Decide which category you're in before you compare vendors, because the comparison criteria differ for each.
Whichever direction you go, test against your own query distribution. Feature lists and vendor benchmarks are a filter for building a shortlist, not a basis for choosing from it. The only evaluation that predicts your production experience is running your queries.
Two paths from here. If you need SERP fidelity or site search over a small set of domains, evaluate the SERP and site-search options above. If you need real-time web retrieval for agents, high-concurrency workloads, or a search layer you can keep independent of your model provider, run your existing Google queries through Octen and compare the output directly.