Close
Nick GrantOctober 5, 202620 min read

How to Keep Negatives Ahead of Google Broad Match Drift

Broad match can create new query variations faster than a manual negative keyword process can evaluate and exclude them, leaving paid search teams continually reacting to waste after spend has already occurred. A scalable negative keyword pipeline fixes this: aggregate search term spend, isolate recurring n-gram waste, map intent, deploy exclusions at the narrowest scope, and audit impression share loss.

One timing note worth flagging up front: starting September 2026, Google is automatically upgrading campaigns that use the campaign-level broad match setting into AI Max, its most aggressive Search matching layer yet. Negative keywords are still respected under AI Max, and AI Max search terms already surface in the standard search terms report. That makes the pipeline below more relevant, not less, as this upgrade rolls out.

Why Don't Negative Keywords Block Close Variant Matches?

Negative keywords are more literal than positive keywords. Google can match a positive keyword to close variants and searches with the same meaning, but negatives do not get that same expansion. A negative that blocks one bad query may therefore miss a slightly different version of the same unwanted intent.

If template is added as a negative broad match, don't assume templates is covered. Google does not automatically extend negatives to close variants, so the operator may need to add both. Negative exact match is narrower still: blocking [free accounting template] does not cover every variation of that unwanted search.

Positive keywords work differently. Google's positive matching can account for close variants and meaning, giving broad match considerably more room to find new searches than negatives have to exclude them.

At scale, that asymmetry makes manual SQR (Search Query Report) audits difficult to sustain. The practical response is to find recurring waste patterns across many queries and build broader negative coverage where the data supports it.

That creates a second problem: all those exclusions have to live somewhere. As negative coverage grows, the operator has to decide which terms belong at the campaign level, which should be shared across campaigns, and which are safe to apply account-wide.

What Are the Practical List Limits You're Managing Against?

Google Ads provides substantial negative capacity but with hard thresholds. Shared negative lists hold up to 5,000 keywords each, account-level negatives are capped at 1,000, and campaign-level negatives can reach 10,000 per campaign. At scale, that means your negative pipeline has to prioritize and place exclusions, not just keep adding them.

Here are the limits that matter:

Negative control

Limit

Use it for

Account-level negatives

1,000

Intent you never want anywhere in the account

Shared negative list

5,000 per list

Exclusions that apply across a defined group of campaigns

Shared negative lists

20 per account

Separating exclusions by product, market, or intent

Campaign-level negatives

10,000 per campaign

Waste specific to one campaign

PMax (Performance Max) campaign-level negatives

10,000 per campaign

Search and Shopping query exclusions specific to that PMax campaign, or via a shared negative keyword list applied to one or more PMax campaigns (limit raised from 100 to 10,000 in March 2025)

Put the negative as high as necessary, but no higher.

If jobs is worthless everywhere, an account-level negative may make sense. If free is junk in one campaign but another product has a legitimate free trial, keep it at the campaign level. If small business and personal should be excluded across five enterprise campaigns but remain eligible in small business campaigns, a shared list is the better fit.

That gets harder as the account grows. Twenty shared lists can disappear quickly across multiple brands, countries, product lines, and different definitions of bad intent. Separate lists for competitors, jobs, support, low-value segments, products, and regional exclusions can leave the operator deciding which campaigns can safely share exclusion rules and which need to stay separate.

I hit this exact wall in a consulting engagement with a mid-market home goods e-commerce brand scaling broad match hard heading into a holiday push. Their team had been adding every ugly long-tail variant straight into one shared list as exact match negatives, and it filled up to the 5,000-keyword ceiling right as we were trying to launch into three new countries. The fix wasn't more negatives, it was fewer, better ones. A quick n-gram pass on the existing list showed that a few hundred well-chosen negative broad match unigrams and bigrams, terms like cheap, used, and secondhand, covered nearly all the same unwanted intent that 5,000 exact match entries were straining to catch. We cut the list to under 400 terms and had room to spare.

The same "keep the negative close to the problem" rule applies to Performance Max.

Google now allows up to 10,000 campaign-level negatives per PMax campaign. So if one PMax campaign keeps spending on used equipment queries while another campaign intentionally sells used inventory, used can stay contained to the PMax campaign where it is a problem instead of becoming an account-wide exclusion.

The extra capacity helps, but it does not change the job. You still have to decide which search patterns are wasting enough money to deserve a negative in the first place.

And once broad match is producing thousands of search terms, finding those patterns one query at a time doesn't scale. That is where n-gram analysis starts to earn its keep.

How Do You Build an N-Gram Pipeline That Finds Waste Automatically?

An n-gram pipeline takes the search terms you're already paying for, breaks them into repeated words and phrases, and rolls the performance back up against those fragments. Instead of finding 50 bad queries one at a time, you can see that the same two-word pattern is responsible for $8,000 of spend across all 50.

An n-gram is just a word or sequence of words pulled from a search term. The point is finding the same expensive language hiding across hundreds of different queries.

Take:

free accounting templates for contractors

The pipeline can break the original query into contiguous sequences like this:

Type

N-grams

Unigrams (1 word)

free · accounting · templates · for · contractors

Bigrams (2 words)

free accounting · accounting templates · templates for · for contractors

Trigrams (3 words)

free accounting templates · accounting templates for · templates for contractors

Common filler words such as for can be ignored when scoring individual unigrams, but preserve the original word order when generating longer n-grams.

Then run the same process across the rest of the search terms:

download accounting templates

free bookkeeping templates

accounting templates for small business

free invoice templates

Now the repeated patterns start to show up:

N-gram

Search terms containing it

Spend

Qualified conversions

templates

63

$4,800

0

free

41

$3,100

1

accounting templates

27

$2,600

0

free accounting

12

$900

0

Take templates. Instead of reviewing 63 search terms individually, the n-gram analysis shows that this one recurring pattern has consumed $4,800 without producing a qualified conversion.

That is the value of the analysis: it surfaces waste that is spread across many different queries and makes the financial impact visible in one place.

But templates should not automatically become a negative. The pattern has earned an operator's attention; the underlying queries still need to be reviewed before anything gets blocked.

From there, the pipeline has four jobs: pull enough data, collect Search and PMax from the right sources, break the queries into useful n-grams, and roll performance up against those patterns.

1. Pull Enough Search-Term Data to See a Pattern

Start with a window long enough to get past daily noise, but short enough to represent the way the account is being matched now. Pull the search term, cost, clicks, conversions and conversion value, along with campaign and ad group context.

For a high-volume account, 30 days may be enough to expose recurring patterns. Lower-volume or high-consideration accounts may need 60 or 90 days.

Don't score recent search terms as if their conversion data is settled.

Before setting the analysis window, check how long conversions actually take to arrive. In Google Ads, you can find this under conversion lag reporting: hover over the conversions column in the "Campaigns" table for conversion lag data directly, or go to the "Campaigns" page, select Segment, then choose Conversions → Days to conversion to see how many days it typically takes to count conversions. That gives you a starting point for deciding how much recent data is too immature to score confidently.

For lead-generation accounts, Google Ads conversion lag may only tell part of the story. Qualified opportunities or revenue can take longer to make it back from the customer relationship management (CRM) system.

Say the account regularly needs about two weeks before enough downstream data has arrived to judge search quality. A pipeline running today should not treat yesterday's $500 of spend and zero qualified opportunities the same as $500 spent six weeks ago with zero qualified opportunities.

Either exclude the immature portion of the date range from scoring or flag it separately so it cannot produce high-confidence negative candidates.

Use the outcome the business actually buys against. If the account is managed to qualified opportunities or revenue, score the n-grams against that mature CRM data rather than Google Ads lead count simply because it is available sooner.

2. Pull Search and PMax Data Separately

If the n-gram analysis covers both Search and Performance Max, you need to pull search-term data from two different places. Standard Search uses search_term_view; Performance Max uses campaign_search_term_view.

This matters because search_term_view does not include Performance Max. A script can run without errors and still leave all of your PMax search terms out of the analysis.

Pull both datasets, normalize them into the same structure, and then run the n-gram analysis across them. Keep campaign type and campaign name attached so you can still see where each waste pattern came from.

The extraction method also has to match the size of the account.

Google Ads Scripts are fine for getting this running quickly, but advertiser-account scripts have a 30-minute execution limit. On a smaller account, there may be plenty of time to pull the search terms, generate the n-grams, aggregate performance, and write the output to a Google Sheet. As search-term volume grows, that same job can start timing out before it finishes.

I watched this happen in real time with a B2B SaaS account that scaled spend hard after a funding round, moving from roughly $40K a month to a quarter-million a month across several countries within a couple of quarters. The daily n-gram script that had run fine for a year started dying partway through the PMax pull, right at the 30-minute mark, which meant we were blind to our biggest source of spend on the days it mattered most. Shortening the lookback window got it running again, but with a sales cycle running about three weeks, the data was too fresh to trust. The real fix was moving off Ads Scripts entirely: pulling raw search terms through the API into BigQuery, where a scheduled SQL job handled the tokenization and rollup in seconds and fed a dashboard that never timed out.

Don't solve that by shortening the date range until the script runs. That gives you a faster pipeline built on less evidence.

At that point, move the heavier processing outside Google Ads Scripts. Pull the raw search-term data through the Google Ads API into a data warehouse or other processing environment, then run the tokenization and aggregation there. BigQuery with SQL or a Python workflow are common options; the specific stack matters less than preserving the full input and making the job repeatable.

The review output can still be simple. Start with the aggregated n-gram candidates so the paid search team can see which patterns deserve attention, but preserve the underlying search terms behind each one. When a pattern rises to the top, the operator should be able to drill into the actual queries, campaigns, spend, and conversions before deciding whether it belongs on a negative list.

There is one limitation the pipeline cannot fix: Google does not expose every search term because of its privacy thresholds. The analysis can only work with the search-term data Google makes available.

3. Break Every Query Into Unigrams, Bigrams, and Trigrams

Use one-, two-, and three-word n-grams because each level gives you a different amount of context. Unigrams are good at finding broad waste patterns, while bigrams and trigrams help determine whether the problem is really the word itself or the intent around it.

Say the pipeline finds these searches:

free accounting software

free accounting templates

free bookkeeping software

accounting software free trial

Breaking them apart shows why the levels matter:

Level

Example

What it tells you

Unigram

free

Is this word associated with waste across the account?

Bigram

free accounting

Is a more specific use of the word causing the problem?

Trigram

free accounting templates

Is there a specific query pattern that can be excluded more safely?

Suppose the unigram free has spent $12,000 at a terrible cost per acquisition. That gets attention, but it does not automatically make free a good negative.

Drill down and you might find that free accounting templates produces nothing, while accounting software free trial converts well.

Now the distinction matters. The unigram found the problem. The bigrams and trigrams showed where the problem actually lives.

That is why the pipeline should preserve the original search terms underneath every n-gram. An operator reviewing free should be able to open it and see the queries, campaigns, spend, and conversions behind the aggregate before deciding what, if anything, to exclude.

4. Roll Performance Up Against Each Fragment

For each n-gram, add up the cost, clicks, conversions, conversion value, and number of search terms containing it. Now you can see how much money is flowing through a repeated piece of language instead of judging each search term in isolation.

A basic output might look like this:

N-gram

Queries

Cost

Conversions

Cost/conv.

Conv. value

free

73

$8,200

2

$4,100

$900

jobs

41

$3,400

0

n/a

$0

for accountants

26

$6,100

19

$321

$14,800

template download

18

$2,700

0

n/a

$0

This is the payoff of the n-gram pipeline. Instead of seeing jobs scattered across 41 different search terms, you can see the total $3,400 spent on that pattern. Instead of noticing one bad free query, you can see what free is doing across 73 queries.

But the output is still evidence, not a negative list. Free could include profitable free trial searches. Jobs might be irrelevant everywhere it appears.

The pipeline finds the patterns and puts the money behind them. The next step is deciding which ones actually deserve to be blocked.

How Should Negative Candidates Get Scored and Routed for Approval?

Rank negative candidates by financial impact, then add operator confidence before anything gets excluded. High-cost, recurring patterns with clearly irrelevant intent should move quickly; ambiguous terms or negatives with a large account-wide blast radius should get more scrutiny.

An n-gram report can surface hundreds of ugly-looking patterns. They should not all get equal attention.

Score Impact First

The impact score decides what gets looked at first. Start with wasted spend, then account for how broadly the pattern is recurring across search terms.

Say an enterprise software campaign surfaces these two patterns:

N-gram

Queries containing it

Spend

Qualified opportunities

jobs

48

$4,200

0

certification course

3

$3,900

0

Both have spent enough to investigate. But jobs appears across 48 different queries, while certification course is concentrated in three. That makes jobs the more persistent pattern and pushes it higher in the review queue.

But the score only tells you what to investigate first. It does not tell you whether the negative is safe to add.

Add Operator Confidence Before Approval

Once a candidate reaches the top of the queue, inspect the search terms underneath it. The question changes from "How much is this costing us?" to "Are we confident this intent is something we don't want?"

Start with jobs. The underlying searches might be:

accounting software jobs

software sales jobs

accounting software implementation jobs

remote accounting jobs

If the campaign is strictly acquiring enterprise software customers, the pattern is consistent. The spend is meaningful, the pattern is recurring, and the underlying intent is clearly employment-related. jobs would carry high confidence as a negative candidate.

Now look at certification course:

accounting certification course

enterprise accounting certification course

software certification course

The performance is bad, but the intent is less obvious. Maybe training searches are irrelevant. Maybe certification is part of the company's enterprise offering. Maybe those leads have a longer sales cycle and the CRM has not reported the downstream opportunities yet.

That candidate might stay medium confidence until the operator checks the business context and conversion lag.

Tag candidates as High, Medium, or Low confidence before pushing them live.

• High: underlying queries consistently represent unwanted intent.

• Medium: performance is poor, but there are legitimate exceptions or unresolved business context.

• Low: there is not enough evidence yet to justify an exclusion.

That gives the review queue two useful dimensions: impact tells you what to investigate first; confidence tells you how safe it is to act.

Route Approval Based on Blast Radius

Once a candidate has enough impact and confidence to act on, the next question is where the negative should live. The wider the negative applies, the more scrutiny it should get.

Go back to jobs. The n-gram analysis showed meaningful spend across searches such as:

accounting software jobs

software sales jobs

accounting software implementation jobs

remote accounting jobs

The pattern has enough impact and the underlying intent is consistently employment-related, so jobs is a high-confidence candidate. The remaining question is where to apply it.

If those searches are only a problem in one enterprise acquisition campaign, keep jobs at the campaign level. The paid search owner can review and deploy it without turning the decision into a committee meeting.

If the same employment intent is wasting spend across a group of acquisition campaigns, jobs may belong on a shared negative list. Now the channel lead should check that every campaign attached to that list should exclude employment searches.

If jobs is irrelevant to every campaign in the account, an account-level negative may make sense. Because that decision affects everything, it deserves one more check from the channel lead and relevant business owner before deployment.

The workflow is:
N-gram candidate
↓
Impact: Is it costing enough to prioritize?
↓
Confidence: Is the underlying intent actually unwanted?
↓
Scope: Where should it be blocked?
↓
Campaign → Paid search owner
Shared list → Channel lead
Account level → Channel lead + relevant business owner
↓
Approve → deploy → log → monitor

Now compare that with certification course. Its impact was meaningful, but confidence was only medium because the business may legitimately want some certification-related traffic. That candidate should not reach deployment simply because the spend looks bad. It stays in review until the intent and downstream conversion data are clear.

That is the operating principle: impact determines priority, confidence determines whether you should act, and scope determines how much approval the action needs.

For every negative that does get deployed, log the term, match type, scope, date, owner, reason, and performance evidence behind the decision. If it gets removed later, log that too.

That history matters because the next failure mode is the opposite of wasted spend: a negative works exactly as configured and quietly blocks demand you actually wanted.

How Do You Know When You've Gone Too Far With Negatives?

Google Ads does not report "impression share lost to negatives," so over-negation has to be diagnosed indirectly. If important demand drops after a negative deployment, rule out budget and Ad Rank first, then trace the loss through change history and every negative layer that can affect the campaign.

Google reports Search impression share (IS), Search Lost IS (budget), and Search Lost IS (rank), but it does not tell you how much eligible traffic a negative keyword removed.

So when important search demand drops after a negative deployment, troubleshoot it in a fixed order.

First, check Search Lost IS (budget) and Search Lost IS (rank). If budget loss increased, the campaign may simply be constrained. If rank loss increased, investigate bids, targets, Ad Rank, and competitive pressure. Auction Insights can add context if the competitive environment changed at the same time.

If budget and rank do not explain the drop, go to Change history and line the timing up with the negative deployment. A large batch added on Tuesday followed by a drop in an important query theme on Wednesday gives you somewhere concrete to investigate.

Then check every negative layer that can affect the campaign:

• Ad group negatives

• Campaign negatives

• Shared negative lists attached to the campaign

• Account-level negatives

The offending negative may not live in the campaign where the loss appears.

Don't stop when you find a negative containing the same word. Check its match type and scope against the searches that disappeared. The problem may not be the negative itself; it may be that a reasonable negative was applied too broadly.

I've seen this exact failure mode play out on an account running separate B2B and B2C lines under one roof, in this case commercial and consumer coffee equipment. The team managing the B2B side noticed a few thousand dollars a month leaking to home-kitchen queries and added terms like home, mini, and kitchen to an account-level negative list, treating it as universally bad intent. Within days, revenue on the flagship consumer espresso machine had dropped by well over a third. Lost IS (budget) and Lost IS (rank) were both flat, which ruled out a bidding problem, so we went straight to change history, lined it up against the traffic drop, and found the account-level update sitting right at the top. The B2B team's own negative had cut off their sister product line's best-performing demand. We pulled the terms off the account-level list and reapplied them as campaign-level negatives on the B2B campaigns only, and volume came back the next day.

That is the over-negation control: rule out budget and rank, trace the timing, find the negative, and roll back the smallest possible piece.

Keep that rollback in the audit log. Otherwise the same pattern can surface in the n-gram report six months later and get excluded again by someone who has no record of why it was removed.

That is also why a negative keyword program cannot just be a growing list of things the account does not want.

Broad match will keep finding new ways to interpret intent. Search behavior will keep changing. And every new negative creates its own risk of blocking something valuable.

A mature negative program runs as a repeatable operating loop:

Detect recurring waste → measure the financial impact → review the underlying intent → apply the negative at the right scope → monitor what happens next.

That keeps the negative program moving at the same speed as query expansion while leaving the final exclusion decision with the operator.

Done well, the program cuts spend on unwanted intent while preserving the demand the account was built to capture.

Frequently Asked Questions

How do I define a negative keyword policy that scales?

Define which types of intent should be excluded, where negatives should live, who can approve them, and what evidence is required before deployment. Use account-level negatives only for universally unwanted intent, shared lists for groups of campaigns, and campaign-level negatives when the exclusion should stay close to a specific source of waste.

What process keeps negatives current as matching changes?

Run search-term analysis on a recurring schedule, use n-grams to identify waste appearing across multiple query variations, and prioritize candidates by financial impact. Review the underlying searches before deployment, apply approved negatives at the narrowest appropriate scope, and monitor results afterward. Higher-volume accounts should run this process more frequently.

How do I prevent negatives from blocking incremental demand?

Review the search terms underneath each negative candidate before adding it, especially for broad patterns that can carry multiple meanings. Keep exclusions at the narrowest appropriate scope and use more specific negatives when necessary. After deployment, monitor important query themes and Search impression share so unintended demand loss can be investigated quickly.

avatar
Nick Grant
Nick Grant is the Marketing Director at Revvim, where his work centers on Reactive Search Management and how near real-time auction intelligence helps brands identify non-incremental spend. Revvim's patented AdAi technology helps advertisers reclaim budget and recapture trapped capital that would otherwise go unnoticed. Since first working with Google Ads in 2005, Nick has spent two decades at the intersection of digital strategy and search innovation. His background includes leading marketing for a high-growth digital agency and more than a decade consulting on marketing strategy for over 20 B2C and B2B businesses. He also co-founded a visual strategy agency, earning Inc. 5000 recognition and four Adobe MAX speaking appearances.

RELATED ARTICLES