Negative Keyword Automation with AI
You can build an AI system with Python that reviews the new search terms in a Google Ads account every day and sends you a summary: how many terms came in, and how many negatives wait for your approval.
Managing negative keywords directly through Claude Code or Codex over MCP works, but a proper system has advantages: it runs on a schedule, tests its output for accuracy, turns rejections into prompt fixes, keeps AI costs under control, and stores all its output in BigQuery. The rejections are what keep the system improving: every rejection is feedback, and every fix is tested on the gold set before it ships, so it gets better where it was wrong without breaking what already worked.
That matters most on accounts with real spend, where one wrong negative blocks searches that convert, and on large accounts, where the new terms each day outnumber what a person can review.
A system like this can be used for a wide variety of use cases, like new keyword research, cross-negation, search term routing and many more.
#1 The system
The system has seven parts. Each lives in one place and does one job.
| Part | Where it lives | What it does |
|---|---|---|
| Search terms | BigQuery, filled daily by the Data Transfer | The account’s search terms and their stats |
| Stats filter | a SQL query | Drops the terms the numbers already settle, and the ones already judged |
| Gold set | a file of terms with known right answers | Scores every prompt change before it ships |
| Prompt | prompts/negatives.md, with the business context at the bottom | Reads each term and decides: ADD_NEGATIVE, LEAVE or TO_REVIEW |
| Pipeline | Python, run daily by CI | Calls the model in batches, validates the output, writes every decision to BigQuery, posts to Slack |
| Review | a Google Sheet to start | A person accepts or rejects each negative |
| Evals | the gold sets, and the rejections in BigQuery | Score the live prompt every day to track accuracy, and turn rejections into tested prompt fixes |
The sections below follow the table. At run time, the parts connect like this:
#2 Search terms and stats
The Google Ads Data Transfer lands the search terms report in BigQuery every day, with clicks, cost and conversions per term. Read stats over a date range from the partitioned p_ads_* table, filtered on segments_date and on the partition (_PARTITIONTIME) so the query scans only the days it needs.
Every decision the model makes is stored in BigQuery, and so is what happened to it next: accepted, rejected, applied. So the first query of every run pulls the search terms from the Data Transfer tables and skips every term that already has a decision (an anti-join), sending only the net-new ones. The model bill then scales with the day’s new terms rather than the size of the account. Then it applies the stats filter:
- Converted: a term with a conversion in the lookback window is kept and never sent.
- Too new: a term whose clicks are still inside the conversion lag waits, so it can still convert before anything judges it.
- Spend floor: a minimum spend before a term is worth a model call. $5 is a sensible start.
These filters are set in the Python configuration:
# config.py: which search terms reach the modelSTATS_FILTER = { "lookback_days": 30, "maturity_days": 7, # the account's conversion lag "drop_if_converted": True, # a converting term is never a negative candidate "min_cost": 5.0, # 0 turns the floor off}#3 The gold set
Build the gold set before the prompt. It is a file of search terms, each with the decision a domain expert would make, and it turns “the prompt looks good” into a number.
Start with about 100 terms from the real account: clear negatives and clear keeps in roughly equal numbers, plus the hard cases, spread evenly across categories (brand terms, competitors, product names the model may not know, how-to searches). A hundred copies of one easy case measure nothing.
Score a prompt by running it over the set and comparing each decision with the expected one. Accuracy, the share that match, is the headline number. For negatives, two kinds of mistake matter, and they cost different things:
- Blocking a good search: the prompt says
ADD_NEGATIVEfor a term that should stay. It costs conversions, and you never see them, because the searches stop. The share of negatives that were right is called precision. - Missing a bad search: the prompt lets through a term that should be a negative. It costs wasted spend, which stays visible in the search terms report. The share of bad searches caught is called recall.
NEG = "ADD_NEGATIVE" def score(decisions: dict[str, str], gold: dict[str, str]) -> dict[str, float]: pairs = [(decisions[term], expected) for term, expected in gold.items()] tp = sum(d == NEG and e == NEG for d, e in pairs) fp = sum(d == NEG and e != NEG for d, e in pairs) fn = sum(d != NEG and e == NEG for d, e in pairs) return { "accuracy": sum(d == e for d, e in pairs) / len(pairs), "precision": tp / max(tp + fp, 1), "recall": tp / max(tp + fn, 1), }One run does not show reliability. The same prompt on the same terms gives slightly different answers each time: in my setup, the daily score on a 250-term gold set moved between 95% and 98% over two weeks with no change to the prompt.
Once the system runs, the gold set is used in two ways, both covered in Evals and prompt improvements: a prod set scored every day, to catch a drop in the live prompt, and a dev set, to test a prompt change before it ships.
#4 The prompt
This is where most of the time goes, and where your domain knowledge comes in: build the prompt, run it, test it against the gold set, improve it, and repeat until the output is what you need. A coding agent does most of that work with you; Agentic Coding for PPC covers that setup.
Many prompt structures work, and every part of this one is yours to change. This is one example that does. It answers one question per term, and returns one of three decisions:
ADD_NEGATIVE: block it, with the negative’s text and match type.LEAVE: a search the account wants, or one it can live with.TO_REVIEW: a call for a human review, including every term the model does not recognise.
Everything else it returns is metadata: intent, category, brand and competitor flags, a one-line reason. The review sheet and the eval triage sort by it.
The structure matters more than the wording. Instructions go at the top, the business context at the bottom:
## ROLE the account, and the one question to answer per term## STEP 0 unknown brand, product or slang? TO_REVIEW, "unknown term"## STEP 1 classify: intent, category, brand, competitor## STEP 2 decide, first match wins: ADD_NEGATIVE, LEAVE or TO_REVIEW## STEP 3 negative text and match type: the full term, EXACT, by default## OUTPUT one JSON object per term, fields in a fixed order## BUSINESS CONTEXT the offer, out of scope, brand and competitors, account rules- Room for unknown: a model asked to pick one of three answers picks one, even for a brand it has never seen. Step 0 gives it a way out before any classification. It does not make the model know what it does not know, so put unfamiliar brand and product names in the gold set and measure how often they land in
TO_REVIEW. - Rules inside their step: each rule sits in the step where the model makes that call. In my setup, the same rule scored better inside its step than in a preamble at the top.
- Exact by default: a phrase negative blocks every search that contains it, which is where precision is lost. Phrase match needs a rule that proves the phrase is safe, and the pipeline enforces the default in code rather than trusting the prompt.
- One shared list: every negative goes to one shared negative keyword list on the non-brand campaigns. That works only when every negative is unwanted in every campaign the list is attached to. A list also holds at most 5,000 negatives, so watch its size and plan a second list. To place negatives per campaign or ad group, add the account structure to the context and a
levelfield to the output, and give every gold set case an expected level too.
The business context is the part each account writes for itself. Without it, the model judges search terms like a new analyst on day one: it knows what a negative keyword is and nothing about the business. Four blocks cover most accounts:
- The offer: what the business sells, in the words its customers use.
- Out of scope: what it does not sell or do, and searches it never wants.
- Brand and competitors: the names that need their own handling.
- Account rules: the calls an experienced account manager makes without thinking.
Keep the context in its own Markdown files and paste them into the prompt’s last block at run time. Most of what changes between accounts is then that block, and the steps change only when the eval loop finds a mistake in them. Review the files like code: a stale product list is worse than none, because the model applies it with full confidence.
#5 The pipeline
The pipeline is one Python entry point that runs the steps in order, each small and testable on its own.
- Fetch: the search terms query from the stats section.
- Build the request: the prompt with its business context, and a batch of search terms.
- Call the model: ask for JSON, and use the provider’s structured output mode where it has one. A reply cut off before the last term is retried with the batch halved.
- Validate: every object against a Pydantic schema, and every returned term against the batch: every term was in the batch, and each comes back exactly once. An invalid row is logged and dropped; its term stays unjudged, so the next run picks it up.
- Write: after each batch, to a staging table and then a MERGE into the decisions table, keyed on the search term, with the run ID on every row.
- Notify: the run’s counts go to its run log, and to Slack when a summary is due.
How many terms go into one call is a setting to test. 100 is a safe start. Bigger batches mean fewer calls and less repeated prompt, but past some size accuracy drops and long replies get cut off. Where that point sits depends on the model and its context window, and it moves with every model release. Run the gold set at 100, 200 and 300 terms per call, several times each, and keep the largest size whose range matches the 100 run.
If cost matters, cache the fixed part of the prompt (the steps and the business context) and send the daily run through the provider’s batch endpoint. Both are billed below the standard rate, and nothing here needs an answer in seconds.
Schedule and Slack
CI runs the job once a day, with the API keys in its secret variables. In GitHub Actions that is a workflow with an on: schedule trigger; GitLab does the same with a pipeline schedule.
# .github/workflows/negatives-daily.ymlname: negatives-dailyon: schedule: - cron: "0 7 * * *" # 07:00 UTC, after the Data Transfer has runconcurrency: group: negatives # one run at a timejobs: run: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: astral-sh/setup-uv@v6 - run: uv sync --frozen - run: uv run python -m negatives --live env: LLM_API_KEY: ${{ secrets.LLM_API_KEY }}Whichever CI runs it, list every job on a clock in docs/scheduled-jobs.md: when it runs, what it writes, how to check it fired. On GitLab that page is the only record, because a pipeline schedule lives in the project settings, not in the repo. Agentic Coding for PPC has the rules for every job on a clock.
Slack keeps the team informed without opening BigQuery: how many terms the job judged, how many negatives wait for review, and a link to the sheet. Post it daily, weekly or monthly, depending on the size of the account and how often you add negatives. Read the numbers from BigQuery, so the post reports what landed rather than what the job meant to do.
Production engineering
- Dry by default: the pipeline writes only with
--live, so a forgotten flag is a dry run. A dry run still calls the model: it skips the writes, not the cost.--limitcaps the input for a first run of a new prompt or query. - Safe reruns: the anti-join and the per-batch MERGE make a rerun safe, as long as the staging rows are deduplicated first (a MERGE does not dedupe its own source) and a lock stops two runs writing at once. A run that crashes halfway keeps the batches it wrote, and the next run picks up the rest.
- Versions: store the prompt, context and model version with every decision, and keep review decisions in their own table. When a rule changes, you can then reprocess the terms it touched without losing their review history. Keyed on the search term alone, the table covers one account; more accounts need the account in the key.
- Run IDs: every run gets an ID, stamped on every row it writes and on its run log entry. One ID finds a run’s rows and its log.
- Fail soft: a rate limit or a timeout is retried with a growing wait, up to three times; a batch that still fails leaves its terms unjudged for the next run.
- Alerts: a failed run posts to Slack, and so does a scheduled run that never started.
- Service account: the schedule runs as a service account, which needs read access to every dataset a query touches. A query that runs on your machine can still fail on the clock.
- CI/CD: fast tests run in CI on every push, including end-to-end runs of the pipeline on frozen fixtures with BigQuery and the model mocked; tests against the real warehouse run before any SQL or schema change. A merge to main is the deploy, because the schedule runs the main branch, so a prompt change merges only after its gates pass.
#6 Review
Nothing reaches the account until a person accepts it. You can start in a Google Sheet: the pipeline writes each day’s proposed negatives as rows (term, negative text, match type, reason, cost), a reviewer marks each one accept or reject, and a sync job reads the decisions back into BigQuery. Build a review app on the same tables once the team outgrows the sheet.
Before a row reaches the sheet, the export drops terms a live negative already blocks and terms that converted since the model judged them. Check those conversions against fresh data: the transfer refreshes only the last seven days by default, so read them from your own conversion data or the Google Ads API, and check again if approval takes days.
Rejections, with a reviewer comment, flow into the eval loop, so give reviewers a short list of preset reasons plus a free-text box. “Wrong” gives the loop nothing to group.
Approved negatives go into the account by hand, pasted into the shared negative keyword list, or through the Google Ads API behind the same approval, which I will cover in a follow-up post.
#7 Evals and prompt improvements
The gold set scores a prompt before it ships. Evals are how you measure and improve it over time: they keep scoring it after it ships, and turn rejections into fixes. The goal is not 100% accuracy. It is knowing where the prompt works, where it fails, and what to fix next.
Daily accuracy check
A CI job scores the live prompt on the prod gold set every morning, and fails when decision accuracy drops under 90%. Nothing in the repo has to change for that to happen: a model alias can move to a new version, or someone edits the context without running the checks. Either one shows up the next morning, not after a week of bad negatives. The threshold is a drift alarm, not the bar a prompt change has to clear.
Prompt improvements
Rejections are where the fixes come from. Each round runs the same four steps:
- Triage: a model labels each new rejection with a failure mode from a short fixed list (
COMPETITOR_MISSED,PRODUCT_READ_AS_UNRELATED,PHRASE_TOO_BROAD) and counts them. The most frequent mode is the next fix. - Error analysis: one person, the domain expert, reads the rejections in the top group and names the pattern behind them. Hamel Husain’s writing on evals is the clearest material on this step.
- Eval the fix: the failing terms go into a dev gold set. Copy the production prompt, change only the step behind the failure, and run two gates, ten runs each, with the unchanged prompt as a control. Dev gold set: did the change fix the failures? Prod gold set: did it break anything that was passing?
- Approve: the change can ship when both gates pass, it adds no new false positives, and a short list of must-pass cases (your brand, your best-selling products) still comes out right. A person approves it, and dev terms that passed ten runs out of ten move into the prod gold set. The rest wait in a backlog that is still scored with every run, so the hard cases stay visible.
A coding agent runs everything but the judgment. A typical session:
You: "Run triage on the new rejections and show me the top failure mode" Claude: → runs the triage job over the rejections not yet labelled → reports: PRODUCT_READ_AS_UNRELATED, 18 of 40, all accessory searches You: "Fix it. Find where the prompt goes wrong." Claude: → reads prompts/negatives.md → finds step 1 matching accessories only on the words in the context → proposes a diff to step 1 and the accessories block → dev gold set, 10 runs against a control: the 18 terms pass 10 of 10 ✓ → prod gold set, 10 runs: 96 to 98%, within 2 points of the baseline ✓ → must-pass cases and false positives: no change ✓ "Both gates pass. Approve to merge." You: approve Claude: → updates the prompt → moves the terms that passed 10 of 10 into the prod gold set → commits the change with both resultsThe part you keep is deciding which failure modes matter. The same two gates decide every other change: a new model, a new provider, a rewritten context block, a bigger batch.
Metrics
Accuracy is one number. These five show where the prompt goes wrong:
- False positives: count the searches the prompt blocked that should stay, and check the negative text and match type as well as the label. A right
ADD_NEGATIVEwith a phrase that is too broad still blocks good searches, and accuracy alone hides it. - Review volume: track how many terms land in
TO_REVIEW. A prompt that sends every hard case to a person scores well and saves nobody time. - The flags: score intent, category, brand and competitor as well as the decision. In my setup, the competitor flag moved between 92% and 100% over the same two weeks in which the decision score held between 95% and 98%.
- Fresh labels: add new terms labelled by someone who did not write the prompt, and sample
LEAVEdecisions as well as negatives. Rejections only show the negatives the prompt proposed wrongly, never the bad searches it let through. - Accept rate: the share of proposed negatives reviewers accept is a quick live check, but reviewers accept what looks right, so it does not replace the gold set.
When new batches stop producing new failure modes, the prompt is done for now.
#8 Extensions
Web search for unknown terms
In my setup, about one term in five comes back unknown on a normal day. Each of those is a review a person has to do by hand, and web search gives the model the missing definition instead. There are two ways to add it:
- Grounding inside the model call: Gemini’s grounding with Google Search and Claude’s web search tool let the model search during the call itself and return the sources it used. One call and no extra client, with each search billed on top of the tokens.
- A search API between two passes: send the unknown terms to a search API like Tavily, then run them through the prompt a second time with the results attached. More code, but you decide exactly what the model sees. My setup works this way.
Either way, search only the unknown terms, not every term, and add a few of them to the gold set so the eval measures whether the lookup helped.
Extra fields
The model already reads every search term that passes the filters, so ask it for more in the same call: commercial intent, product category, competitor, brand or non-brand. Each field costs a few tokens and lands in the decisions table next to the decision. Reporting on intent by category or on competitor presence then needs no second pass over the data. It describes the terms the model saw, not the whole account: an account-wide branded share needs every term classified.
#9 The bigger picture
The same system runs other decisions: keyword research (from search terms, Search Console queries or a keyword research tool’s data), cross-negation and search term routing. The pipeline, the review step and the eval loop carry over unchanged. What each new decision needs is its own evals: a way to see where it works, where it fails and what to fix, and the loop that keeps improving it.
That changes the job. Instead of working through tasks by hand, you manage AI systems that do them for you: you build the pipeline with a coding agent, write and test the prompts, review the output at scale, and make the system better with every round of feedback. As models take on more of the daily decisions in paid search, that is where the work is going.