Are Those News Alerts the Same Story? A News Search API Deduplication Guide

By ··

When your news search API returns two links about one development, a correction or later update still deserves its own review. Collapse exact repeated URLs first; then assign story IDs with a human review step while preserving each publisher URL and every merge or split decision.

I reviewed the News Search documentation, the documented response shape and official NewsAPI material on October 6, 2026. The records below are fictional examples. This documentation review did not include an authenticated production call or a measured coverage test. The News Search overview explains the product surface.

Decide what counts as one story

A useful alert is an event or development a reader can act on, not a count of search-result rows. Make the rule explicit before writing code: an identical publisher URL is the same record; two publishers reporting the same announcement may join one story; a later recall or correction remains a separate development.

Candidate signalUse it forDo not infer
Canonicalized urlExact repeat across runs or queriesThat different URLs are different events
title and snippetQueue likely same-event reports for reviewThat similar headlines prove duplication
sourceRetain publisher diversityThat the source string identifies the rights holder
publishedTimeCalendar-day grouping when presentA precise publication timestamp

Walk through a fictional recall desk

Imagine you monitor product safety for a retailer. At 08:10, a publisher reports that fictional Northstar Batteries recalled Model A. At 08:20, your search returns the same article with a tracking parameter; at 08:35, a trade publication reports that same announcement. At 14:00, a regulator says Model C is now included. You want one alert for the original recall and another for the expansion, while keeping both publishers' links under the first alert.

Fictional rowDecisionWhy
08:10 publisher URLCreate story S1 and keep the raw URL.It is the first observed report of the Model A recall.
08:20 same URL with utm_sourceSuppress the repeat observation.The normalized URL is the same page; keep the second observation in its history.
08:35 trade publication URLQueue a possible merge into S1.A second publisher may corroborate the same event, but a title match alone is insufficient.
14:00 regulator updateCreate story S2, linked to S1.Adding Model C changes the action a reader must take.

This is an editorial rule, not a claim about real search results. If your alerts trigger trading, safety or legal decisions, show a person the underlying source page before grouping reports. A quiet queue is valuable only when it does not hide a new action.

Fetch rows and make a durable exact-match key

Request format=full to keep title, URL, source, snippet, image and publication date. The common seven article fields are present with null where unavailable. Google rows also carry google_news_url. Prefer a publisher url when the row has one, but retain both links and never assume a resolver link identifies a distinct event. Save the request query, engine, country and observation time beside every row.

Install requests with python -m pip install requests and set SERPENT_API_KEY before running this illustrative script.

import json, os
from datetime import datetime, timezone
from urllib.parse import urlsplit, urlunsplit, parse_qsl, urlencode
import requests

API = "https://apiserpent.com/api/news"
KEY = os.environ["SERPENT_API_KEY"]
DROP = {"fbclid", "gclid", "ref"}

def url_key(raw):
    if not raw: return None
    u = urlsplit(raw)
    if u.scheme not in ("http", "https") or not u.hostname: return None
    host = u.hostname.lower().removeprefix("www.")
    path = (u.path.rstrip("/") or "/")
    qs = [(k, v) for k, v in parse_qsl(u.query)
          if not k.lower().startswith("utm_") and k.lower() not in DROP]
    return urlunsplit(("https", host, path, urlencode(sorted(qs)), ""))

r = requests.get(API, params={"q": "battery recall", "engine": "google",
                              "country": "us", "pages": 1, "format": "full"},
                 headers={"X-API-Key": KEY}, timeout=60)
r.raise_for_status()
data = r.json()
if data.get("success") is not True: raise ValueError("No usable news answer")
payload = data.get("results")
if not isinstance(payload, dict): raise ValueError("Missing results object")
rows = payload.get("articles")
if not isinstance(rows, list): raise ValueError("articles was not a list")
seen, review = set(), []
for row in rows:
    if not isinstance(row, dict): continue
    key = url_key(row.get("url"))
    if key and key in seen: continue
    if key: seen.add(key)
    review.append({"url_key": key, "url": row.get("url"),
                   "title": row.get("title"), "source": row.get("source"),
                   "published_date": row.get("publishedTime"),
                   "snippet": row.get("snippet")})
print(json.dumps({"observed_at": datetime.now(timezone.utc).isoformat(),
                  "row_count": len(rows),
                  "needs_coverage_review": not rows or bool((data.get("meta") or {}).get("partialResults")),
                  "delivery": data.get("delivery"), "review": review}, indent=2))

Illustrative decision record: {"url_key":"https://publisher.example/recall","published_date":"2026-10-05","source":"Example Publisher"}. That fictional record shows the fields to persist; it is not a captured article. In a scheduled job, put keys in an atomic store rather than a process-local set, and keep the unmodified response for audit.

Make URL cleanup conservative

You can safely remove known tracking parameters that your own publishers use, but a query string can also select a different article, language or edition. The example drops common campaign parameters and retains the others. Test it against a labeled set of your actual publisher URLs before making the normalized value a database key. Keep the original URL even after you compute the key.

Redirects, mobile URLs and syndicated copies need another decision. Two addresses that redirect to the same page can be candidates for one record, but do not make a live redirect request inside a time-critical alert path unless you have a clear timeout and a way to retain the original link. Syndicated articles can share text while carrying different publisher context. Keep the separate publisher rows beneath one reviewed story instead of deleting them.

When a source edits a page in place, an exact URL key will correctly prevent a duplicate new-link alert, but it will also suppress news about the edit. Compare a compact hash of the headline and snippet between observations, then queue a changed row for review. Treat the hash as a change detector, not proof that the underlying article changed meaning. A person should open the source before an update alert goes out.

Cluster likely same-event reports without losing evidence

  1. Bucket candidates within a short review window using normalized terms from the headline and, when present, the calendar date. Treat a missing date as unknown, not today.
  2. Show the analyst each candidate’s publisher, title, snippet, URL and first-seen time. Open the source page before merging. Similar wording may cover a new development.
  3. Assign a story ID after review. Keep one row per publisher URL beneath it and record why two records were joined. When a correction changes the event, split the cluster without losing history.

Budget a sampling run and compare the alternatives

At the Serpent Default News rate of $0.20 per 1,000 units, two one-page queries every hour for 30 days request 1,440 units, or $0.288 in listed usage. Storage, analyst time and any extra pages are separate. Multi-page News charging depends on tier or live balance, so use the billing ledger as the invoice authority. num is capped at 50 and is a ceiling, not a delivered-row promise.

OptionUseful forUnit or access on October 6, 2026Decision
Serpent News SearchSearch-result observations across documented enginesDefault $0.20/1,000 News units; one-page example aboveUse when you need the search-result view and will own clustering.
NewsAPIArticle search with source and date filteringFree Developer plan is development-only, delayed by 24 hours and limited to 100 requests/dayUse its official filters when they match your product and choose a suitable commercial plan for production.
Direct publisher feedsA known publisher list and direct article linksEach publisher sets its own feed and reuse termsUse when completeness for specific publishers matters more than broad discovery.

I worked through the listed Serpent arithmetic as 2 × 24 × 30 = 1,440 one-page requests and 1,440 × $0.0002 = $0.288. That figure buys searches, not verified distinct stories. If your two queries return many repeated publisher URLs, divide your actual monthly bill by the number of reviewed, useful developments to learn the cost per useful alert. Do not use HTTP success or raw row count as the denominator.

NewsAPI documents exact from and to parameters for its Everything endpoint, so it may suit a historical date-bounded search better than a rolling freshness filter. Its published pricing page also offers enterprise story clustering. If you need the provider to supply clusters, compare that enterprise feature directly with your own review labor. A cheaper request price would not settle the decision if your team spends hours correcting false merges.

Keep a reviewable story record, including corrections

Use two IDs: a stable URL key for repeated observations of one publisher page, and a human-assigned story ID for reports about the same development. Do not generate the story ID from headline similarity alone. For example, “manufacturer announces recall” and “regulator expands recall” may share most words while requiring two alerts. A later correction to one publisher page should update that URL’s observation history; it should not erase the first text your alert used.

Queue fieldWhat to saveWhy
Exact-page evidenceOriginal URL, normalized key, publisher, title, snippet, publication date if present and first-seen UTC time.A reviewer can reopen the actual result and trace repeat appearances.
Story decisionStory ID, reviewer, review time and a short merge-or-split reason.Similar headlines can describe separate developments.
Alert historySent time, recipients within your system, and the row versions used.A corrected source can be handled as an update without silently rewriting an old alert.

A practical triage policy is to send an automatic alert for a new exact URL only after its query sample is usable, then group likely same-event reports into a review queue. If the topic is high stakes, have a person check the source page before any outward alert. Count both duplicate URL suppression and false story merges in a labeled weekly sample; a lower alert count is not useful if separate developments were hidden.

Decide whether the deduplication is safe to launch

Before using the queue for real decisions, label a week of results yourself. Mark the true event, whether a row repeats an existing URL, whether it is the same development as another publisher's row, and whether it contains a correction or a new action. Include quiet periods and short responses, not only busy news days.

CheckHow to calculate itExample launch rule to adapt
Exact-repeat suppressionRepeat URL observations suppressed ÷ all repeat URL observations in the labeled set.100% after documented URL normalization, with manual inspection of collisions.
False story mergesDistinct developments incorrectly combined ÷ all proposed story merges reviewed.Below 1% for low-stakes internal triage; require human approval for consequential alerts.
Development captureLabeled distinct developments that entered the alert or review queue ÷ all labeled developments visible in your sampled result windows.At least 95% in the chosen query set before reducing human review.
Coverage caveatRuns with no parsed rows or with a short-delivery signal ÷ scheduled runs.Investigate every such run before claiming that a topic was quiet.

These percentages are proposed acceptance targets, not measured Serpent performance. The denominator for development capture is what your sampled windows and manual source checks exposed; it cannot prove that a search engine indexed every article. If the false-merge threshold fails, keep exact-URL suppression but move cross-publisher grouping back to review.

What can make the alert misleading?

An HTTP success alone does not prove a complete article set. Inspect parsed results.articles, delivery and meta.partialResults. A short or empty sample may reflect the selected query or delivery limit; do not convert absence into a negative claim. A publication date is only YYYY-MM-DD, while meta.timestamp and your own clock record observation time.

A useful acceptance test is a week of manually labeled events: track how many genuine developments made it into the queue, how many duplicate alerts remained, and which clusters merged distinct developments. Recheck official pricing and content rights before using article text or images beyond internal triage.

For a broader starting point, see the News Search API overview; the media monitoring workflow covers a related task.

The rolling news freshness guide covers repeat snapshots and first-seen dates.

Build a reviewable news alert queue

Start with a small query set, keep the raw rows, and review uncertain clusters before sending alerts.

Get an API key

Try the playground · Read the API reference

FAQ

Should I deduplicate news by headline alone?

No. Headlines change, and different publishers may use similar wording for separate reports. Use a stable URL key for exact repeats, then review likely same-event clusters with publisher, date and article context.

Does a missing publication date mean the article is old?

No. Serpent returns publishedTime as a calendar date or null when the search result does not carry one. Record the observation time separately and leave the publication date unknown.

Will this find every story about my topic?

No. It samples a selected query, engine, country and result window. Keep those inputs with the rows and treat zero or short delivery as limited evidence.

How are news searches billed?

Serpent lists News at $0.20 per 1,000 Default units. A one-page call uses one unit; multi-page charging depends on tier or live balance. Check the current pricing page and billing ledger before scaling.