Was That News Story Actually New? A News Search API Freshness Check
A link first seen in your news search API today may have been published yesterday. To explain a late alert, keep the source’s reported publication date, API response time and your monitor’s first-seen time in separate fields. An unknown date stays unknown; none proves the exact posting minute.
I reviewed the News Search documentation, the documented response shape and official NewsAPI pages on October 6, 2026. The scenario and output below are fictional; I did not make an authenticated production call or measure coverage for this article. If repeated links are your next problem, use the story grouping guide.
Pick a window, cadence and claim you can defend
Use freshness=h or 1h for an hourly window, d or 1d for a day, 7d or w for a week, m or 1m for a month, and y or 1y for a year. These are rolling filters, not exact date bounds. The date-only publishedTime may be null. Run an overlapping window so a temporarily absent row can be seen later.
For example, if you poll hourly, a one-hour filter leaves almost no overlap. Start with a day window for hourly snapshots, then narrow it only after your own labeled runs show that the shorter window catches the stories your team needs. Store a first-seen key across polls; overlap without deduplication simply creates repeated alerts.
| Field | Meaning | Use in the monitor |
|---|---|---|
publishedTime | Calendar date YYYY-MM-DD or null | Group by day; never compute minute-level lag. |
meta.timestamp | Time returned with the API response | Record the call context, not publisher time. |
Local observed_at | Your UTC ingestion time | Track when your system first saw a URL. |
delivery / meta.partialResults | Short-answer context when present | Flag the snapshot as incomplete for review. |
Make an hourly snapshot with a stable first-seen record
Schedule this script hourly. Each run requests a rolling day of results and persists URL keys to a JSON file, so successive windows overlap. It prints newly observed links, not newly published stories. It is a small starting point for one process; use an atomic database upsert for concurrent workers. The API fields and filters come from the public contract, but this exact query was not live-tested for this article.
Install requests with python -m pip install requests and set SERPENT_API_KEY before running this illustrative script.
import json, os
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlsplit, urlunsplit, parse_qsl, urlencode
import requests
STATE = Path("news-first-seen.json")
seen = json.loads(STATE.read_text()) if STATE.exists() else {}
now = datetime.now(timezone.utc).isoformat()
response = requests.get("https://apiserpent.com/api/news",
params={"q": "battery recall", "engine": "google", "country": "us",
"freshness": "1d", "pages": 1, "format": "full"},
headers={"X-API-Key": os.environ["SERPENT_API_KEY"]}, timeout=60)
response.raise_for_status()
data = response.json()
if data.get("success") is not True: raise ValueError("No usable news answer")
payload = data.get("results")
if not isinstance(payload, dict): raise ValueError("Missing results object")
rows = payload.get("articles")
if not isinstance(rows, list): raise ValueError("Missing article list")
new = []
for row in rows:
if not isinstance(row, dict) or not row.get("url"): continue
u = urlsplit(row["url"])
if u.scheme not in ("http", "https") or not u.hostname: continue
query = urlencode(sorted((k, v) for k, v in parse_qsl(u.query)
if not k.lower().startswith("utm_")
and k.lower() not in {"fbclid", "gclid"}))
key = urlunsplit(("https", u.hostname.lower(),
u.path.rstrip("/") or "/", query, ""))
if key in seen: continue
seen[key] = {"first_seen": now, "published_date": row.get("publishedTime")}
new.append({"url": row["url"], "title": row.get("title"),
"published_date": row.get("publishedTime"), "first_seen": now})
STATE.write_text(json.dumps(seen, indent=2, sort_keys=True))
delivery = data.get("delivery") or {}
requested = delivery.get("requested")
returned = delivery.get("returned")
sample_short = (returned < requested if isinstance(returned, (int, float))
and isinstance(requested, (int, float)) else None)
if (data.get("meta") or {}).get("partialResults") or not rows:
sample_short = True
print(json.dumps({"new_links": new,
"sample_short": sample_short,
"row_count": len(rows)}, indent=2))
Illustrative output: {"new_links":[{"title":"Example recall update","published_date":"2026-10-05","first_seen":"2026-10-05T09:00:00+00:00"}],"sample_short":null,"row_count":1}. The title and timestamps are fictional; null means this response did not provide a short-delivery comparison, not that coverage was complete. A page can appear in a later snapshot even if its publication date is earlier; label that “first seen,” not “published just now.”
A fictional timeline you can use to word the alert
Imagine you monitor service disruptions for a fictional transit operator. A publisher page says only “October 3” for a report about a station closure. Your monitor first sees its URL on October 4 at 09:00 UTC and sends an alert at 09:05. The honest alert is: “New to our monitor at 09:00 UTC on October 4; publisher date: October 3.” You cannot infer whether the page was posted at 23:59 on October 3, whether it entered the search index later, or whether yesterday’s query window simply missed it.
If the same URL appears in five later hourly snapshots, it remains one discovered link. If a publisher changes the headline on that URL to announce reopening, a URL-only key will hide the update. For topics where revisions matter, retain a compact title or snippet history and flag meaningful changes for human review. If a separate page announces reopening, treat it as a new development even if the wording resembles the closure report.
Make short and empty samples visible to operators
A scheduled monitor should save the query, engine, country, freshness value, row count and any delivery metadata on every run. On a short answer, send a coverage warning to the operator while retaining the rows you did receive. Do not silently mark missing expected publishers as absent. Inspect the source page when a claim depends on exact time or revisions.
Choose the API for your required time controls
| Option | Freshness controls | Billing unit / limit | Best fit |
|---|---|---|---|
| Serpent News Search | Documented rolling freshness values; date-only publication field | Default $0.20/1,000 News units; one-page calls use one unit | Search-result monitoring with your own first-seen store. |
| NewsAPI | Its article endpoints document from, to, sorting and sources | Developer plan is development-only, 24-hour delayed and 100 requests/day | A date-bounded article query where its plan and licensing fit. |
| Publisher feeds | Source-specific entries and timestamps where provided | No common cross-publisher unit | Known sources whose own publication feed is authoritative. |
At one one-page call each hour for 30 days, the plan requests 720 Serpent News units. At the Default listed rate of $0.20/1,000, that is $0.144 in metered usage before storage and review work. More pages can change units according to tier and balance. The account billing ledger decides the actual charge. Recheck both products’ plan terms before launching.
I checked the arithmetic as 1 × 24 × 30 = 720 one-page requests and 720 × $0.0002 = $0.144. If you run three countries, the illustrative request count triples before any extra pages. This is a listed-use estimate, not an observed bill or a cost per useful article. Use your account ledger and the parsed results from a labeled trial to calculate those.
The choice turns on the question you must answer. If you need to ask “what did the search results show in my selected rolling window?”, Serpent's documented News endpoint gives you that sample. If you need an exact historical interval and source or domain selectors, the NewsAPI Everything documentation describes those controls and a UTC publishedAt timestamp. Its Developer plan is for development and testing, includes a 24-hour article delay, and caps requests at 100 per day. Choose a production plan if that product fits a live workflow. If your requirement is every item from three known publishers, their own feeds may be a better ground truth than either broad search sample.
Recover late discoveries without inventing publication times
Use an overlapping window and store both first_seen and the source’s date-only publishedTime. A row first seen at 10:00 with a publication date of yesterday is a late discovery by this monitor; it is not proof that the publisher posted it at 10:00 or that search indexing took a measured number of hours. A null date remains unknown even when the row is new to your database.
| Monitor event | Operator action | Report wording |
|---|---|---|
| New URL with today’s publication date | Open the source if a precise timeline matters. | “First observed at 10:00 UTC; source date is today.” |
| New URL with an older or null date | Keep the row and inspect the article before calling it breaking news. | “New to this monitor,” with source date shown as supplied or unknown. |
| Short, partial or empty scheduled run | Save the run record and check the following overlapping window; alert the operator about coverage. | “Sample incomplete or empty”; no assertion that the topic had no articles. |
The example JSON file keeps first-seen state for one process. In a shared worker setup, use an atomic insert keyed by normalized URL and query so two workers cannot both send “new” alerts. Retain a run log even when zero rows arrive, and recheck the publisher page when an alert depends on exact timing or a later correction.
Keep the clocks and query context together
A useful run record includes the exact query, engine, country, freshness value, scheduled time, request start, response time, parsed row count and any short-delivery marker. For each URL, retain the first-seen UTC instant, the latest-seen instant, source date if supplied, and the raw title and snippet you used for an alert. If you later change the query or country, start a new coverage series. Otherwise a “late” result may simply have entered because you broadened the search.
Calendar dates also need careful presentation. If a publisher shows October 3 without a time zone, do not turn that into midnight UTC. Keep the date as a date. If the publisher offers an actual timestamp, store its original offset as well as a normalized UTC value and note where you obtained it. That lets you compare clocks without inventing precision that the search response never supplied.
Decide who gets an alert before increasing frequency. A daily digest can report “first seen since the last digest,” while an incident desk may need a rapid notification plus a manual source check. Each extra poll samples the selected result window; it does not turn the search index into a complete publisher feed. For named publishers where missing one story is costly, monitor their own feeds or pages alongside the broader search sample.
Acceptance check before calling it a freshness monitor
- For a week, maintain a small manual log of publication URLs and dates from publishers you care about. Mark sources with a real timestamp separately from date-only pages.
- Compare each API snapshot by URL and parsed publication date, not HTTP status. Count rows first seen later than the source date, repeated alerts and runs with no rows or short delivery.
- Calculate capture as manually logged links discovered within your follow-up window divided by all manually logged links in scope. Set a target that fits your use case, such as 90% within 24 hours for a broad internal digest. This is a proposed gate, not a measured result.
- Track the unknown-date share and short-run share alongside capture. Investigate each miss: query wording, ranking window, date filter, source revision and ingestion failure can have different fixes.
- Report minute-level delay only for sources with trustworthy full timestamps and matching time zones. A calendar date alone cannot establish a median lag.
If capture is too low, widen the rolling window or refine the query set before increasing polling frequency. More frequent calls to an unhelpful query add cost without making its results complete.
For a broader starting point, see the News Search API overview; the media monitoring workflow covers a related task.
Start a dated news snapshot
Keep the raw results and observation time so later alerts can distinguish new findings from newly published work.
Get an API keyFAQ
Can I pass exact from and to timestamps to Serpent News?
No. The documented News endpoint accepts rolling freshness values such as h, d, 7d, m and y. It does not document arbitrary from/to parameters.
Is publishedTime a full timestamp?
No. It is a YYYY-MM-DD calendar date or null. Store your own UTC observed_at timestamp if you need a monitor clock.
Does freshness guarantee only newly published articles?
No. Search ranking, source dates and the rolling filter can change between runs. Deduplicate on a durable URL key and review the source page for time-sensitive claims.
How should I handle a null publication date?
Keep it unknown, retain first_seen separately and let a reviewer inspect the publisher page if the exact timing affects a decision.






