One Clip, Many Links: Build a Video Search API Inventory

By Anurag Pathak··

To build a video search API inventory across sites, save every result as an observation, then connect observations to a cautiously identified video item. Two links to the same upload can share an item; a re-edited clip deserves its own item even when the title looks identical. Your finished inventory is a reviewable sample of search visibility, not a count of everything that exists online.

If you monitor training videos, competitor clips or publisher coverage, this distinction saves you from announcing three “new” videos when one upload simply appears at three URLs. I reviewed Serpent's Video Search API documentation and documented endpoint shape on October 6, 2026, and worked through the cost example below. The official YouTube search reference was also checked for this article on October 6. These are documentation and code reviews, not a hands-on comparison of delivered results or cross-site coverage.

First decide what “one video” means to you

A search result tells you that a particular query surfaced a link at a particular time. It does not prove that the link still plays, that the listed uploader owns the work, or that the video is unique. Keep the result's own identity as an observation: query, engine, country, position, timestamp, returned URL and raw fields. Give a separate item identity to the underlying upload only when you have enough evidence. This lets your team revise a duplicate decision without erasing what the search actually returned.

For a practical, fictional case, suppose you track “urban garden tutorial.” Search returns a YouTube watch link with video ID abc123 and later a short youtu.be/abc123 link. Those are two observations of one YouTube upload. A gardening magazine page then embeds that same ID: you can record a third viewing location and link it to the upload after inspecting the page. A 45-second highlight that the magazine edited and uploaded under a different video ID is a related version, not the same item. Matching titles, thumbnails or publishers alone should only create a possible-duplicate review task.

Fictional resultProvisional itemWhy
youtube.com/watch?v=abc123 and youtu.be/abc123One upload; two observationsThe parsed YouTube ID agrees. Keep both original URLs.
Magazine article embedding abc123Same upload after page reviewThe page is another location, not automatically another upload.
Magazine's edited highlight with a new IDSeparate item, linked as a versionThe clip has a different edit and upload identity.

Decide whether your unit is an upload, a creative work, or a viewing page before you deduplicate. A brand team may want one row per creative work with all cuts beneath it. An engineering catalog usually needs one item per platform upload plus a related_version_of link. Neither model can be inferred safely from titles alone.

Read the fields without inventing what is missing

Serpent's documented Videos endpoint searches video results from Google, DuckDuckGo, Yahoo, Bing or Brave; ddg is the default. For this inventory, request format=full. The full row names position, title, url, duration, source, views, thumbnail, description, publisher, embedUrl and publishedTime. Several values may be null. Use the returned watch url as a candidate identifier; treat source as a location or host hint and publisher as an uploader or channel label when present.

FieldUseful roleSafe handling
urlOriginal watch link and identity candidateIf absent, retain the observation but queue identity for review.
source, publisherLocation versus uploaderKeep them separate; a platform name does not identify a creator.
views, duration, publishedTimeOptional context for a reviewerNull means unknown, not zero or “just uploaded.”
thumbnail, embedUrlReview aids and possible player linksNeither confirms playback, ownership or embedding permission.

The num parameter asks for up to a chosen count, with 100 as its documented ceiling. A shorter parsed list is not proof that no other matching videos exist. Save the requested count, actual row count, meta.partialResults and delivery alongside the run. If you need to inspect how dates change across runs, the video freshness audit explains the separate publication-date and observation-time questions.

Store items and observations in separate tables

Your video_items table needs an item key, the original watch URLs you have linked, a review state and perhaps a confirmed platform ID. Your search_observations table needs a run ID, query controls, position, observed time, raw response row and an optional item key. One item can have many observations. A row without a usable URL still belongs in observations; it just cannot be confidently attached to an item yet.

Make the merge reversible. Store a reviewer, date and reason when you link two URLs; retain the original observations even if a later reviewer splits an item. Keep uncertain cross-platform matches in a queue such as possible_same_work. If you overwrite the raw source URL with a normalized URL, you lose the evidence needed to undo a bad merge or investigate a broken page.

The following Python is an illustrative local example for one result window. Set SERPENT_API_KEY, install requests with python -m pip install requests, and replace the query. It parses an obvious YouTube ID when present and otherwise uses a lightly cleaned URL. It deliberately makes no automatic cross-site merge. The output shape below is fictional; I have not used this script to measure live result quality.

import json, os
from datetime import datetime, timezone
from urllib.parse import urlsplit, parse_qs, parse_qsl, urlencode, urlunsplit
import requests

def item_key(raw):
    try:
        u = urlsplit(raw)
        if u.scheme not in ("http", "https") or not u.hostname: return None
        host = u.hostname.lower().removeprefix("www.")
        vid = (parse_qs(u.query).get("v", [None])[0]
               if host == "youtube.com" and u.path.rstrip("/") == "/watch" else None)
        if host == "youtu.be": vid = u.path.strip("/").split("/")[0]
        if vid: return "youtube:" + vid
        query = [(k, v) for k, v in parse_qsl(u.query, keep_blank_values=True)
                 if not k.lower().startswith("utm_") and k.lower() not in {"fbclid", "gclid"}]
        return urlunsplit(("https", host, u.path.rstrip("/") or "/",
                           urlencode(sorted(query)), ""))
    except ValueError:
        return None

params = {"q": "urban garden tutorial", "engine": "ddg",
          "country": "us", "num": 25, "format": "full"}
response = requests.get("https://apiserpent.com/api/videos", params=params,
    headers={"X-API-Key": os.environ["SERPENT_API_KEY"]}, timeout=60)
response.raise_for_status()
data = response.json()
if not isinstance(data, dict) or data.get("success") is not True:
    raise ValueError("No usable video answer")
result = data.get("results")
if not isinstance(result, dict): raise ValueError("Missing full-format results")
rows = result.get("videos")
if not isinstance(rows, list): raise ValueError("Missing video list")
if not rows: raise ValueError("Empty parsed list; inspect delivery before treating it as no videos")
observed_at = datetime.now(timezone.utc).isoformat()
items, observations, exceptions = {}, [], []
for position, row in enumerate(rows, 1):
    key = (item_key(row.get("url"))
           if isinstance(row, dict) and isinstance(row.get("url"), str) else None)
    observations.append({"observation_id": f"{observed_at}:{position}",
                         "position": position, "item_key": key,
                         "observed_at": observed_at, "raw_row": row})
    if key is None:
        exceptions.append({"position": position, "reason": "no usable URL"})
        continue
    item = items.setdefault(key, {"key": key, "watch_urls": [], "observation_count": 0})
    item["observation_count"] += 1
    if row["url"] not in item["watch_urls"]: item["watch_urls"].append(row["url"])
warnings = (["Fewer parsed rows than requested; review coverage and delivery"]
            if len(rows) < params["num"] else [])
print(json.dumps({"request": params, "observed_at": observed_at,
                  "items": list(items.values()), "observations": observations,
                  "exceptions": exceptions, "warnings": warnings,
                  "partialResults": (data.get("meta") or {}).get("partialResults"),
                  "delivery": data.get("delivery")}, indent=2))

Abbreviated fictional output: {"items":[{"key":"youtube:abc123","observation_count":2}],"observations":[{"item_key":"youtube:abc123"},{"item_key":"youtube:abc123"}],"warnings":["Fewer parsed rows than requested; review coverage and delivery"]}. In a real run, retain the full raw rows and inspect exceptions. A short list can reflect the result window or delivery; it is a prompt to review, not a license to manufacture missing rows.

Measure whether the sample helps your actual job

Choose a small, labeled reference set before automating a large catalog. For example, assemble 20 known, currently accessible videos across two sites for your topic. Write down their URLs and which edited versions should remain separate. Run fixed queries with the same engine, country and date window; repeat at a known cadence. Open a sample of returned watch pages to check playback and inspect possible duplicates. Keep the search terms and controls with every run so that a change in query does not masquerade as a change in coverage.

Acceptance measureHow you calculate itWhat it tells you
Known-link recoveryKnown items found ÷ known items eligible for this query. In a fictional trial, 14 ÷ 20 = 70%.Coverage of your labeled set for this query and time, not web-wide recall.
Duplicate-decision precisionCorrect proposed merges ÷ proposed merges reviewed. Fictional example: 18 ÷ 20 = 90%.How much human correction the identity rules need.
Usable-row shareRows with a usable URL and required fields ÷ parsed rows.Whether returned volume translates into a reviewable inventory.
Observed playback shareWatch pages that play in your sampled manual review ÷ pages opened.Current accessibility of the reviewed subset only.

Compare those measures separately by source host and query, because a single percentage can hide a weak part of the inventory. If a result is missing, try a later controlled run before marking a known video as removed. If your team needs complete coverage of a known channel or publisher, define that source's own catalog as the denominator and use its authorized platform API or export. A general search sample cannot promise that completeness.

When YouTube's official API is the better fit

If your project is YouTube-only, start with YouTube's official search.list documentation and its search result resource. They describe YouTube resource types and platform IDs, including video and channel identifiers, that are useful for a YouTube catalog. For known channels, decide whether search discovery is even the right source of truth for your task. That API does not inventory other sites for you.

ApproachUse it whenWhat to budget
Serpent VideosYou need a normalized sample of video results surfaced across search engines.Listed calls plus storage, URL checks and human duplicate review.
YouTube's official APIYou need YouTube resource IDs, channel context or a YouTube-specific workflow.The current project quota rules and your full request sequence.
Authorized publisher catalogYou know the source set and need its own complete inventory.Its access terms, permissions, pagination and update work.

The official YouTube quota documentation has changed over time, so I would use its current project console and documentation for a new estimate rather than copy an older per-search unit figure. YouTube quota units and Serpent billable calls are different units; compare the workflow you actually need.

For a Serpent example, five one-page queries per day for 30 days mean 5 × 30 = 150 calls. At the listed Default Videos rate of $0.10 per 1,000 calls, 150 × $0.10 ÷ 1,000 = $0.015 of listed usage. I checked that arithmetic on October 6, 2026. It excludes storage, manual verification and any plan-specific billing details; it is not a measured cost per usable video.

Start with a reviewable first week

Pick one topic, two or three deliberately different queries, a fixed engine and country, and one daily run. Preserve every raw observation. Let the script create only the high-confidence same-ID merges, then review cross-site pairs and edited versions by hand. After a week, inspect your known-link recovery, merge precision, nullable-field rates and short-result runs. Expand the query set only when you know which gap it addresses. The competitor monitoring guide covers how to compare repeated runs; the Python tutorial is a smaller starting point if you first need to make a single request.

Try one video inventory sample

Fetch a small result window, save its raw observations and review your first duplicate decisions before scheduling more calls.

Get an API key

Try the playground · Read the Videos reference

FAQ

Can a video search API find sites beyond YouTube?

Serpent's documented Videos endpoint searches five engines and returns video links with a source field. It can surface sites represented in those results, but it is a search sample rather than a full catalog of every platform.

Is publisher the same as source?

No. Source is a location or host hint; publisher is an uploader or channel label when the result provides one. Keep both fields, and leave either null when unknown.

Should two pages with the same clip be one inventory item?

Only when your item means the underlying upload or work and you have evidence that the pages refer to it. Keep each page as its own observation. Treat an edited or newly uploaded version as a separate item linked to the original.

When should I use YouTube's official API instead?

Use it for YouTube-specific resource IDs, channel context or platform workflows. Check current quota documentation for your exact request sequence, and use another source if you must cover sites outside YouTube.