We Audited 14 AI Rank Trackers Against 7 Transparency Questions. Most Fail.

By Serpent API Team · · 15 min read

"If you use three different tools and give them the same prompts, you get three different answers." That's Paul Dyer, CEO of the agency /prompt, in Digiday's May 2026 investigation of AI visibility tools.

He's not exaggerating. Agency testing cited by Brainlabs found "less than a 1 in 100 chance that any of these AI tools will give you the same list of brand recommendations in any two runs of the same prompt." Meanwhile, per Digiday's reporting, investors have poured some $227 million into the category.

So instead of writing another listicle, we did the thing nobody in the "best AI rank trackers" search results has done: we read all 14 vendors' official documentation — pricing pages, help centers, methodology docs, changelogs — and scored every tool against the same 7 transparency questions. Then we cross-checked what real users say on G2, Trustpilot and OMR. Every claim below carries a source, checked on July 21, 2026.

TL;DR: No tool of the 14 fully discloses how its numbers are made. Zero publish confidence intervals; zero commit to announcing methodology changes; 8 of 14 never name the model versions they query. The most transparent docs belong to GetMint and Peec AI; the least to BrandRank.AI (no public pricing at all). And the advertised-vs-real price gap runs up to 4–5× — Ahrefs' "$199/mo" becomes $828/mo configured for all platforms.

What we did (and deliberately didn't)

This is a documentation-and-evidence audit, not a hands-on trial. We say that up front because "tested" claims in this category are usually vendor marketing.

What we did: on July 21, 2026 we read the official public surfaces of all 14 tools — homepage, pricing, help center, developer docs, changelog — and scored each against the 7 questions below. If a vendor's docs are silent on a question, the score says "not disclosed." Silence is the finding; we inferred nothing from marketing tone. In parallel, we pulled ratings and verbatim quotes from live review platforms (G2's own structured data, Trustpilot, Germany's OMR), and flagged every number we could not verify rather than publishing it.

What we didn't do: run the 14 dashboards side-by-side ourselves. That's a different (and worthwhile) project — but it wouldn't answer the question this audit asks, which is: can a buyer tell, from what the vendor publishes, how the numbers on the dashboard are made?

The tools: Otterly.AI, Scrunch AI, AthenaHQ, Knowatoa, ZipTie.dev, Rankscale, Peec AI, Profound, Evertune, BrandRank.AI, Semrush AI Toolkit, Ahrefs Brand Radar, SE Ranking (AI Results Tracker / SE Visible) and GetMint.

The 7 transparency questions

The rubric comes from Metehan Yesilyurt's July 2026 analysis of how AI visibility tools collect data, which argues transparency "reduces to whether the company will answer seven questions in writing":

1. Which collection method do you use per channel — API, UI execution, or hybrid?
2. Which exact model or product version is queried per channel?
3. From which countries and languages are prompts executed, and can I control that?
4. How many runs per prompt per period, and how are rates calculated from them?
5. How is a "mention" detected (exact string, alias list, fuzzy match) — and can I see the raw answers?
6. When your method changes, do you announce it and annotate historical data?
7. Can I export the raw data and reproduce your aggregates myself?

As Yesilyurt puts it: "A vendor with real sampling will answer with concrete numbers. A vendor without it will answer with adjectives." We asked the questions of their public docs instead of their sales teams — a vendor that answers them in writing, in public, is a vendor you don't have to take on faith.

The scorecard: 14 tools × 7 questions

Yes = documented concretely · Partial = some elements documented · ND = not disclosed (docs are silent) · No = docs state the opposite. Scored from official vendor pages only, accessed July 21, 2026.

ToolQ1 MethodQ2 ModelsQ3 GeoQ4 Runs+mathQ5 Mentions+rawQ6 ChangesQ7 Export
GetMintYesYesYesPartialPartialPartialYes
Peec AIYesPartialYesYesYesPartialPartial
SE RankingPartialNDYesYesYesNDPartial
EvertuneYesYesNDYesNDNoPartial
Ahrefs Brand RadarYesPartialPartialPartialPartialNDPartial
ProfoundPartialNDYesYesPartialNDPartial
RankscalePartialYesPartialPartialNDPartialPartial
Otterly.AINDNDYesPartialPartialPartialYes
Scrunch AIPartialNDPartialPartialPartialPartialPartial
Semrush AI ToolkitPartialPartialPartialPartialPartialNDPartial
KnowatoaNDPartialPartialPartialPartialNoYes
ZipTie.devPartialNDPartialPartialPartialNDYes
AthenaHQNDNDPartialPartialPartialNDPartial
BrandRank.AINDNDPartialPartialNDNDND

Coded coarsely (Yes = 2, Partial = 1, ND/No = 0; max 14), the field looks like this:

Who's transparent, who isn't

The leaders earn it in different ways. GetMint's developer docs split every channel into interface models (collected from the product surface) versus native-API models with exact IDs — down to strings like claude-sonnet-4-5-20250929 — and even publish which markets are empirically verified or blocked per model. Peec documents its collection technique, runs each prompt once every 24 hours, publishes its visibility formula with a worked denominator, and is the only tool with customer-controlled mention detection (tracked names, aliases, even RegEx) spelled out in docs.

The single best scoring disclosure belongs to SE Ranking's SE Visible FAQ: 3 runs per prompt per week, equal weight per response, the full visibility formula with a worked example. That's what "answering with concrete numbers" looks like — and it's the exact denominator information most rivals withhold.

Honorable mention: Evertune publishes the best model-version mapping in the audit (a dedicated doc pairs each surface with versions like Gemini 2.5 and Claude Sonnet 4.6, with per-model reasons for API-vs-app collection) and discloses real sampling: 100 runs per prompt per model. Its weak spots are the other half of the rubric — geography undisclosed, no export API, and a stated policy of adapting methodology "in real time" without annotations.

At the bottom: BrandRank.AI's /pricing URL returns a 404 — there is no public pricing, no docs subdomain, and the deepest official material is a marketing FAQ. AthenaHQ publishes its formulas but nowhere discloses run frequency, which makes its credit-based plans ("1 credit = 1 AI response", 3,600 credits at $295/mo) impossible to convert into a guaranteed prompt count. Daily runs across its 9 engines would be ~13 prompts; weekly would be ~111. That's an 8× uncertainty in what you're buying.

One pattern worth naming: the tools that disclose the most also tend to admit the awkward parts. ZipTie's FAQ openly says AI results are personalized and "you may see different results… than what our tracker shows." That candor costs them nothing and buys real trust — the opposite of dashboards that present a constructed metric with three decimal places. As Arber Xhindoli put it in a widely-shared June 2026 critique: "Without the prompt list, runs per prompt, geography, model, account state, and scoring formula, the dashboard is showing a constructed metric."

True configured cost: advertised vs real

The second half of the audit priced a realistic configuration for each tool from its official pricing pages. The gaps are not rounding errors.

ToolAdvertised entryRealistic configured costThe catch
Ahrefs Brand Radar"from $199/mo"$828/mo (Lite $129 + all-platforms $699)$199 buys ONE of 7 AI indexes ("$199/month per index"); API-grade export needs Enterprise at $1,499 base
Profound$99/mo Starter$399/mo Growth for 3 enginesHeadline is billed yearly, single-engine, zero exports; no monthly price is published at all
Otterly.AI$29/mo Lite$189/mo Standard + engine add-onsGemini, AI Mode and Claude are paid add-ons ($9–$439/mo each by tier) on top of 4 base engines
Semrush AI Toolkit$99/mo$159/mo for 75 prompts, 1 domainPrice is per domain; +$60/mo per extra 50 prompts; exports capped at 1,000 rows, 10/day
SE Ranking$89/mo add-on$218/mo ($129 Core base + $89)The add-on requires a base plan; 1 check = 1 prompt × 1 engine, so 5-engine coverage divides your quota by 5
Rankscale$20/mo Essentials$99–$385/mo realistically"0.25 credits per engine" is typical, not universal — Claude costs 2 credits (8× the headline); API only from $385
Scrunch AI"$250/mo"$300/mo month-to-monthThe advertised number IS the annual-billed rate; raw-export API is custom-priced Enterprise
GetMint$80/mo Starter$349/mo beyond 50 prompts4.4× jump to tier 2 with nothing in between; the (well-documented) API is sales-gated
Knowatoa$59/mo StarterNot computableNo prompt quotas published on any plan — volumes live behind "Schedule a call"
AthenaHQ$295/mo StarterNot computableCredits without a disclosed run cadence = indeterminate unit cost (8× spread)
ZipTie.dev$69/mo Basic$69–$159/mo (honest!)Cleanest unit economics in the audit: 1 search = all 3 engines, no add-ons. But 1 seat, no API
Peec AI$95/mo Starter (US)$130/mo with one extra modelPer-model add-ons ($35–$165/mo each, scaling with plan); currency display is geo-localized
Evertune$800/mo Pro$800/mo+ (Enterprise custom above)No low tier exists — but you get disclosed 100-runs-per-prompt sampling for the money
BrandRank.AINone publishedUnknownPricing page 404s; fully demo-gated

Every figure above comes from the vendor's live official pages on July 21, 2026 — links in the references. Prices in this category move constantly; treat any un-dated pricing table (including ones in other roundups) as stale until you've clicked through.

What real users say

Documentation tells you what a vendor promises. Reviews tell you what happens after the invoice. Three themes dominate the verified ones.

1. Pricing surprises are the #1 complaint. The single best example is an AthenaHQ reviewer on G2 (June 2026, 5/5) explaining why they switched: "Scrunch started adding fees once we needed Claude and Gemini coverage. AthenaHQ published what the plan included, we signed up, and we got exactly that." The add-on-engine pattern our cost table documents is exactly what this buyer hit. An Otterly reviewer (4/5, June 2026): "At this price point, I would expect more concrete guidance and strategic recommendations rather than just reporting and tracking."

2. Data stability is a real, documented problem. The most serious verified review in our research is a June 19, 2026 Trustpilot report from a $517/mo Semrush subscriber who found that already-exported historical data from the AI Brand Performance module changed between two exports of the same period: "Gemini share of voice dropped from 8.2% to 0%… This is the same historical period showing different data on two different dates. Historical data should not change." We can't know the root cause from outside — but recall that zero of 14 vendors commit to annotating historical data when methodology changes. This review is what that blind spot looks like from the customer's side.

3. Accuracy claims don't survive spot checks. A Writesonic review of Ahrefs Brand Radar (July 2025, updated January 2026 — note Writesonic competes in this space) ran its own brand through it: Brand Radar reported 3 ChatGPT mentions globally where their own logs showed 123. Their verdict: "That's not a small discrepancy. That's a completely different picture."

And a reality check on social proof: of the 14 tools, five have effectively zero third-party reviews (we confirmed live: Evertune, BrandRank.AI, GetMint and Rankscale show no G2 rating; Knowatoa has no G2 listing at all). Profound has by far the largest base (4.5/5 from 1,124 G2 reviews), followed by Scrunch (4.6/5, 73), Otterly (4.8/5, 52) and AthenaHQ (4.9/5, 34). If you see a "4.9 on G2" badge for a tool whose live G2 page shows no rating — and we found exactly that circulating for one vendor — treat the badge as marketing.

Three blind spots the whole industry shares

Blind spot #1: nobody commits to change announcements. Zero of 14 tools promise to announce collection-method changes or annotate historical data. The best have active changelogs that do it in practice (Rankscale's even logs collection-cost changes); Evertune states the opposite policy outright — its methodology "adapts in real time."

Blind spot #2: model versions. 8 of 14 tools never say which model versions produce their data — most docs stop at "ChatGPT." Only Evertune, GetMint and Rankscale publish actual model IDs. This matters because assistants change underneath the trackers: what "your ChatGPT visibility" means shifts every time the default model does, and how each assistant searches the web differs enough to change who gets cited.

Blind spot #3: no denominators, no variance. Only three tools disclose runs-per-prompt (Evertune 100/month, SE Visible 3/week, Peec 1/day). None publish confidence intervals. Given the measured run-to-run instability of LLM answers — the Brainlabs 1-in-100 figure above, and the same non-determinism we've documented for ordinary Google results — a "12% visibility" score with no run count behind it is a vibe wearing a percent sign.

The buyer's checklist

If you're evaluating any tool in this category, do this before paying:

1. Email them the 7 questions and ask for answers in writing. Concrete numbers = real sampling. Adjectives = walk away.
2. Price your real configuration, not the headline: your engine list × your prompt count × your markets × seats. Our table above shows where the multipliers hide.
3. Demand raw-answer access on your tier — 10 of 14 tools show raw answers somewhere, but unrestricted export is usually Enterprise-gated. If you can't see the answers behind the score, you can't audit the score.
4. Screenshot your dashboard monthly. Since no vendor annotates historical data, your own export trail is the only stable record you'll have.
5. Distrust review badges — check the live G2/Trustpilot page yourself; several tools in this space have zero real reviews.

The API alternative (disclosure: ours)

Full disclosure: we sell an API in this space, so read this section as the vendor talking. Serpent's AI Ranking API returns how the major assistants respond to a prompt — the answer, the citations with positions, and where a brand appears — as raw JSON, priced per call with no subscription. We deliberately did not score our own product in the audit above: it's an API rather than a dashboard tracker, and grading our own homework would be exactly the kind of thing this post exists to complain about.

The honest positioning: if you need team workflows, alerts, recommendations and reporting, buy one of the dashboards above — the transparent ones first. If what you actually need is the underlying data on your own prompt set, on your own schedule, with the raw responses in your own warehouse, an API is the cheaper and more auditable route. We've published the full build — prompts, storage, scoring — in the build-vs-buy teardown, and the metric design in AI visibility scoring, explained. For the tool-by-tool feature landscape, our directory comparison and AI Overview tracker roundup stay current, and the engine-specific pages cover Perplexity, ChatGPT and Gemini tracking.

Track AI answers from the source

One API call returns the assistant's answer, its citations and your brand's position — raw JSON you can audit, on prompts you control. 10 free calls, pay-as-you-go after, no subscription.

Get Your Free API Key

Explore: AI Ranking API · Documentation · Pricing

FAQ

Why do AI rank trackers give different answers for the same prompts?

Because every tool makes different undisclosed choices: which model version it queries, how (interface vs API), from where, how many runs per prompt, and how a "mention" is counted. LLM answers are also non-deterministic — agency testing found less than a 1-in-100 chance two runs of the same prompt return the same brand list. Without the methodology, two dashboards legitimately disagree.

Which AI rank tracker is the most transparent about its methodology?

In our July 2026 docs-only audit, GetMint and Peec AI tied for first (11 of 14 points): both disclose how data is collected per engine, and Peec publishes its run cadence, visibility formula and customer-editable mention detection. SE Ranking's SE Visible FAQ has the single best scoring disclosure — 3 runs per prompt weekly with a worked formula.

What does an AI rank tracker really cost after configuration?

Usually far more than the headline. Verified from official pricing pages in July 2026: Ahrefs' "$199/mo" buys one AI platform index — all platforms cost $699/mo plus a $129 base plan, so a realistic setup is $828/mo. Profound's $99 headline is yearly-billed and single-engine; real multi-engine entry is $399/mo. Semrush multiplies per domain.

Do any AI visibility tools publish confidence intervals?

No. Zero of the 14 tools we audited quantify variance on their scores. Only three even disclose how many runs per prompt their numbers are built from: Evertune (100 per model per month), SE Ranking's SE Visible (3 per week) and Peec (1 per 24 hours). Everyone else reports rates without a denominator.

Can I build my own AI rank tracker instead of paying for a dashboard?

Yes — if you mainly need the raw data rather than a team UI. An API that returns how the assistants answer a prompt, which sources get cited and where your brand appears lets you run your own prompt set on your own schedule and keep the raw responses, typically for a fraction of dashboard pricing. Dashboards earn their keep on team workflows, alerts and reporting.