We Audited 14 AI Rank Trackers Against 7 Transparency Questions. Most Fail.
"If you use three different tools and give them the same prompts, you get three different answers." That's Paul Dyer, CEO of the agency /prompt, in Digiday's May 2026 investigation of AI visibility tools.
He's not exaggerating. Agency testing cited by Brainlabs found "less than a 1 in 100 chance that any of these AI tools will give you the same list of brand recommendations in any two runs of the same prompt." Meanwhile, per Digiday's reporting, investors have poured some $227 million into the category.
So instead of writing another listicle, we did the thing nobody in the "best AI rank trackers" search results has done: we read all 14 vendors' official documentation — pricing pages, help centers, methodology docs, changelogs — and scored every tool against the same 7 transparency questions. Then we cross-checked what real users say on G2, Trustpilot and OMR. Every claim below carries a source, checked on July 21, 2026.
TL;DR: No tool of the 14 fully discloses how its numbers are made. Zero publish confidence intervals; zero commit to announcing methodology changes; 8 of 14 never name the model versions they query. The most transparent docs belong to GetMint and Peec AI; the least to BrandRank.AI (no public pricing at all). And the advertised-vs-real price gap runs up to 4–5× — Ahrefs' "$199/mo" becomes $828/mo configured for all platforms.
What we did (and deliberately didn't)
This is a documentation-and-evidence audit, not a hands-on trial. We say that up front because "tested" claims in this category are usually vendor marketing.
What we did: on July 21, 2026 we read the official public surfaces of all 14 tools — homepage, pricing, help center, developer docs, changelog — and scored each against the 7 questions below. If a vendor's docs are silent on a question, the score says "not disclosed." Silence is the finding; we inferred nothing from marketing tone. In parallel, we pulled ratings and verbatim quotes from live review platforms (G2's own structured data, Trustpilot, Germany's OMR), and flagged every number we could not verify rather than publishing it.
What we didn't do: run the 14 dashboards side-by-side ourselves. That's a different (and worthwhile) project — but it wouldn't answer the question this audit asks, which is: can a buyer tell, from what the vendor publishes, how the numbers on the dashboard are made?
The tools: Otterly.AI, Scrunch AI, AthenaHQ, Knowatoa, ZipTie.dev, Rankscale, Peec AI, Profound, Evertune, BrandRank.AI, Semrush AI Toolkit, Ahrefs Brand Radar, SE Ranking (AI Results Tracker / SE Visible) and GetMint.
The 7 transparency questions
The rubric comes from Metehan Yesilyurt's July 2026 analysis of how AI visibility tools collect data, which argues transparency "reduces to whether the company will answer seven questions in writing":
1. Which collection method do you use per channel — API, UI execution, or hybrid?
2. Which exact model or product version is queried per channel?
3. From which countries and languages are prompts executed, and can I control that?
4. How many runs per prompt per period, and how are rates calculated from them?
5. How is a "mention" detected (exact string, alias list, fuzzy match) — and can I see the raw answers?
6. When your method changes, do you announce it and annotate historical data?
7. Can I export the raw data and reproduce your aggregates myself?
As Yesilyurt puts it: "A vendor with real sampling will answer with concrete numbers. A vendor without it will answer with adjectives." We asked the questions of their public docs instead of their sales teams — a vendor that answers them in writing, in public, is a vendor you don't have to take on faith.
The scorecard: 14 tools × 7 questions
Yes = documented concretely · Partial = some elements documented · ND = not disclosed (docs are silent) · No = docs state the opposite. Scored from official vendor pages only, accessed July 21, 2026.
| Tool | Q1 Method | Q2 Models | Q3 Geo | Q4 Runs+math | Q5 Mentions+raw | Q6 Changes | Q7 Export |
|---|---|---|---|---|---|---|---|
| GetMint | Yes | Yes | Yes | Partial | Partial | Partial | Yes |
| Peec AI | Yes | Partial | Yes | Yes | Yes | Partial | Partial |
| SE Ranking | Partial | ND | Yes | Yes | Yes | ND | Partial |
| Evertune | Yes | Yes | ND | Yes | ND | No | Partial |
| Ahrefs Brand Radar | Yes | Partial | Partial | Partial | Partial | ND | Partial |
| Profound | Partial | ND | Yes | Yes | Partial | ND | Partial |
| Rankscale | Partial | Yes | Partial | Partial | ND | Partial | Partial |
| Otterly.AI | ND | ND | Yes | Partial | Partial | Partial | Yes |
| Scrunch AI | Partial | ND | Partial | Partial | Partial | Partial | Partial |
| Semrush AI Toolkit | Partial | Partial | Partial | Partial | Partial | ND | Partial |
| Knowatoa | ND | Partial | Partial | Partial | Partial | No | Yes |
| ZipTie.dev | Partial | ND | Partial | Partial | Partial | ND | Yes |
| AthenaHQ | ND | ND | Partial | Partial | Partial | ND | Partial |
| BrandRank.AI | ND | ND | Partial | Partial | ND | ND | ND |
Coded coarsely (Yes = 2, Partial = 1, ND/No = 0; max 14), the field looks like this:
Who's transparent, who isn't
The leaders earn it in different ways. GetMint's developer docs split every channel into interface models (collected from the product surface) versus native-API models with exact IDs — down to strings like claude-sonnet-4-5-20250929 — and even publish which markets are empirically verified or blocked per model. Peec documents its collection technique, runs each prompt once every 24 hours, publishes its visibility formula with a worked denominator, and is the only tool with customer-controlled mention detection (tracked names, aliases, even RegEx) spelled out in docs.
The single best scoring disclosure belongs to SE Ranking's SE Visible FAQ: 3 runs per prompt per week, equal weight per response, the full visibility formula with a worked example. That's what "answering with concrete numbers" looks like — and it's the exact denominator information most rivals withhold.
Honorable mention: Evertune publishes the best model-version mapping in the audit (a dedicated doc pairs each surface with versions like Gemini 2.5 and Claude Sonnet 4.6, with per-model reasons for API-vs-app collection) and discloses real sampling: 100 runs per prompt per model. Its weak spots are the other half of the rubric — geography undisclosed, no export API, and a stated policy of adapting methodology "in real time" without annotations.
At the bottom: BrandRank.AI's /pricing URL returns a 404 — there is no public pricing, no docs subdomain, and the deepest official material is a marketing FAQ. AthenaHQ publishes its formulas but nowhere discloses run frequency, which makes its credit-based plans ("1 credit = 1 AI response", 3,600 credits at $295/mo) impossible to convert into a guaranteed prompt count. Daily runs across its 9 engines would be ~13 prompts; weekly would be ~111. That's an 8× uncertainty in what you're buying.
One pattern worth naming: the tools that disclose the most also tend to admit the awkward parts. ZipTie's FAQ openly says AI results are personalized and "you may see different results… than what our tracker shows." That candor costs them nothing and buys real trust — the opposite of dashboards that present a constructed metric with three decimal places. As Arber Xhindoli put it in a widely-shared June 2026 critique: "Without the prompt list, runs per prompt, geography, model, account state, and scoring formula, the dashboard is showing a constructed metric."
True configured cost: advertised vs real
The second half of the audit priced a realistic configuration for each tool from its official pricing pages. The gaps are not rounding errors.
| Tool | Advertised entry | Realistic configured cost | The catch |
|---|---|---|---|
| Ahrefs Brand Radar | "from $199/mo" | $828/mo (Lite $129 + all-platforms $699) | $199 buys ONE of 7 AI indexes ("$199/month per index"); API-grade export needs Enterprise at $1,499 base |
| Profound | $99/mo Starter | $399/mo Growth for 3 engines | Headline is billed yearly, single-engine, zero exports; no monthly price is published at all |
| Otterly.AI | $29/mo Lite | $189/mo Standard + engine add-ons | Gemini, AI Mode and Claude are paid add-ons ($9–$439/mo each by tier) on top of 4 base engines |
| Semrush AI Toolkit | $99/mo | $159/mo for 75 prompts, 1 domain | Price is per domain; +$60/mo per extra 50 prompts; exports capped at 1,000 rows, 10/day |
| SE Ranking | $89/mo add-on | $218/mo ($129 Core base + $89) | The add-on requires a base plan; 1 check = 1 prompt × 1 engine, so 5-engine coverage divides your quota by 5 |
| Rankscale | $20/mo Essentials | $99–$385/mo realistically | "0.25 credits per engine" is typical, not universal — Claude costs 2 credits (8× the headline); API only from $385 |
| Scrunch AI | "$250/mo" | $300/mo month-to-month | The advertised number IS the annual-billed rate; raw-export API is custom-priced Enterprise |
| GetMint | $80/mo Starter | $349/mo beyond 50 prompts | 4.4× jump to tier 2 with nothing in between; the (well-documented) API is sales-gated |
| Knowatoa | $59/mo Starter | Not computable | No prompt quotas published on any plan — volumes live behind "Schedule a call" |
| AthenaHQ | $295/mo Starter | Not computable | Credits without a disclosed run cadence = indeterminate unit cost (8× spread) |
| ZipTie.dev | $69/mo Basic | $69–$159/mo (honest!) | Cleanest unit economics in the audit: 1 search = all 3 engines, no add-ons. But 1 seat, no API |
| Peec AI | $95/mo Starter (US) | $130/mo with one extra model | Per-model add-ons ($35–$165/mo each, scaling with plan); currency display is geo-localized |
| Evertune | $800/mo Pro | $800/mo+ (Enterprise custom above) | No low tier exists — but you get disclosed 100-runs-per-prompt sampling for the money |
| BrandRank.AI | None published | Unknown | Pricing page 404s; fully demo-gated |
Every figure above comes from the vendor's live official pages on July 21, 2026 — links in the references. Prices in this category move constantly; treat any un-dated pricing table (including ones in other roundups) as stale until you've clicked through.
What real users say
Documentation tells you what a vendor promises. Reviews tell you what happens after the invoice. Three themes dominate the verified ones.
1. Pricing surprises are the #1 complaint. The single best example is an AthenaHQ reviewer on G2 (June 2026, 5/5) explaining why they switched: "Scrunch started adding fees once we needed Claude and Gemini coverage. AthenaHQ published what the plan included, we signed up, and we got exactly that." The add-on-engine pattern our cost table documents is exactly what this buyer hit. An Otterly reviewer (4/5, June 2026): "At this price point, I would expect more concrete guidance and strategic recommendations rather than just reporting and tracking."
2. Data stability is a real, documented problem. The most serious verified review in our research is a June 19, 2026 Trustpilot report from a $517/mo Semrush subscriber who found that already-exported historical data from the AI Brand Performance module changed between two exports of the same period: "Gemini share of voice dropped from 8.2% to 0%… This is the same historical period showing different data on two different dates. Historical data should not change." We can't know the root cause from outside — but recall that zero of 14 vendors commit to annotating historical data when methodology changes. This review is what that blind spot looks like from the customer's side.
3. Accuracy claims don't survive spot checks. A Writesonic review of Ahrefs Brand Radar (July 2025, updated January 2026 — note Writesonic competes in this space) ran its own brand through it: Brand Radar reported 3 ChatGPT mentions globally where their own logs showed 123. Their verdict: "That's not a small discrepancy. That's a completely different picture."
And a reality check on social proof: of the 14 tools, five have effectively zero third-party reviews (we confirmed live: Evertune, BrandRank.AI, GetMint and Rankscale show no G2 rating; Knowatoa has no G2 listing at all). Profound has by far the largest base (4.5/5 from 1,124 G2 reviews), followed by Scrunch (4.6/5, 73), Otterly (4.8/5, 52) and AthenaHQ (4.9/5, 34). If you see a "4.9 on G2" badge for a tool whose live G2 page shows no rating — and we found exactly that circulating for one vendor — treat the badge as marketing.
Three blind spots the whole industry shares
Blind spot #1: nobody commits to change announcements. Zero of 14 tools promise to announce collection-method changes or annotate historical data. The best have active changelogs that do it in practice (Rankscale's even logs collection-cost changes); Evertune states the opposite policy outright — its methodology "adapts in real time."
Blind spot #2: model versions. 8 of 14 tools never say which model versions produce their data — most docs stop at "ChatGPT." Only Evertune, GetMint and Rankscale publish actual model IDs. This matters because assistants change underneath the trackers: what "your ChatGPT visibility" means shifts every time the default model does, and how each assistant searches the web differs enough to change who gets cited.
Blind spot #3: no denominators, no variance. Only three tools disclose runs-per-prompt (Evertune 100/month, SE Visible 3/week, Peec 1/day). None publish confidence intervals. Given the measured run-to-run instability of LLM answers — the Brainlabs 1-in-100 figure above, and the same non-determinism we've documented for ordinary Google results — a "12% visibility" score with no run count behind it is a vibe wearing a percent sign.
The buyer's checklist
If you're evaluating any tool in this category, do this before paying:
1. Email them the 7 questions and ask for answers in writing. Concrete numbers = real sampling. Adjectives = walk away.
2. Price your real configuration, not the headline: your engine list × your prompt count × your markets × seats. Our table above shows where the multipliers hide.
3. Demand raw-answer access on your tier — 10 of 14 tools show raw answers somewhere, but unrestricted export is usually Enterprise-gated. If you can't see the answers behind the score, you can't audit the score.
4. Screenshot your dashboard monthly. Since no vendor annotates historical data, your own export trail is the only stable record you'll have.
5. Distrust review badges — check the live G2/Trustpilot page yourself; several tools in this space have zero real reviews.
The API alternative (disclosure: ours)
Full disclosure: we sell an API in this space, so read this section as the vendor talking. Serpent's AI Ranking API returns how the major assistants respond to a prompt — the answer, the citations with positions, and where a brand appears — as raw JSON, priced per call with no subscription. We deliberately did not score our own product in the audit above: it's an API rather than a dashboard tracker, and grading our own homework would be exactly the kind of thing this post exists to complain about.
The honest positioning: if you need team workflows, alerts, recommendations and reporting, buy one of the dashboards above — the transparent ones first. If what you actually need is the underlying data on your own prompt set, on your own schedule, with the raw responses in your own warehouse, an API is the cheaper and more auditable route. We've published the full build — prompts, storage, scoring — in the build-vs-buy teardown, and the metric design in AI visibility scoring, explained. For the tool-by-tool feature landscape, our directory comparison and AI Overview tracker roundup stay current, and the engine-specific pages cover Perplexity, ChatGPT and Gemini tracking.
Track AI answers from the source
One API call returns the assistant's answer, its citations and your brand's position — raw JSON you can audit, on prompts you control. 10 free calls, pay-as-you-go after, no subscription.
Get Your Free API KeyExplore: AI Ranking API · Documentation · Pricing
FAQ
Why do AI rank trackers give different answers for the same prompts?
Because every tool makes different undisclosed choices: which model version it queries, how (interface vs API), from where, how many runs per prompt, and how a "mention" is counted. LLM answers are also non-deterministic — agency testing found less than a 1-in-100 chance two runs of the same prompt return the same brand list. Without the methodology, two dashboards legitimately disagree.
Which AI rank tracker is the most transparent about its methodology?
In our July 2026 docs-only audit, GetMint and Peec AI tied for first (11 of 14 points): both disclose how data is collected per engine, and Peec publishes its run cadence, visibility formula and customer-editable mention detection. SE Ranking's SE Visible FAQ has the single best scoring disclosure — 3 runs per prompt weekly with a worked formula.
What does an AI rank tracker really cost after configuration?
Usually far more than the headline. Verified from official pricing pages in July 2026: Ahrefs' "$199/mo" buys one AI platform index — all platforms cost $699/mo plus a $129 base plan, so a realistic setup is $828/mo. Profound's $99 headline is yearly-billed and single-engine; real multi-engine entry is $399/mo. Semrush multiplies per domain.
Do any AI visibility tools publish confidence intervals?
No. Zero of the 14 tools we audited quantify variance on their scores. Only three even disclose how many runs per prompt their numbers are built from: Evertune (100 per model per month), SE Ranking's SE Visible (3 per week) and Peec (1 per 24 hours). Everyone else reports rates without a denominator.
Can I build my own AI rank tracker instead of paying for a dashboard?
Yes — if you mainly need the raw data rather than a team UI. An API that returns how the assistants answer a prompt, which sources get cited and where your brand appears lets you run your own prompt set on your own schedule and keep the raw responses, typically for a fraction of dashboard pricing. Dashboards earn their keep on team workflows, alerts and reporting.

