Methodology

SiteIndiq reports mix two very different kinds of data. Knowing which is which is the difference between using these numbers well and being misled by them.

Measured live

Every scan connects to the domain directly. These figures are facts observed at scan time, not estimates:

  • DNS records — A, AAAA, MX, NS, TXT and SOA resolved through public resolvers (1.1.1.1, 8.8.8.8, 9.9.9.9).
  • SSL certificate — read from a real TLS handshake: issuer, validity window, covered hostnames, protocol and cipher.
  • HTTP behaviour — the full redirect chain, status codes, response headers, compression and time to first byte.
  • On-page SEO — title, meta description, canonical, headings, word count, image alt coverage and structured data, parsed from the homepage HTML.
  • WHOIS — registrar, creation, update and expiry dates queried from the registry.
  • robots.txt — fetched and evaluated per crawler.
  • Hosting — IP geolocation from the DB-IP City Lite database and network owner via Team Cymru’s ASN service. When a site sits behind a CDN we say so, and label the location as an edge node rather than claiming the site is hosted there.
  • Origin server — for CDN-fronted sites, candidate origin addresses recovered from public records (SPF entries, mail servers, unproxied subdomains, Certificate Transparency logs). Each candidate is then contacted directly, and only reported as confirmed if it answers for that domain.

Modelled, not measured

Nobody outside a website can see its analytics. Traffic, revenue and valuation are therefore statistical estimates, derived in three steps.

1. Popularity rank

We use the Tranco list, a research ranking that averages several public popularity lists over a 30-day window to resist the day-to-day noise that affects any single source. A domain outside the top 1,000,000 has no published rank; how those are handled is set out below.

2. Rank to traffic

Web popularity follows a power law, so monthly visits are modelled as:

monthly_visits = 8 x 10^10 x rank^-1.15
daily_visitors = monthly_visits / 30
daily_pageviews = daily_visitors x 1.8

The two constants are fitted against published traffic figures for well-known sites at ranks 1, 100, 10,000 and 1,000,000. The fit reproduces those anchors within roughly 25% — good enough for order-of-magnitude comparison, not for media buying.

3. Traffic to value

daily_revenue = daily_pageviews / 1000 x eCPM
estimated_worth = daily_revenue x 1825

The eCPM starts at a blended display rate and is adjusted two ways. First by audience geography, proxied from the hosting country and the ccTLD — a US audience monetises several times better than a South Asian one. Second by a monetisation damping factor: sites in the global top few thousand are search engines, CDNs, ad networks and app back-ends that do not run display ads at all, so a flat eCPM would overstate them enormously.

The 1,825-day multiplier is five years of projected earnings — the convention these calculators use, and roughly in line with how content sites actually trade (typically 30–45x monthly profit).

Domains with no rank

Most of the web sits outside the top 1,000,000 and has no published rank. Rather than leave those reports blank, we estimate where the domain would sit on the same curve, using signals we actually measured on it:

  • How many pages it publishes, counted from its XML sitemap — the strongest size signal available without analytics.
  • Registration age from WHOIS: a domain renewed for fifteen years is rarely dormant.
  • Technical health, its score on the audit below.
  • TLD, HTTPS and indexability, which shift the estimate up or down at the margin.

The result is a modelled rank, usually somewhere between 1 and 30 million, which then feeds the same traffic and revenue formulas. These reports are labelled low confidence and show a value range, because a modelled rank is a much weaker input than a measured one. Treat them as an order of magnitude, not a price.

A domain that does not respond at all is reported as having no measurable traffic or advertising value, and its report is excluded from search engines rather than published as a thin page.

SEO health score

The score is fully measured — no estimation. 27 weighted checks cover reachability, HTTPS and certificate validity, title and meta description presence and length, heading structure, mobile viewport, canonical tags, structured data, robots.txt, indexability, sitemap availability, compression, response time, security headers, image alt text, content volume, redirect count, SPF/DMARC and domain expiry. The score is the share of weight earned, mapped to a letter grade.

Limitations you should assume

  • Estimates describe advertising potential only. A SaaS, marketplace or subscription business will be worth far more than its report suggests; a parked or affiliate domain, often less.
  • Only the homepage is fetched. A site with a weak homepage but strong interior pages will score lower than it deserves.
  • Rankings update roughly monthly, so a site that grew or collapsed in the last few weeks will lag.
  • Subdomains inherit nothing from their parent domain — they are almost always unranked.
  • WHOIS output is heavily redacted for most gTLDs under GDPR, so registrant details are frequently unavailable.

← Analyze a website