User-agent: *
Disallow:
Disallow: /language/ro
Disallow: /language/ru
Disallow: /language/en
Disallow: /en/widget/*
Disallow: /ro/widget/*
Disallow: /ru/widget/*
# Login-walled company funnels. Both answer an anonymous request with
# 307 -> /{locale}/login, so they can never be indexed, yet they are linked
# from every company page and so form a second ~382M x 16-locale crawlable
# surface that spends crawl budget on nothing. Measured 2026-07-31: 3,779
# requests, 1.9% of ALL traffic, in one 200,000-request nginx window.
# Two lines per prefix: robots.txt matches a literal prefix, so the bare form
# and the locale-prefixed form must both be named.
# See lib/crawlBudgetGuard.ts.
Disallow: /claim-company/
Disallow: /*/claim-company/
Disallow: /edit-company/
Disallow: /*/edit-company/
# The `?backUrl=` auth funnel. Every public company page renders an anonymous
# "create a free account" link whose href carries THAT PAGE as a query
# parameter, so /register alone has the cardinality of the company index
# (~382M x 16 locales) and /search adds one per distinct query. Measured
# 2026-08-03 in one ~21.5h nginx window: 155,867 GET /*/register?...backUrl=
# plus 23,586 GET /*/login?...backUrl= — 179,453 bot SSR renders, against 129
# real registrations. None of it can ever be indexed (/login and /register both
# ship noindex,follow), so every fetch displaces a real company page.
# One wildcard line covers /register, /login, /pricing and /payorder in the
# bare form and in all 16 locale-prefixed forms; the bare parameterless
# /{locale}/login and /{locale}/register stay crawlable, as they cost nothing.
# See lib/authFunnelCrawlGuard.ts.
Disallow: /*backUrl=
# The anon-view gate beacon. Since 2026-08-04 its payload rides in the QUERY
# STRING (a body-carrying POST was being truncated and answered 400 by nginx,
# losing ~1,500 views/day — lib/anonViewBeaconPayload.ts), which made the URL
# unique per company: ~382M of them. A body is invisible to a crawler, a query
# string is not, and JS-rendering crawlers re-request the URLs a page fetches.
# Measured: 0 GETs on this route on 08-03 and 0 before the 21:52 deploy on
# 08-04, then 37 from bingbot carrying 35 distinct companyKey values in the
# ~40min after it. Blocking this ONE route costs a crawler nothing — a
# recognised crawler is already short-circuited out of metering before the
# route does any work (lib/crawlerRequest.ts) — which is why it is safe to
# disallow here and NOT safe to disallow /api/ wholesale, since Googlebot
# applies robots.txt to subresources during rendering.
# See lib/anonViewBeaconCrawl.ts.
Disallow: /api/anon-view/
# The /compare company deep-link. Every company page's header renders
# Compare, so the `c` parameter
# gives /compare the cardinality of the company index — ~382M x 16 locales.
# Measured 2026-08-05 over a 23.6h nginx window (7,154,530 requests): 131,799
# GET /*/compare, 121,752 of them carrying exactly ONE `c=` param, i.e. the
# followable header link; the two deep-links that were already nofollowed emit
# the multi-`c=` shape and account for the other 10,040. Of the 87,279 distinct
# IPs that asked, 128 (0.15%) ever fetched a JS chunk — this is a synthetic
# fleet, not visitors. None of it is indexable: /compare?c=… already
# self-canonicalises to the bare /compare (lib/canonicalPath), so every one of
# those ~0.57s SSR renders displaces a fetch of a real company page.
# Two lines per prefix: robots.txt matches a literal prefix, so the bare form
# and the locale-prefixed form must both be named. The parameterless
# /{locale}/compare stays crawlable — it is the canonical target and costs
# nothing. See lib/compareDeepLinkCrawl.ts.
Disallow: /compare?c=
Disallow: /*/compare?c=
# The /contact-us removal deep-link. From 2026-08-09 the officer page's privacy
# footer links to /contact-us?reason=1&personId={this person}, so that the
# person whose record it is arrives with the removal reason chosen and their own
# record already resolved instead of re-finding themselves in an index of 207M
# people. The href is therefore unique per officer page x 16 locales, which is
# the same cardinality trap as the four rules above — declared here BEFORE it
# can be crawled rather than after. Baseline measured in the 09/Aug window, on
# the build that shipped without it: 348 requests to /contact-us, 339 bare and
# the 9 with a query carrying only utm_source / __jsfix. Nothing indexable is
# lost: the parameterless /{locale}/contact-us is the canonical form, is the one
# listed in sitemap-pages.xml, and stays crawlable — a query-string form of it
# self-canonicalises there anyway (lib/canonicalPath). Two lines per prefix,
# same reason as above. The bounded pre-existing deep-links (/export-data's
# volume quote, the cancel popup) are covered by the same rules and lose
# nothing, being reachable only by a human already on those pages.
# See lib/contactUsPrefill.ts and components/OfficerV2/PrivacyFooter.tsx.
Disallow: /contact-us?
Disallow: /*/contact-us?
# The faceted /search surface. Every option in the filter rail — jurisdiction,
# status, legal form, year, contact axis, size band — plus every pagination
# cell is an whose href is the CURRENT filter set plus one more axis, so a
# fetched search URL hands out SEVENTY more in its own HTML (measured live) and
# each of those hands out seventy more. The space is combinatorial: country x
# types x statuses x years x activities x contacts x size x sort x page x q, in
# 16 locales. Measured 2026-08-10 over one nginx day: 300,287 requests to
# /{locale}/search — 4.9% of everything served — of which 300,259 carry a query
# string, 269,661 carry a facet parameter, and 287,232 are DISTINCT URLs (95.6%
# fetched exactly once). That is enumeration, not searching; the referers agree,
# with 75 arrivals from Google web search against 231,008 with no referer.
# Not one of these URLs can ever be indexed: every /search response ships
# noindex,follow and canonicalises to the bare /{locale}/search, and no sitemap
# entry anywhere carries a query string. Yet Googlebot — verified real, its 169
# addresses reverse-resolving into googlebot.com — spent 45,770 requests here,
# 10.8% of its entire crawl of this site, second only to /company. rel="nofollow"
# was already on every link inside the search page's own filter rail and
# pagination and did not stop it, because nofollow has been a hint rather than a
# directive to Google since 2019 — while the SEED links, on the company and
# officer pages that hand the frontier its first URLs, carried no rel at all
# (five files, fixed in the same change). robots.txt is what still binds
# Googlebot, and /search was the one high-cardinality surface with no rule at
# all. Two lines per prefix, same reason as the rules above. The
# parameterless /{locale}/search stays crawlable — it is the canonical target,
# it is what every variant already points at, and it took 28 of the 300,287
# requests. See lib/searchFacetCrawlGuard.ts.
Disallow: /search?
Disallow: /*/search?
# The archive day page's keyset continuation cursor. A day with more than 900
# registrations (73,035 (country, day) rows across 59 jurisdictions, measured
# 2026-08-11) used to end at page 30 with a bare 404 for anyone who tried to
# read further; `?after=` continues from a known row instead, at
# constant cost. Every continuation is `noindex,follow` and self-canonicalizes
# to the bare day URL, and — the part that actually binds the fleet that forges
# its user-agent — THE CURSOR IS NEVER IN SERVER-RENDERED HTML: it is written
# in from an effect after mount, so a client that does not execute JS sees only
# the ordinary archive links it already had. This line is kept for the
# compliant minority, exactly like the five rules above it, and is not what
# makes the surface safe. See lib/archiveContinuation.ts.
Disallow: /*after=
# The two self-healing recovery flags. When edge-cached HTML outlives the build
# whose content-hashed assets it references, a missing stylesheet renders the
# page unstyled and a missing script renders it inert; lib/staleStylesheetRecovery.ts
# re-fetches the page under `?__cssfix=1` and lib/staleScriptRecovery.ts
# re-navigates under `?__jsfix=1`. Both flags force a CDN MISS on purpose, so
# every one of them is a full uncached SSR render at the origin.
#
# Measured 2026-08-11 in one partial nginx day: 14,128 `?__cssfix=1` and 22,671
# `?__jsfix=1` requests. 10,411 of the cssfix ones came from DECLARED crawlers
# (9,422 GoogleOther), 316.7 MB of renders that bought those crawlers nothing —
# a crawler does not paint and does not click, and the SSR HTML it came for is
# already complete. The source of that is now fixed: BOTH modules skip crawlers
# outright, so these URLs stop being generated.
#
# This rule is for the ones ALREADY discovered. Unlike the forged-user-agent
# fleets the rules above cannot bind, the clients here are compliant declared
# crawlers — GoogleOther is 9,422 of them — which is exactly the population a
# Disallow does work on. Both flags self-canonicalize to the clean URL
# (Head.Component strips the query), so nothing indexable is lost.
Disallow: /*__cssfix=
Disallow: /*__jsfix=
# The three sitemaps, and what each one is FOR — they do not overlap:
#
# sitemap_index.xml.gz company pages only. 3,984 child sitemaps, all named
# sitemap_{country}_{n}.xml.gz, 100% /company/ URLs
# (verified 2026-08-07 by enumerating the whole index
# and opening four children across four jurisdictions).
# Generated by b2bhint-cron-scripts, not by this repo.
# sitemap.xml a hand-maintained static file: the 15 locale
# homepages. nginx aliases this URL straight to
# public/sitemap.xml, so it can never be a Next route.
# sitemap-pages.xml EVERYTHING ELSE, added 2026-08-07 because it was in
# NO sitemap in ANY locale: /pricing, /about, /archive,
# /compare, /export-data, /list-match, /openindex,
# /api-reference, /data-sources (index + every
# per-registry citation page), /contact-us, the legal
# pages, every country landing page, and the six
# cross-portfolio exposure searches. Built from the
# app's own canonical slug builders so no listed URL is
# one 301 away from its canonical form; noindex and
# session-gated routes are excluded by construction.
# See lib/staticSitemap.ts.
Sitemap: https://b2bhint.com/sitemap_index.xml.gz
Sitemap: https://b2bhint.com/sitemap.xml
Sitemap: https://b2bhint.com/sitemap-pages.xml