User-agent: * Disallow: Disallow: /language/ro Disallow: /language/ru Disallow: /language/en Disallow: /en/widget/* Disallow: /ro/widget/* Disallow: /ru/widget/* # Login-walled company funnels. Both answer an anonymous request with # 307 -> /{locale}/login, so they can never be indexed, yet they are linked # from every company page and so form a second ~382M x 16-locale crawlable # surface that spends crawl budget on nothing. Measured 2026-07-31: 3,779 # requests, 1.9% of ALL traffic, in one 200,000-request nginx window. # Two lines per prefix: robots.txt matches a literal prefix, so the bare form # and the locale-prefixed form must both be named. # See lib/crawlBudgetGuard.ts. Disallow: /claim-company/ Disallow: /*/claim-company/ Disallow: /edit-company/ Disallow: /*/edit-company/ # The `?backUrl=` auth funnel. Every public company page renders an anonymous # "create a free account" link whose href carries THAT PAGE as a query # parameter, so /register alone has the cardinality of the company index # (~382M x 16 locales) and /search adds one per distinct query. Measured # 2026-08-03 in one ~21.5h nginx window: 155,867 GET /*/register?...backUrl= # plus 23,586 GET /*/login?...backUrl= — 179,453 bot SSR renders, against 129 # real registrations. None of it can ever be indexed (/login and /register both # ship noindex,follow), so every fetch displaces a real company page. # One wildcard line covers /register, /login, /pricing and /payorder in the # bare form and in all 16 locale-prefixed forms; the bare parameterless # /{locale}/login and /{locale}/register stay crawlable, as they cost nothing. # See lib/authFunnelCrawlGuard.ts. Disallow: /*backUrl= # The anon-view gate beacon. Since 2026-08-04 its payload rides in the QUERY # STRING (a body-carrying POST was being truncated and answered 400 by nginx, # losing ~1,500 views/day — lib/anonViewBeaconPayload.ts), which made the URL # unique per company: ~382M of them. A body is invisible to a crawler, a query # string is not, and JS-rendering crawlers re-request the URLs a page fetches. # Measured: 0 GETs on this route on 08-03 and 0 before the 21:52 deploy on # 08-04, then 37 from bingbot carrying 35 distinct companyKey values in the # ~40min after it. Blocking this ONE route costs a crawler nothing — a # recognised crawler is already short-circuited out of metering before the # route does any work (lib/crawlerRequest.ts) — which is why it is safe to # disallow here and NOT safe to disallow /api/ wholesale, since Googlebot # applies robots.txt to subresources during rendering. # See lib/anonViewBeaconCrawl.ts. Disallow: /api/anon-view/ # The /compare company deep-link. Every company page's header renders # Compare, so the `c` parameter # gives /compare the cardinality of the company index — ~382M x 16 locales. # Measured 2026-08-05 over a 23.6h nginx window (7,154,530 requests): 131,799 # GET /*/compare, 121,752 of them carrying exactly ONE `c=` param, i.e. the # followable header link; the two deep-links that were already nofollowed emit # the multi-`c=` shape and account for the other 10,040. Of the 87,279 distinct # IPs that asked, 128 (0.15%) ever fetched a JS chunk — this is a synthetic # fleet, not visitors. None of it is indexable: /compare?c=… already # self-canonicalises to the bare /compare (lib/canonicalPath), so every one of # those ~0.57s SSR renders displaces a fetch of a real company page. # Two lines per prefix: robots.txt matches a literal prefix, so the bare form # and the locale-prefixed form must both be named. The parameterless # /{locale}/compare stays crawlable — it is the canonical target and costs # nothing. See lib/compareDeepLinkCrawl.ts. Disallow: /compare?c= Disallow: /*/compare?c= # The /contact-us removal deep-link. From 2026-08-09 the officer page's privacy # footer links to /contact-us?reason=1&personId={this person}, so that the # person whose record it is arrives with the removal reason chosen and their own # record already resolved instead of re-finding themselves in an index of 207M # people. The href is therefore unique per officer page x 16 locales, which is # the same cardinality trap as the four rules above — declared here BEFORE it # can be crawled rather than after. Baseline measured in the 09/Aug window, on # the build that shipped without it: 348 requests to /contact-us, 339 bare and # the 9 with a query carrying only utm_source / __jsfix. Nothing indexable is # lost: the parameterless /{locale}/contact-us is the canonical form, is the one # listed in sitemap-pages.xml, and stays crawlable — a query-string form of it # self-canonicalises there anyway (lib/canonicalPath). Two lines per prefix, # same reason as above. The bounded pre-existing deep-links (/export-data's # volume quote, the cancel popup) are covered by the same rules and lose # nothing, being reachable only by a human already on those pages. # See lib/contactUsPrefill.ts and components/OfficerV2/PrivacyFooter.tsx. Disallow: /contact-us? Disallow: /*/contact-us? # The faceted /search surface. Every option in the filter rail — jurisdiction, # status, legal form, year, contact axis, size band — plus every pagination # cell is an whose href is the CURRENT filter set plus one more axis, so a # fetched search URL hands out SEVENTY more in its own HTML (measured live) and # each of those hands out seventy more. The space is combinatorial: country x # types x statuses x years x activities x contacts x size x sort x page x q, in # 16 locales. Measured 2026-08-10 over one nginx day: 300,287 requests to # /{locale}/search — 4.9% of everything served — of which 300,259 carry a query # string, 269,661 carry a facet parameter, and 287,232 are DISTINCT URLs (95.6% # fetched exactly once). That is enumeration, not searching; the referers agree, # with 75 arrivals from Google web search against 231,008 with no referer. # Not one of these URLs can ever be indexed: every /search response ships # noindex,follow and canonicalises to the bare /{locale}/search, and no sitemap # entry anywhere carries a query string. Yet Googlebot — verified real, its 169 # addresses reverse-resolving into googlebot.com — spent 45,770 requests here, # 10.8% of its entire crawl of this site, second only to /company. rel="nofollow" # was already on every link inside the search page's own filter rail and # pagination and did not stop it, because nofollow has been a hint rather than a # directive to Google since 2019 — while the SEED links, on the company and # officer pages that hand the frontier its first URLs, carried no rel at all # (five files, fixed in the same change). robots.txt is what still binds # Googlebot, and /search was the one high-cardinality surface with no rule at # all. Two lines per prefix, same reason as the rules above. The # parameterless /{locale}/search stays crawlable — it is the canonical target, # it is what every variant already points at, and it took 28 of the 300,287 # requests. See lib/searchFacetCrawlGuard.ts. Disallow: /search? Disallow: /*/search? # The archive day page's keyset continuation cursor. A day with more than 900 # registrations (73,035 (country, day) rows across 59 jurisdictions, measured # 2026-08-11) used to end at page 30 with a bare 404 for anyone who tried to # read further; `?after=` continues from a known row instead, at # constant cost. Every continuation is `noindex,follow` and self-canonicalizes # to the bare day URL, and — the part that actually binds the fleet that forges # its user-agent — THE CURSOR IS NEVER IN SERVER-RENDERED HTML: it is written # in from an effect after mount, so a client that does not execute JS sees only # the ordinary archive links it already had. This line is kept for the # compliant minority, exactly like the five rules above it, and is not what # makes the surface safe. See lib/archiveContinuation.ts. Disallow: /*after= # The two self-healing recovery flags. When edge-cached HTML outlives the build # whose content-hashed assets it references, a missing stylesheet renders the # page unstyled and a missing script renders it inert; lib/staleStylesheetRecovery.ts # re-fetches the page under `?__cssfix=1` and lib/staleScriptRecovery.ts # re-navigates under `?__jsfix=1`. Both flags force a CDN MISS on purpose, so # every one of them is a full uncached SSR render at the origin. # # Measured 2026-08-11 in one partial nginx day: 14,128 `?__cssfix=1` and 22,671 # `?__jsfix=1` requests. 10,411 of the cssfix ones came from DECLARED crawlers # (9,422 GoogleOther), 316.7 MB of renders that bought those crawlers nothing — # a crawler does not paint and does not click, and the SSR HTML it came for is # already complete. The source of that is now fixed: BOTH modules skip crawlers # outright, so these URLs stop being generated. # # This rule is for the ones ALREADY discovered. Unlike the forged-user-agent # fleets the rules above cannot bind, the clients here are compliant declared # crawlers — GoogleOther is 9,422 of them — which is exactly the population a # Disallow does work on. Both flags self-canonicalize to the clean URL # (Head.Component strips the query), so nothing indexable is lost. Disallow: /*__cssfix= Disallow: /*__jsfix= # The three sitemaps, and what each one is FOR — they do not overlap: # # sitemap_index.xml.gz company pages only. 3,984 child sitemaps, all named # sitemap_{country}_{n}.xml.gz, 100% /company/ URLs # (verified 2026-08-07 by enumerating the whole index # and opening four children across four jurisdictions). # Generated by b2bhint-cron-scripts, not by this repo. # sitemap.xml a hand-maintained static file: the 15 locale # homepages. nginx aliases this URL straight to # public/sitemap.xml, so it can never be a Next route. # sitemap-pages.xml EVERYTHING ELSE, added 2026-08-07 because it was in # NO sitemap in ANY locale: /pricing, /about, /archive, # /compare, /export-data, /list-match, /openindex, # /api-reference, /data-sources (index + every # per-registry citation page), /contact-us, the legal # pages, every country landing page, and the six # cross-portfolio exposure searches. Built from the # app's own canonical slug builders so no listed URL is # one 301 away from its canonical form; noindex and # session-gated routes are excluded by construction. # See lib/staticSitemap.ts. Sitemap: https://b2bhint.com/sitemap_index.xml.gz Sitemap: https://b2bhint.com/sitemap.xml Sitemap: https://b2bhint.com/sitemap-pages.xml