Kapiway Research · August 2026

The AI search land grab has already started. We have the receipts.

On August 7, 2026 we ran 50 of the biggest SaaS marketing sites through Kapiway's own audit engine: the same crawl-based checks our free AI Readiness scan runs on any site. We expected to find neglect. We found the opposite: the leaders are quietly positioning for AI search, and the real gap is not where most people think.

31 of 50
ship a real llms.txt on their main domain
4 of 50
restrict AI crawlers in robots.txt
29 of 50
score weak (50 or less) on structured data
37 of 50
have images missing alt text

Finding 1: llms.txt went mainstream while nobody was looking

62% of the sites we scanned (31 of 50) serve a real llms.txt file on their main domain: a plain-text index that tells AI crawlers what the site is and where its important pages live. Not docs subdomains, not auto-generated stubs. Real, curated files on the primary marketing domain, at Stripe, Slack, Salesforce, Shopify, HubSpot, Notion, Webflow, Zapier and 23 more. Slack's file literally opens with "Welcome, humans and bots alike."

For calibration: when we checked 20 well known SEO tools four days earlier, only 35% had one on their main domain. The SaaS leaders have out-adopted the SEO industry itself.

Finding 2: almost nobody locks AI out, but four sites are curating it

Only 4 of the 48 sites whose robots.txt we could read restrict any AI crawler (two sites, gusto.com and mercury.com, refuse automated fetches entirely). Loom blocks GPTBot outright. Calendly blocks CCBot everywhere except two pages it clearly wants read: /llms.txt and /pricing. Canva curates five training crawlers with allow exceptions and fully blocks Bytespider. And Figma has written the most deliberate AI policy in the study: the training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) are blocked entirely, while the answer engines' live agents (OAI-SearchBot, CHATGPT-User, PerplexityBot, Claude-SearchBot) are granted access through roughly 8,700 allow rules covering hand-picked pages: home, pricing, release notes and localized versions. That is not opting out of AI search. That is curating exactly what AI may read: feed the answers, starve the training sets. The remaining 44 readable sites simply leave the door open.

Finding 3: the real moat is structured data, and most sites are weak there

Here is the gap nobody talks about. 58% of these sites (29 of 50) score 50 or below on structured data: the schema markup that helps machines actually understand what a page is, who it is for, and what it offers. The file that says "AI, please read me" is everywhere. The markup that makes the reading useful is not. Add the hygiene gaps (37 of 50 have images missing alt text, 45 of 50 have title tag length problems) and the pattern is clear: adoption of the visible signal is ahead of the fundamentals.

Every site, every number

Scores are from Kapiway's crawl-based audit (up to 6 pages per site, public pages only, Aug 7 2026). llms.txt = real text file served at /llms.txt on the main domain. Blocked = AI crawlers disallowed sitewide in robots.txt. Rerun any site through the free scan to check the current state.

Sitellms.txtBlocks AI botsOn-pageStructured dataTechnical
15five.comnono75100100
activecampaign.comnono7550100
airtable.comnono755094
asana.comyesno645089
basecamp.comnono62100100
box.comnono6250100
brex.comyesno100100100
buffer.comnono73100100
calendly.comyesCCBot7958100
canva.comnoGPTBot, ClaudeBot, CCBot, anthropic-ai, Applebot-Extended, Bytespider505033
clickup.comyesno6258100
deel.comyesno625067
docusign.comnono6250100
dropbox.comyesno75100100
figma.comnoGPTBot, ClaudeBot, Google-Extended, CCBot75100100
framer.comyesno100100100
freshworks.comyesno505033
ghost.orgnono799294
grammarly.comnono8850100
gusto.comnono505033
hootsuite.comyesno8850100
hubspot.comyesno8850100
intercom.comyesno10050100
kit.comnono7550100
klaviyo.comyesno7550100
lattice.comyesno7550100
linear.appyesno7150100
loom.comyesGPTBot88100100
mailchimp.comyesno83100100
make.comyesno505033
mercury.comnono505033
miro.comnono717589
monday.comyesno425072
notion.soyesno8850100
pipedrive.comnono505033
ramp.comnono8610094
rippling.comyesno88100100
salesforce.comyesno75100100
shopify.comyesno75100100
slack.comyesno7550100
sproutsocial.comyesno87100100
squarespace.comyesno75100100
stripe.comyesno5250100
surveymonkey.comyesno100100100
typeform.comnono75100100
webflow.comyesno755089
wix.comyesno88100100
zapier.comyesno71100100
zendesk.comyesno8850100
zoom.usnono7758100

Method, honestly

  • Sites: 50 widely known B2B SaaS marketing domains, chosen before scanning, none excluded afterward.
  • Engine: Kapiway's production crawl audit (sitemap discovery, up to 6 public pages per site), plus direct checks of /llms.txt and /robots.txt. No logins, no private data.
  • llms.txt counts only if the main domain serves a real text file, not an HTML error page. We spot-verified positives by hand.
  • Scores are category averages of pass rates on standard checks (titles, meta, headings, canonical, schema presence, alt text, HTTPS, status codes and more).
  • Robots analysis covers 48 of 50 sites; two (gusto.com, mercury.com) block automated fetches of robots.txt itself.
  • Point-in-time snapshot: sites change. Run any of them through the free scan for today's answer.

Where does your site stand?

The same engine that produced these numbers grades any site free, in about 2 minutes.

Get your free score