Brandon Lazovic

02 · Pillar

How Machines Read the Web

Agents and crawlers read structure, not pixels. Semantic HTML, server-rendered schema, and embeddable item data decide what gets understood and surfaced.

Search Console counts your AI impressions and cannot connect one of them to a visitGoogle's generative AI report shipped worldwide with impressions and no clicks. Three separate changes hit AI citations in the same two weeks, and the report cannot price any of them.How Machines Read the WebI passed every agent-readiness check while three statements in my own metadata were falseA public scanner rated this site level 4 of 4 for AI agents while three of its machine-facing claims were false. Structural checks cannot see a wrong value.How Machines Read the WebFetch is not parse: why broken schema still returns a clean crawlGoogle now applies one pass of HTML unescaping to JSON-LD, and its crawlers do not parse JSON at all. A successful fetch says nothing about whether your data parsed.How Machines Read the WebGoogle's goto Redirect Breaks URL Precision While Rank Tracking SurvivesGoogle confirmed google.com/goto redirects are rolling out. Domain-level rank tracking survives the change, and URL-level attribution is what degrades.How Machines Read the WebMore than half of these ChatGPT answers never searched the web at allAcross 270 alternatives queries, ChatGPT searched the live web on 44%. On the rest nothing you publish this quarter can change the answer.How Machines Read the WebBlocking the AI crawlers did not stop the citations, and the bot table explains whyAround 75% of sites blocking OpenAI or Google AI bots were still cited. Their own table shows why: nobody blocked the retrieval bot, only the training bots.How Machines Read the WebCitation half-life: a refresh trigger worth having, and a number worth checkingA vendor now plots per-URL citation decay and says half of cited content is under 13 weeks old. The metric is useful. The figure has no published method.How Machines Read the WebCloudflare wrote a real AI agent checklist. Don't trust the score next to it.Cloudflare published a real, checkable AI-agent specification. A 15-retailer test shows why the score placed beside it should not be trusted.How Machines Read the WebMost AI-Driven Site Visits Carry No Referral Tag At AllA panel study finds AI mentions lift site visits 1.5 to 2.5x, but for ChatGPT, only about 2.5% of those visits carry a trackable referral tag.How Machines Read the WebGoogle's Sign-In Test Turns a CAPTCHA Wall Into an Identity WallGoogle is testing a sign-in wall on deep search results weeks after a DMCA loss to SerpApi, swapping an anonymous CAPTCHA for an identity check.How Machines Read the WebWhen an AI Agent Logs In, Your Analytics See a HumanThree AI agents now log into accounts and complete purchases, and this site's own crawler log proves no user-agent bucket can catch them.How Machines Read the WebWhy an AI watermark can convict your content but never clear itAnthropic now watermarks Claude's text. Google has watermarked Gemini's since May 2024 with no ranking signal to show for it. Here is what the mark proves.How Machines Read the WebGoogle Wants a 304. 11 of 13 Retail Homepages Send the Full Page Anyway.Google's crawl-budget guide now recommends serving 304 responses, but testing 13 live retail homepages found 11 still send the full page every time.How Machines Read the WebDisallow Stops the Crawl. It Doesn't Stop the Index.Google indexed Claude's blocked share pages via inbound links, then said noindex needs crawl access too. Same mechanism, opposite failure.How Machines Read the WebValidate Fix in Search Console: why it only speeds recrawling for batch fixesValidate Fix samples a few pages before recrawling the rest, which only pays off for batch issues affecting many URLs.How Machines Read the WebChatGPT's source selection lives in an editable prompt OpenAI can rewrite overnightChatGPT's source selection runs on an editable server-side policy that OpenAI can rewrite at will, and the underlying fields already changed twice in July 2026.How Machines Read the WebGoogle's billions-of-clicks claim: why it can't be checked against your own site dataGoogle says AI Search sends billions of clicks weekly and calls the number stable, but Cloudflare and Wikipedia both report reproducible traffic declines.How Machines Read the WebAnnotated Page Content: the structured tree Chrome hands Gemini instead of your HTMLChrome builds a structured, typed node tree from a page's render tree and hands that to Gemini when a tab is shared, per Dejan AI's teardown.How Machines Read the WebDoes AI prefer Reddit? Why the answer depends on which AI you askOpenAI rejects 99.4% of the Reddit pages it retrieves, Anthropic cites zero, and only Google keeps Reddit near its baseline citation rate.How Machines Read the WebShopify's hreflang toggle relocates the risk instead of fixing itShopify's new toggle turns off auto-generated hreflang tags, but it can't verify your replacement is correct or see CDN-level conflicts.How Machines Read the WebHow Brave Search finds pages, and why chasing every engine is a trapBrave has no Search Console, no sitemap ingestion, and no IndexNow. It rides on Googlebot-crawlability, and Claude's web search runs on Brave's index, not Bing's.How Machines Read the WebYour 'are you a bot' screen can deindex your real page, and you'll never see it happenGoogle's John Mueller confirmed bot-challenge screens served to Googlebot with a 200 status get indexed, and dedup can hand your canonical to a stranger's copy.How Machines Read the WebGoogle's 'nearest trusted seed' ranking patent expires this year. The leak shows the idea did not.A 2006 Google patent ranks pages by distance from a few trusted seed sites. It expires in 2026, but the leak shows the idea survived as PageRank-NearestSeeds.How Machines Read the WebThe EU AI Act just made "AI-generated" a required, machine-readable label. Here is who has to use it.The EU AI Act's Article 50 makes AI-content disclosure enforceable on August 2, 2026. Here is what gets marked, who has to disclose, and what it means for search.How Machines Read the WebWhat RAG Actually Does (and When You Shouldn't Use It)RAG lets an AI look answers up instead of guessing from training. Here is how it works, what you can build with it, and when a simpler tool wins.How Machines Read the WebGoogle's new AI-slop defense is a video-platform paper, not a Search reveal, and that's the useful partGoogle's S-CTS paper targets coordinated AI-slop channels on video platforms, not Search rankings. The real lesson is coordinated footprints, not your pages.How Machines Read the WebAnthropic's Economic Index is a demand signal, not a jobs report, and the unit of demand is now the artifactAnthropic's June 2026 Economic Index shows the unit of demand shifting from keyword to artifact. Read as a demand signal, it maps which intents AI is absorbing.How Machines Read the WebSonnet 5 brings near-Opus agents at a fraction of the Opus price, and that is a deadline for machine-readable commerceClaude Sonnet 5 matches or comes within two points of Opus 4.8 on most agentic benchmarks, at a fraction of the price. For commerce, the story is the cost curve.How Machines Read the WebProduction agents read the accessibility tree firstProduction browser agents like Playwright MCP read the accessibility tree before pixels. For e-commerce, clean semantic HTML and schema decide machine readability.How Machines Read the WebSoft tokens: how single embeddings change item representationGoogle is testing 'soft tokens', single embeddings that replace multi-token item descriptions, for cheaper representation with cold-start implications.How Machines Read the WebGoogle's crawler split: why SSR vs. CSR decides Merchant Center complianceThe Shopping crawler does not render JavaScript; it validates price and availability. SSR vs. CSR is the real determinant of Merchant Center compliance.How Machines Read the Web