Brandon Lazovic

02 · Pillar

How Machines Read the Web

Agents and crawlers read structure, not pixels. Semantic HTML, server-rendered schema, and embeddable item data decide what gets understood and surfaced.

ChatGPT's source selection lives in an editable prompt OpenAI can rewrite overnightChatGPT's source selection runs on an editable server-side policy that OpenAI can rewrite at will, and the underlying fields already changed twice in July 2026.How Machines Read the WebGoogle's billions-of-clicks claim: why it can't be checked against your own site dataGoogle says AI Search sends billions of clicks weekly and calls the number stable, but Cloudflare and Wikipedia both report reproducible traffic declines.How Machines Read the WebAnnotated Page Content: the structured tree Chrome hands Gemini instead of your HTMLChrome builds a structured, typed node tree from a page's render tree and hands that to Gemini when a tab is shared, per Dejan AI's teardown.How Machines Read the WebDoes AI prefer Reddit? Why the answer depends on which AI you askOpenAI rejects 99.4% of the Reddit pages it retrieves, Anthropic cites zero, and only Google keeps Reddit near its baseline citation rate.How Machines Read the WebShopify's hreflang toggle relocates the risk instead of fixing itShopify's new toggle turns off auto-generated hreflang tags, but it can't verify your replacement is correct or see CDN-level conflicts.How Machines Read the WebHow Brave Search finds pages, and why chasing every engine is a trapBrave has no Search Console, no sitemap ingestion, and no IndexNow. It rides on Googlebot-crawlability, and Claude's web search runs on Brave's index, not Bing's.How Machines Read the WebYour 'are you a bot' screen can deindex your real page, and you'll never see it happenGoogle's John Mueller confirmed bot-challenge screens served to Googlebot with a 200 status get indexed, and dedup can hand your canonical to a stranger's copy.How Machines Read the WebGoogle's 'nearest trusted seed' ranking patent expires this year. The leak shows the idea did not.A 2006 Google patent ranks pages by distance from a few trusted seed sites. It expires in 2026, but the leak shows the idea survived as PageRank-NearestSeeds.How Machines Read the WebThe EU AI Act just made "AI-generated" a required, machine-readable label. Here is who has to use it.The EU AI Act's Article 50 makes AI-content disclosure enforceable on August 2, 2026. Here is what gets marked, who has to disclose, and what it means for search.How Machines Read the WebWhat RAG Actually Does (and When You Shouldn't Use It)RAG lets an AI look answers up instead of guessing from training. Here is how it works, what you can build with it, and when a simpler tool wins.How Machines Read the WebGoogle's new AI-slop defense is a video-platform paper, not a Search reveal, and that's the useful partGoogle's S-CTS paper targets coordinated AI-slop channels on video platforms, not Search rankings. The real lesson is coordinated footprints, not your pages.How Machines Read the WebAnthropic's Economic Index is a demand signal, not a jobs report, and the unit of demand is now the artifactAnthropic's June 2026 Economic Index shows the unit of demand shifting from keyword to artifact. Read as a demand signal, it maps which intents AI is absorbing.How Machines Read the WebSonnet 5 brings near-Opus agents at a fraction of the Opus price, and that is a deadline for machine-readable commerceClaude Sonnet 5 matches or comes within two points of Opus 4.8 on most agentic benchmarks, at a fraction of the price. For commerce, the story is the cost curve.How Machines Read the WebProduction agents read the accessibility tree firstOpenAI Atlas, Playwright MCP, and Perplexity Comet read the accessibility tree before pixels. For e-commerce, clean semantic HTML and schema matter more.How Machines Read the WebSoft tokens: how single embeddings change item representationGoogle is testing 'soft tokens', single embeddings that replace multi-token item descriptions, for cheaper representation with cold-start implications.How Machines Read the WebGoogle's crawler split: why SSR vs. CSR decides Merchant Center complianceThe Shopping crawler does not render JavaScript; it validates price and availability. SSR vs. CSR is the real determinant of Merchant Center compliance.How Machines Read the Web