Pulse’s semantic HTML check reported “no h1, no h2, no main” on zu2b.com despite the tags being visibly present in view-source. Testing revealed three sequential, independently plausible bugs stacked on top of each other. Each “fix” successfully resolved one layer and revealed the next one underneath.
⚡ Quick Fix (TL;DR)
The Culprit: Bug 1: Minified HTML without spaces before attributes (`<h1class=`) slipped through the regex pattern. Bug 2: The fetcher never decompressed gzip responses, so raw binary noise was being pattern-matched. Bug 3: The body accumulation was capped at 150KB, and the real h1 sat at byte 467,365, past a massive wall of Autoptimize-inlined CSS in the head.
The Fix: All three bugs were fixed: (1) regex updated to catch any valid tag-terminator character; (2) zlib-based decompression added before text processing; (3) body accumulation cap raised to 500KB. Broader implication: this likely affected every regex-based check on any gzip-compressed site (Cloudflare, Sucuri, etc.), not an edge case.
Tagged in :