Read any public web page as clean, LLM-ready Markdown plus metadata (title, description, author, dates, canonical). Respects robots.txt.