TechnicalRobots.txt, sitemap.xml and llms.txt set up right
Act as a technical SEO engineer. Set up robots.txt, sitemap.xml and llms.txt for {{website_url}}, built with {{tech_stack}}. Our most important pages: {{key_pages}}. First, read the current files and tell me what's wrong, including any leftover staging rule that blocks the whole site.
robots.txt:
- Allow everything by default; disallow only low-value crawl paths such as internal search results, cart, checkout and API routes
- Don't use it to hide private pages, since anyone can read it; protect those with sign-in and a noindex header
- AI crawlers: list the main ones, including each company's separate crawlers for model training and for answering searches, then ask me which to allow. Don't decide for me
- End with the sitemap's absolute URL
- Staging and preview builds stay out of search through a password or noindex header, never a copied robots file
sitemap.xml:
- Generated from routes and the database at build time or on request, never written by hand
- Only canonical URLs that return 200 and are indexable: no redirects, noindex pages or duplicates
- lastmod only when the content really changed; skip priority and changefreq, which search engines ignore
- A sitemap index past 50,000 URLs, with separate sitemaps for posts, products or locations
llms.txt:
- A Markdown file at the root: an H1 with our name, a one-paragraph summary, then sections linking key pages with one line each on what they cover
- Honest and current, generated from the same page data where possible
- Tell me plainly that it's a proposed convention, and that search engines don't use it for ranking
Finish with all three files, how to submit the sitemap in Google Search Console and Bing Webmaster Tools, and what to check after deploying.