# robots.txt — mindlynk.ca # Format: RFC 9309 (https://www.rfc-editor.org/rfc/rfc9309) # Content-Signal: https://contentsignals.org/ # # The policy in one sentence: we want to be found, read and cited by search # engines and AI assistants; we do not offer our writing as model training data. # # To change that stance, edit two things together so they stay consistent: # 1. the Content-Signal line in the default group below, and # 2. the "training crawlers" groups near the bottom. # --------------------------------------------------------------------------- # Default policy — every crawler not named further down # --------------------------------------------------------------------------- User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /_old-site/ # --------------------------------------------------------------------------- # Search and answer engines — allowed in full. # Being read and cited here is the entire point of the site. # --------------------------------------------------------------------------- User-agent: Googlebot Allow: / Disallow: /_old-site/ User-agent: Bingbot Allow: / Disallow: /_old-site/ User-agent: DuckDuckBot Allow: / Disallow: /_old-site/ # OpenAI's search index (distinct from GPTBot, which is training) User-agent: OAI-SearchBot Allow: / Disallow: /_old-site/ # OpenAI, fetching a page because a user asked for it in a conversation User-agent: ChatGPT-User Allow: / Disallow: /_old-site/ # Anthropic, fetching a page because a user asked for it User-agent: Claude-User Allow: / Disallow: /_old-site/ # Anthropic's search index User-agent: Claude-SearchBot Allow: / Disallow: /_old-site/ User-agent: PerplexityBot Allow: / Disallow: /_old-site/ # --------------------------------------------------------------------------- # Model-training crawlers — declined, to match "ai-train=no" above. # These are the groups to flip if the position on training ever changes. # --------------------------------------------------------------------------- # OpenAI's training crawler User-agent: GPTBot Disallow: / # Anthropic's general crawler User-agent: ClaudeBot Disallow: / # Legacy Anthropic agent name, kept for older deployments User-agent: Claude-Web Disallow: / # Controls Gemini and Vertex AI grounding and training. # Note the trade-off: this does NOT remove the site from Google Search or from # AI Overviews, which follow Googlebot above — but it does keep these pages out # of Gemini's grounding set. Allow it if Gemini visibility matters more than the # training position. User-agent: Google-Extended Disallow: / # Apple Intelligence training User-agent: Applebot-Extended Disallow: / # Meta's AI crawler User-agent: meta-externalagent Disallow: / # Common Crawl — a public corpus widely used as training data User-agent: CCBot Disallow: / # ByteDance User-agent: Bytespider Disallow: / # Amazon User-agent: Amazonbot Disallow: / # --------------------------------------------------------------------------- Sitemap: https://mindlynk.ca/sitemap.xml # Agent capability manifest (ARD) — see also on pages Agentmap: https://mindlynk.ca/.well-known/ai-catalog.json