# Horamsa — crawl policy. # # Three jobs here: # 1. Let ordinary search engines index the public marketing, guide and legal pages, so # the site is findable. The signed-in app is disallowed — it is per-user, needs a # token, and has nothing a search result should ever show. # 2. Let AI ANSWER engines (ChatGPT search, Perplexity, Claude search) fetch those same # public pages, so they can cite and link to us when someone asks them about Vedic # astrology. These agents retrieve a page to answer a live question with a link back; # their operators state they are not used to train models. Each is named below with # the same app disallows as the wildcard group. # 3. Reserve our content against text-and-data mining in MACHINE-READABLE form. # terms.html section 6 forbids using our output to train AI, but a crawler never # agrees to Terms — and the EU DSM Directive Art. 4 permits TDM of public content # *unless* the rightsholder has reserved it in a way a machine can read. The TRAINING # crawlers below are that reservation. The two must say the same thing: if the Terms # clause is ever changed, change this file in the same commit. # # The split between (2) and (3) is per AGENT, not per company: OpenAI's OAI-SearchBot is # allowed and its GPTBot is not; Anthropic's Claude-SearchBot is allowed and its ClaudeBot # is not. Google's AI Overviews ride on Googlebot (group 1); Google-Extended, which only # governs training, stays reserved. # # A Disallow is a REQUEST. Well-behaved crawlers honour it; a determined scraper ignores # it. This is not a security control — the app's auth and rate limits are. Nor does # blocking Google-Extended or ClaudeBot below affect our own use of those companies' # APIs to generate readings (see privacy.html section 5): a training crawler and a # paid API are different doors. Sitemap: https://horamsa.com/sitemap.xml # ── Search engines: index the public pages, stay out of the app ─────────────── User-agent: * Disallow: /api/ Disallow: /auth/ Disallow: /admin Disallow: /chat Disallow: /readings Disallow: /charts Disallow: /reports Disallow: /workspace Disallow: /favourites Disallow: /refer Disallow: /support Disallow: /settings Disallow: /desk Disallow: /whiteboard Allow: / # ── AI answer engines: may fetch public pages to cite them, same app disallows ── # A named agent reads ONLY its own group, so each repeats the wildcard group's rules. User-agent: OAI-SearchBot Disallow: /api/ Disallow: /auth/ Disallow: /admin Disallow: /chat Disallow: /readings Disallow: /charts Disallow: /reports Disallow: /workspace Disallow: /favourites Disallow: /refer Disallow: /support Disallow: /settings Disallow: /desk Disallow: /whiteboard Allow: / User-agent: ChatGPT-User Disallow: /api/ Disallow: /auth/ Disallow: /admin Disallow: /chat Disallow: /readings Disallow: /charts Disallow: /reports Disallow: /workspace Disallow: /favourites Disallow: /refer Disallow: /support Disallow: /settings Disallow: /desk Disallow: /whiteboard Allow: / User-agent: PerplexityBot Disallow: /api/ Disallow: /auth/ Disallow: /admin Disallow: /chat Disallow: /readings Disallow: /charts Disallow: /reports Disallow: /workspace Disallow: /favourites Disallow: /refer Disallow: /support Disallow: /settings Disallow: /desk Disallow: /whiteboard Allow: / User-agent: Perplexity-User Disallow: /api/ Disallow: /auth/ Disallow: /admin Disallow: /chat Disallow: /readings Disallow: /charts Disallow: /reports Disallow: /workspace Disallow: /favourites Disallow: /refer Disallow: /support Disallow: /settings Disallow: /desk Disallow: /whiteboard Allow: / User-agent: Claude-SearchBot Disallow: /api/ Disallow: /auth/ Disallow: /admin Disallow: /chat Disallow: /readings Disallow: /charts Disallow: /reports Disallow: /workspace Disallow: /favourites Disallow: /refer Disallow: /support Disallow: /settings Disallow: /desk Disallow: /whiteboard Allow: / User-agent: Claude-User Disallow: /api/ Disallow: /auth/ Disallow: /admin Disallow: /chat Disallow: /readings Disallow: /charts Disallow: /reports Disallow: /workspace Disallow: /favourites Disallow: /refer Disallow: /support Disallow: /settings Disallow: /desk Disallow: /whiteboard Allow: / # ── AI training crawlers: reserved, no mining ───────────────────────────────── # One block per agent — the wildcard group above does NOT apply to a named agent, so a # bot listed here reads only its own block and would otherwise be unrestricted. User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: meta-externalagent Disallow: / User-agent: FacebookBot Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: cohere-ai Disallow: / User-agent: cohere-training-data-crawler Disallow: / User-agent: Diffbot Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Timpibot Disallow: / User-agent: AI2Bot Disallow: / User-agent: Kangaroo Bot Disallow: / User-agent: PanguBot Disallow: / User-agent: Scrapy Disallow: /