In Nepal, “AEO” means three different things
On 12 August 2026 I pulled the Nepali Google results for “aeo nepal”. Nine organic results came back carrying three different meanings of the same acronym.
- Six were answer engine optimization, sold by agencies and trainers.
- Second place was New Business Age on Nepal and China aligning over Authorised Economic Operator recognition — a customs designation for trusted traders.
- Seventh place offered “SEO/AEO”, glossed on its own page as App Engine Optimization.
Say the full words out loud when you brief anyone.
Three labels circulate for the answer engine version — AEO, GEO and LLMO — and I have covered the labels themselves in AI and the future of SEO in Nepal. This page is the work underneath them: making sure a machine can fetch your pages, and that when it does it finds statements worth repeating.
The order matters more than the vocabulary. Most Nepali guides start with content and schema, which is the wrong end: content advice does nothing if the machine cannot reach the page, and some Nepali sites are blocking the crawlers their owners most want.
Step one: check whether the machines can fetch you at all
Thirty seconds, no tools. Type your domain into a browser with /robots.txt on the end:
yourbusiness.com.np/robots.txt
Look for lines beginning User-agent: followed by the name of an AI system, and whether the line beneath says Allow or Disallow. Three outcomes are common:
- Nothing about AI at all. Usually a WordPress default. This is fine — a wildcard Allow: / covers every crawler you have not named, AI systems included.
- A list of AI names with Disallow beneath them. Someone opted you out, and it may not have been you.
- A wall of comment text about “content signals”. That is Cloudflare, and it is the case below.
One caution: robots.txt is a request, not a wall. As Cloudflare puts it, “robots.txt compliance is voluntary” and “does not prevent crawlers from accessing your content at a technical level.” The operators named here do honour it. Not everyone does.
The Cloudflare switch that blocks crawlers you never decided to block
Cloudflare is common in front of Nepali sites, usually for the free SSL and caching. It has a dashboard setting called Set your preference to block training in robots.txt. Switch it on and, in Cloudflare’s words, it “will prepend our managed robots.txt before your existing robots.txt, combining both into a single response.” Available on all plans, the free one included.
The managed block disallows Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, CloudflareBrowserRenderingCrawler, Google-Extended, GPTBot and meta-externalagent, and sets a content signal of search=yes, ai-train=no, use=reference.
Two live Nepali examples, both fetched on 12 August 2026. Nepal Airlines serves that exact block at nepalairlines.com.np/robots.txt: every traveller asking an assistant about domestic flights in Nepal is querying systems told not to train on the flag carrier’s own site.
Bytecode Developers, a Kathmandu agency, sells AEO work: section nine of its SEO services page is headed “AEO and GEO Optimization”, under the title “SEO Services in Nepal | Get Found Online on Google & AI”. Its robots.txt serves the same managed block. I read that as the trap rather than as dishonesty — the switch lives in a security dashboard, not in any SEO tool, and nobody thought to look. If it can happen to an agency writing on this topic, it can happen to you.
Now read the list again, because this is the part almost every commentator gets wrong. The block names GPTBot and ClaudeBot. It does not name OAI-SearchBot, Claude-SearchBot or PerplexityBot — the retrieval crawlers. A site behind that block is therefore still eligible to appear in ChatGPT search, Claude’s search results and Perplexity. What it has given up is the training crawlers and Google-Extended.
That is a trade-off, not a catastrophe. If you object to your archive training models, keep it; you are not sacrificing your place in AI answers by doing so. Just decide it deliberately, rather than finding out later that a security setting made the decision for you.
Three kinds of bot, and only one kind decides whether you get cited
Every major operator runs separate crawlers for separate jobs, each controlled independently. Block the wrong one, it costs nothing. Block the right one, you are invisible.
| User-agent | Operator | What it governs | Cost of disallowing it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Appearing in ChatGPT’s search features | High. Opted-out sites “will not be shown in ChatGPT search answers” |
| GPTBot | OpenAI | Content that may train foundation models | None for search visibility. A training decision only |
| ChatGPT-User | OpenAI | Fetches when a user asks a question | None for search. Being user-initiated, “robots.txt rules may not apply” |
| Claude-SearchBot | Anthropic | Indexing content for search quality | High. “May reduce your site’s visibility and accuracy in user search results” |
| ClaudeBot | Anthropic | Collecting web content for model training | None for search visibility. Signals exclusion from training datasets |
| Claude-User | Anthropic | Fetches when a Claude user asks a question | Moderate. “May reduce your site’s visibility for user-directed web search” |
| PerplexityBot | Perplexity | Appearing and being linked in Perplexity results | High. It is “not used to crawl content for AI foundation models” — purely visibility |
| Perplexity-User | Perplexity | Fetches in response to a user question | Little. It “generally ignores robots.txt rules” |
| Google-Extended | Gemini training, and grounding in Gemini Apps and Vertex AI | Moderate. See below — it does not touch Google Search | |
| Googlebot | “Google Search (including Discover and all Google Search features)” | Total. Everything, AI features in Search included |
Quotations come from the crawler documentation of OpenAI, Anthropic, Perplexity and Google. Read the top two rows together, because they overturn the most repeated claim in Nepali AI-search content. Disallowing GPTBot does not remove you from ChatGPT. OpenAI states the settings are independent: a webmaster “can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot”. That is a supported configuration, not a contradiction.
The practical minimum is short: allow OAI-SearchBot, Claude-SearchBot and PerplexityBot. Those three are your visibility. GPTBot, ClaudeBot, CCBot and Google-Extended are a separate question about training, and you can answer it however you like.
What Google-Extended does not do
Google-Extended gets described in Nepali guides as the switch that gets you into Google’s AI answers. It is not. Google’s crawler documentation says the opposite in one sentence:
Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.
It is a standalone product token governing whether your content trains future Gemini models and whether it can ground answers in Gemini Apps and Vertex AI. It has no user-agent string of its own; existing Google agents do the crawling.
If your goal is to appear in Google’s results, AI features included, the crawler deciding that is Googlebot, exactly as it was five years ago. Which is the unglamorous point underneath all of this: the work that gets you into AI answers is mostly the work that got you into search results. If pages are slow, blocked or duplicated across two URLs, no amount of AEO vocabulary fixes it — these are the faults I list in why a Nepali website is not ranking.
How to test your AI visibility so the answer means something
Most “proof” of AI visibility in Nepal is a screenshot of a chatbot naming a client — no prompt, no date, no way to repeat it. These systems are non-deterministic; ask twice and you may get two answers.
Worse, the obvious test gives false positives. Paste your own URL into ChatGPT and the fetch is made by ChatGPT-User, and because a user asked for it, robots.txt rules may not apply. Perplexity is explicit that Perplexity-User “generally ignores robots.txt rules” for the same reason. So the chatbot reads a site its own search crawler is forbidden to index, and you conclude you are visible when you are not. A method you can defend:
- Fix your prompts in writing first. Ten to fifteen, phrased as a customer would: “best trekking company in Kathmandu for Annapurna Base Camp”, not “is [your brand] good”. Never paste your own URL.
- Run them logged out, so you are not shown your own history.
- Log every run in a sheet. Date, system, exact prompt, whether you were named, which sources were cited. That last column is the valuable one: it tells you which Nepali pages the system trusts on your topic.
- Repeat monthly, same prompts. One run tells you nothing; a trend across three months tells you something.
- Check your server logs. The objective half. If OAI-SearchBot, PerplexityBot and Claude-SearchBot never appear there, nothing above will improve until that changes.
Give changes time. OpenAI notes it can take around 24 hours from a robots.txt update for its systems to adjust, and Perplexity gives the same figure. Editing the file this morning and testing at lunchtime proves nothing.
One warning that matters in Nepal, where cheap shared hosting and blunt firewall rules are common: robots.txt is not the only thing that can block a crawler. Perplexity’s documentation devotes a section to web application firewalls and publishes IP ranges so you can write allow rules. If your logs show nothing while robots.txt says Allow, look at the firewall next.
What actually earns a Nepali business a citation
Once the machines can reach you, the question is what they find. In my experience these systems quote specific, checkable, attributable statements and mostly skip marketing language. “Leading provider of world-class solutions” is unquotable; there is no fact in it. What is quotable tends to be dull and precise:
- Prices in NPR, as numerals on the page. Not “affordable packages”, not a PDF, not an enquiry form. A range in text, with what changes it.
- A full address including ward number. Kathmandu Metropolitan City has 32 wards, so “Lainchaur-26, Kathmandu” is a location and “Kathmandu, Nepal” is not. Ward numbers are how Nepali addresses work, and they separate you from every other business on the road.
- One postcode, used everywhere. Genuinely awkward here. Nepal Post’s published postal code table lists Kathmandu Metropolitan City as 30608, with per-ward codes 3060801 to 3060832. The code most Nepali businesses and online forms write is 44600, which does not appear in that table at all. Pick one and use only that, identically, across your site, Google Business Profile and every directory listing. Inconsistency stops a machine treating them as one business.
- Registration and licence numbers. Company registration, PAN or VAT, a Nepal Tourism Board or trekking licence number if you hold one. Among the strongest trust signals a Nepali business can publish, and almost nobody does.
- Dates and named authorship. Published and last-updated, plus a real biography for whoever wrote it.
The inverse holds too. Round unverifiable numbers — “500+ businesses served” — are a bad trade, because a machine will repeat them to a prospect who then asks you to prove them. If a claim cannot survive that question, do not publish it.
llms.txt and schema: useful, but not the first move
The most common technical recommendation in Nepali AEO content is to add an llms.txt file, presented as the step that puts you ahead. I publish one, so I am not against it — but the priority is inverted. llms.txt is a proposed convention, not a standard adopted by the major operators, and I am not aware of any published commitment from OpenAI, Google, Anthropic or Perplexity to read it as an input to search visibility. It costs an hour and may pay off later. It does nothing on a site whose robots.txt blocks the crawler that would fetch it.
Structured data is the stronger investment, because it is genuinely standardised and does double duty in ordinary search: Organization or LocalBusiness with address and phone in machine-readable fields, Article with author and dates, FAQPage for questions you really answer, Product or Service with real prices. Mark up what is visible and nothing else.
The running order:
- Crawler access — robots.txt and firewall. Without it nothing else registers.
- Technical health — indexing, speed, one canonical URL per page.
- Specific, checkable facts written into the visible content.
- Structured data describing those facts.
- llms.txt, last, as a cheap bet on a convention that may or may not matter.
Steps one and two are ordinary technical SEO. Not a disappointing conclusion but a useful one: the work is well understood, and you can check whether it was done.
What this site does, described accurately
It would be poor form to write this without stating my own configuration, and you can check it by adding /robots.txt to this domain. This site uses a wildcard rule: User-agent: * then Allow: /. There are no per-agent Allow blocks, because none are needed — the wildcard permits every crawler I have not named, OAI-SearchBot, Claude-SearchBot, PerplexityBot, GPTBot, ClaudeBot and Google-Extended included. A comment records that this is deliberate, but a comment is documentation for humans, not a directive. Two agents are named and disallowed, Bytespider and Amazonbot — my judgement about crawlers that take content without sending traffic back, and you may disagree. I publish an llms.txt too, and have no evidence any system reads it.
Two more Nepali files read well as templates. aviralacharya.com.np groups GPTBot, ChatGPT-User, Google-Extended, PerplexityBot and others under one Allow, commented “Explicitly welcoming AI for training and real-time citation” — though it names no retrieval-only crawler. rejishkhanal.com.np does name OAI-SearchBot and Claude-SearchBot, while keeping /admin/ and /private/ closed. Both were deliberate, which puts them ahead of most.
Frequently asked questions
The short version
Open your robots.txt. If a Cloudflare block or an old plugin is disallowing the crawlers, nothing else here can help you. Fix that, confirm OAI-SearchBot, Claude-SearchBot and PerplexityBot are permitted, and check your firewall is not blocking them behind your back. Decide separately whether you want the training crawlers.
After that the work is familiar: pages that load and index, prices in NPR as numerals, an address with a ward number, one postcode used consistently, dates, named authors, structured data describing facts actually on the page. Then test with a fixed set of prompts, logged and dated, month after month — never by pasting your own URL into a chatbot, since that fetch bypasses the rules you are measuring.
Six of the nine Nepali results I pulled for “aeo nepal” were selling AEO training or AEO services. The market for the vocabulary is well ahead of the market for the work, and the thing that decides whether a machine can quote your business is a text file most owners have never opened.