llms.txt and AI Crawlers: What ERP Consulting Firms Should Allow

For thirty years, one small text file has decided who gets to read your website. The robots.txt convention is older than Google itself, and for most of its life the visitors it governed were search crawlers whose bargain was obvious: let us read you, and we will send you traffic.

Now there are new visitors knocking. The crawlers that feed AI models and answer engines have arrived, and the bargain is less obvious. Which is why so many ERP consulting firms have made their decision about them entirely by accident:

  • A security plugin shipped with AI bots blocked by default.
  • A CDN toggled a setting nobody was told about.
  • An article about content scraping scared someone into a blanket ban in 2023, and nobody has looked since.

Months later, a controller in Texas asks ChatGPT which NetSuite partners actually understand process manufacturing, and the firm that wrote the definitive guide on exactly that topic is nowhere in the answer. We keep finding the same solution to this mystery in audit engagements: the front door was locked, by nobody, on purpose.

This post is the deliberate version of that decision. Who the crawlers are, what the tradeoff actually is for an ERP services firm, what we recommend allowing, and how llms.txt, the newest file in the conversation, fits in.

The short answer

llms.txt is a proposed convention for giving AI systems a curated, plain-text map of your website: what your firm is, stated in one paragraph, and where your most important pages live. It complements robots.txt, which controls what crawlers may access, and sitemap.xml, which lists your URLs for discovery. For ERP consulting firms, our standing recommendation is simple: allow AI crawlers on all public marketing content, keep genuinely private client material blocked at every layer, publish a short llms.txt because it costs an hour and doubles as your positioning record, and write the whole policy down so no plugin update can quietly reverse it.

The rest of this post is the reasoning, the mechanics, and the 30-minute implementation.

Three files, three jobs: robots.txt vs sitemap.xml vs llms.txt

Before the access question, get the vocabulary straight. These three files sit at your site root and do different work, and conflating them is where most confusion in partner conversations starts.

File What it does Who reads it Status
robots.txt Tells crawlers which paths they may and may not access Every reputable crawler, search and AI alike Established convention since 1994, formalized as an internet standard in 2022
sitemap.xml Lists your URLs so crawlers can discover everything you publish Search engines primarily, some AI crawlers Established standard
llms.txt A curated, annotated map written for language models: who you are and which pages matter most Emerging and uneven; some AI tools read it, others do not yet Proposed convention, introduced in late 2024

One sentence each, for the record:

  • robots.txt is the lock. It decides access.
  • sitemap.xml is the inventory. It decides discoverability.
  • llms.txt is the introduction. It decides how a machine that already got in understands your firm.

Most of this post is about the lock, because that is where ERP firms are quietly hurting themselves. The introduction comes after, because it only matters once the door is open.

Who is actually crawling your ERP firm’s website?

The AI crawler population sorts into three types with different jobs, and the distinction matters because you can treat them differently in robots.txt. Each operator documents its bots publicly, including OpenAI, Anthropic, and Perplexity.

1. Training crawlers

These collect text that may be used to train future models. Allowing them influences what tomorrow’s models know in their bones, the memory layer we described in how AI assistants build their ERP recommendations. When a model “just knows” that a certain firm specializes in NetSuite rescues for distributors, training data is usually why.

2. Retrieval and search-index crawlers

These build and refresh the indexes that AI assistants search at answer time, so they can cite live sources. Allowing them influences whether you can be cited today. This is the citation pipeline, and it includes the classic search crawlers whose indexes now feed AI surfaces: Bingbot behind Copilot, documented at Bing Webmaster Tools, and Googlebot behind Google’s AI answers, per Google Search Central.

3. User-triggered fetchers

These grab one specific page because a human asked an assistant to read it. Picture a CFO pasting your NetSuite implementation cost guide into ChatGPT and asking for a summary. Blocking these mostly frustrates your own prospects mid-research.

Here is the current roster that matters, sorted by job:

User agent Operator Job What allowing it means for an ERP firm
GPTBot OpenAI Training Your expertise informs future OpenAI models
ClaudeBot Anthropic Training Your expertise informs future Claude models
Google-Extended Google Training control Governs AI training use of Google’s crawl, separate from Search
CCBot Common Crawl Training (open dataset) Your content enters the public dataset many models have historically used
OAI-SearchBot OpenAI Retrieval index Eligible for citation in ChatGPT search answers
Claude-SearchBot Anthropic Retrieval index Eligible for citation in Claude’s web answers
PerplexityBot Perplexity Retrieval index Eligible as a source in Perplexity answers
Bingbot / Googlebot Microsoft / Google Classic search, feeds AI surfaces Visibility in Copilot and Google AI answers, plus ordinary search
ChatGPT-User, Claude-User, Perplexity-User OpenAI / Anthropic / Perplexity User-triggered fetch A prospect can hand your page to their assistant and it works

Two honest caveats before you laminate that table.

The cast changes. Bots are added and renamed continually, and the trade press, Search Engine Land among others, tracks the roster better than any static list can. Govern by category and review quarterly rather than memorizing names.

Compliance varies. Reputable operators document their bots and honor robots.txt. The disreputable fringe ignores it. That is an argument for infrastructure-level controls where it matters, not an argument that the polite signals are pointless.

What does blocking AI crawlers actually cost a consulting firm?

For publishers who monetize pageviews, the AI bargain is genuinely contested: models ingest their content and answer readers who then never visit, which is why news sites block training bots in droves. Their content is their product. Reasonable position.

Your situation is structurally different, and this is the paragraph that should drive your policy:

A consulting firm’s website content is not the product. It is the advertisement. You publish NetSuite implementation guides and manufacturing case studies precisely so the right people learn that you exist, what you do, and why you can be trusted. A machine that reads all of it and then tells a Texas manufacturer “consider this firm, they specialize in exactly your situation” has not stolen your content. It has completed your content’s mission, with distribution you could not have bought.

Run the ledger both ways:

Cost of allowing Cost of blocking
Content reuse Public marketing pages may inform models and answers without a click None
Control You surrender some theoretical control over how public pages are reused You keep control over pages fewer machines ever read
Future models None Absence from the memory of tomorrow’s models
Citations today None Ineligible as a source when assistants answer live questions
Shortlists None Invisible in the channel where, as we covered in GEO vs SEO, ERP shortlists increasingly form

For a firm whose entire growth motion is being known and trusted by a small market, the blocking column is not a rounding error. It is the business. Selection committees at mid-market companies are already asking assistants to draft their vendor longlists, a shift the whole discipline of GEO for ERP consulting firms exists to address. Absence from the index is absence from the longlist.

The access policy we recommend for ERP partners

Given that ledger, our standing recommendation has held up across every client situation so far. Four rules:

  1. Allow, by default, for all public marketing content. Blog, service pages, case studies, About, team pages, the whole persuasive surface, open to training and retrieval crawlers alike. This content exists to be found and repeated. Let it be.
  2. Protect what was never public anyway. Client portals, gated resources, staging and development environments, internal search results, and anything under NDA stay disallowed, for AI bots and everyone else. This is not AI policy. It is the same site hygiene an ERP consulting website should already have.
  3. Decide the edge cases deliberately. Proprietary frameworks or paid research sit in a gray zone. Some firms gate them, accepting invisibility for exclusivity. Others publish them precisely to be cited as the source. Either is defensible. What is not defensible is never having had the conversation.
  4. Write the decision down. One paragraph in your marketing documentation: what is open, what is closed, who decided, when it gets reviewed. The whole failure mode this post exists to prevent is policy-by-plugin-default, and documentation is the vaccine.

Applied to the content types on a typical partner site:

Content type Examples Policy
Persuasive surface Blog posts, service pages, case studies, industry pages, About Open to all crawlers
Decision content Cost guides, comparison pages, methodology overviews Open; this is what assistants quote
Genuinely private Client portal, project workspaces, staging, NDA material Blocked for everyone, at the infrastructure level where possible
Gray zone Proprietary frameworks, paid research, premium templates Deliberate choice, documented either way

The mechanics: robots.txt without tears

Implementation is a fifteen-minute job wearing a scary reputation.

Here is the quiet secret of the open policy: allowing is the default. A crawler you never mention in robots.txt is allowed. So the recommended posture is mostly about deleting hostile lines, not adding welcoming ones. A clean file for the policy above looks like this:

# Everyone, search and AI alike: public marketing surface is open,

# private paths are not.

User-agent: *

Disallow: /client-portal/

Disallow: /staging/

Disallow: /wp-admin/

 

Sitemap: https://www.yourfirm.com/sitemap_index.xml

If you prefer the hedge some firms choose, blocking training while staying citable, the two categories use different user agents, so the middle position is easy to express:

# Opt out of model training only

User-agent: GPTBot

Disallow: /

 

User-agent: ClaudeBot

Disallow: /

 

User-agent: Google-Extended

Disallow: /

 

User-agent: CCBot

Disallow: /

 

# Everyone else, including retrieval and search crawlers

User-agent: *

Disallow: /client-portal/

Disallow: /staging/

One technical gotcha worth knowing: a bot that gets its own named section ignores the general section entirely. If a plugin generated a specific block for GPTBot, your private-path rules must be repeated inside that block, or GPTBot never sees them. This is the kind of detail that makes reading the actual file, not the plugin’s summary screen, worth five minutes.

And that is where the surprises live, because robots.txt is only one of three layers that can bounce a crawler:

Layer What to check Typical failure
robots.txt file The raw file at yoursite.com/robots.txt Old blanket disallows aimed at AI user agents, left over from a 2023 scare
Plugins Security and SEO plugin settings “Block AI bots” shipped as a checkbox someone once ticked and forgot
CDN / firewall Bot management rules Providers like Cloudflare offer one-click AI crawler blocking, and have moved toward blocking by default for new domains, overriding anything your file politely says

The file can welcome a bot the firewall is bouncing at the door. Verify at every layer, then confirm reality in your server logs: the crawlers you allowed should actually appear fetching pages within days.

Now, llms.txt: what it is and is not

Alongside the access question sits the newer idea. llms.txt is a convention proposed by the team at Answer.AI and documented at llmstxt.org: a markdown file at your site root giving language models a curated map of your site. What your firm is, in plain words, and where the pages that matter live. Optionally accompanied by a fuller llms-full.txt containing the content itself in clean text.

The honest status report, which we owe you because most coverage skips it:

  • It is an emerging convention, not a standard.
  • Adoption by the major model providers is uneven and shifting, and some have said they do not read the file yet.
  • Nobody should promise you that publishing one changes your AI visibility next week.

So why do we still recommend it? Three reasons that survive the hype discount:

  1. It is nearly free. An hour of work for a small firm, less if your positioning is already sharp.
  2. Its downside is zero and its option value is real. If consumption grows, early publishers inherit the benefit without lifting a finger again.
  3. Writing it is a forcing function. An llms.txt is your canonical entity record and your site’s greatest hits, stated plainly, which is exactly the exercise the entity SEO sprint demands anyway. Firms that struggle to write their llms.txt have discovered a positioning problem, not a file-format problem.

A good one for an ERP consultancy runs a page or two. The shape:

# Yourfirm Consulting

 

> NetSuite Solution Provider serving mid-market manufacturers and

> distributors across the US Midwest. 60+ implementations since 2014.

> Fixed-fee rescue engagements for stalled projects a specialty.

 

## Services

– [NetSuite Implementation](https://…/netsuite-implementation/): Scoping to go-live for manufacturing and distribution

– [NetSuite Rescue](https://…/netsuite-rescue/): Recovery for stalled or failed projects, fixed fee

 

## Proof

– [Case Study: 9-week WMS rollout](https://…/case-studies/wms/): Three warehouses, one distribution client

– [Case Study: Process manufacturing go-live](https://…/case-studies/process-mfg/): Batch costing and compliance

 

## Guides

– [NetSuite Implementation Cost Guide](https://…/cost-guide/): Real ranges by company size and complexity

The one-sentence description, a short paragraph of context, then curated links with one-line annotations to your core service pages, flagship case studies, and best guides. Not the sitemap; the tour.

Reading your logs: five-minute literacy

Once the policy is live, the server log is your truth serum, and reading it takes less skill than it looks. Filter requests by user agent and you will see the population directly: the classic search crawlers on their steady rounds, the training bots sweeping periodically in bursts, the retrieval bots arriving in small, query-shaped visits, often hitting exactly the pages that answer a question someone just asked. That last pattern is worth savoring: it is the closest thing GEO has to watching demand happen, a theme we pull apart in how Perplexity picks its sources.

What the patterns mean:

What you see What it means What to do
Allowed bots fetching marketing pages within days of the change Healthy; the policy is live end to end Nothing; diary the next quarterly review
Silence from a bot you thought you welcomed A layer above the file, plugin, firewall, or CDN rule, is still bouncing it Re-check all three layers, then the logs again
Bots fetching disallowed private paths Either a misconfigured rule or a non-compliant crawler Fix the rule, or move that path behind authentication where it belonged anyway
A flood of requests from agents you have never heard of The disreputable fringe Rate limiting and firewall rules, a server hygiene matter rather than robots etiquette

Questions ERP partners actually ask

Does allowing AI crawlers hurt our Google rankings? No. No mechanism connects the two. The AI training controls are separate from search crawling, and Google-Extended specifically exists so you can govern AI training use without touching Googlebot. Blocking Googlebot itself would hurt everything, which is exactly why the controls were separated.

Can we allow retrieval but not training? Yes, and it is a coherent middle position: block the training agents, allow the search and retrieval agents, and you stay citable today while opting out of tomorrow’s model memory. We think full openness is the better trade for a services firm, but this hedge is defensible and easy to implement, since the two categories use different user agents. The second code block above is that posture.

What about our images and downloadable assets? Same logic, applied per asset. Marketing imagery and public one-pagers gain from being readable. Anything you would not hand a stranger at a conference stays behind the gate it should already have.

Does llms.txt replace robots.txt or our sitemap? No. Three files, three jobs: robots.txt decides access, sitemap.xml decides discoverability, llms.txt explains you to machines that already got in. Publish all three.

We are a twelve-person NetSuite partner with a forty-page site. Is this worth our time? Especially you. Small partner sites are the ones most likely to be running on plugin defaults nobody has reviewed, and the whole exercise is thirty minutes. A firm your size lives on being known by a small market, which makes the blocking column of the ledger proportionally more expensive, not less.

How do we find out what our current policy is? Read yoursite.com/robots.txt in a browser, open your security and SEO plugin settings, check your CDN’s bot rules, and then look at thirty days of server logs filtered by user agent. Or have it done for you: the access check is a standard early page of our AI visibility audit.

What about crawlers that ignore robots.txt entirely? Handle them at the infrastructure layer with rate limits and firewall rules. Their existence is not an argument against the polite signals; the reputable operators, who run the assistants your buyers actually use, honor the file, and they are the ones your policy is for.

The 30-minute setup, end to end

Minutes Task Output
0 to 10 Read your current robots.txt, plugin settings, and CDN bot rules A written note of what the accidental policy has been
10 to 20 Implement the deliberate policy at whichever layer was blocking: open marketing surface, protected private paths A robots.txt and rule set that match the decision table above
20 to 30 Draft the llms.txt from your canonical entity record, publish both files, diary the log check for next week and the policy review for next quarter Two files live at the site root, two calendar entries

Fold the quarterly review into the standing audit cadence and it stops being a project at all. This is the cheapest infrastructure work in the entire GEO series, and for firms that were accidentally locked down, the single highest-leverage half hour available, because every other investment, the citation magnets, the schema layer, the content itself, was being built behind a closed door.

The bottom line

The AI crawler question feels like a technology decision and is actually a strategy declaration: is your public content an asset you rent out for pageviews, or an ambassador you send to wherever your buyers ask their questions?

Publishers can reasonably answer the first way. An ERP consulting firm that answers the first way has misidentified its own business model. Open the marketing surface, guard the genuinely private, write the policy down, and let the machines read the story you spent years building. The alternative is letting them tell your market a story with you missing from it.

Not sure what your accidental policy has been? Get in touch and we will run the access check. It has yet to come back boring.

ABOUT THE AUTHOR

Zees Zeeshan

Founder of IgnitX · SEO & Growth Strategist for ERP Consulting Firms

Zees has spent years in the ERP world working with NetSuite, SAP, Dynamics, Acumatica, Odoo, and many other partners, and founded IgnitX to help consulting firms win the quiet research phase, when ERP deals are actually decided.

 

Share This Post :

Table of Contents