Guides

Check Agent Readiness

Monitor the things AI agents look for on your site: robots.txt rules, llms.txt, a Markdown version of the page, and discovery files. What each probe checks, how the score works, alerting, and the OpenTelemetry event.

View as Markdown

Before an AI agent renders your page, and often instead of rendering it, it looks for a handful of machine-readable things: what robots.txt lets it do, whether there is an llms.txt, whether the page is available as Markdown, and whether the site publishes any discovery files.

These are easy to get right once and easy to break later. A deploy replaces llms.txt with the app's HTML shell. A docs restructure leaves its links pointing at nothing. The Markdown version of the pricing page keeps last quarter's prices. None of it shows up in an uptime check.

Agent readiness is an option on HTTP monitors. On the runs you choose, the monitor probes its own site for these things, reports each one, and gives the run a score. You decide which probes, if any, should fail the run.

One-off scanners will give you a score today. This gives you the same checks on a schedule, from every location the monitor runs in, with history, alerting and events in your own OpenTelemetry backend.

What gets checked

Every probe reports one of four outcomes: pass, fail (present but wrong), not served (absent, which is often a choice rather than a fault), or not checked (could not be evaluated on this run).

Access

ProbeIDPasses when
robots.txtrobots_txt/robots.txt is served as a text type, is not an HTML page, and can be parsed. The result lists which AI agents it disallows for the monitored page, and any Content-Signal directives.
robots.txt matches realityrobots_consistencyNo agent that robots.txt allows was blocked or challenged on this run. Needs agent personas on the same monitor.
SitemapsitemapThe sitemap named in robots.txt, or /sitemap.xml, is served as XML. A sitemap on another host is not fetched and is reported as not checked.

Content for agents

ProbeIDPasses when
llms.txtllms_txt/llms.txt is served as a text type, has a title line, and the links on this site that were checked on this run resolve. Up to 10 links are checked per run; with more than that, each run checks 10 starting from a random point in the list, so the whole file is covered over successive runs.
Markdown on requestmarkdown_negotiationThe monitored page returns Markdown when requested with Accept: text/markdown. Checked for GET monitors only.
Markdown file alongside the pagemarkdown_siblingA .md version sits next to the page: /pricing.md for /pricing, /index.md for /.
Markdown agrees with the pagemarkdown_parityEvery price and percentage in the Markdown version also appears on the page.

Discovery files

ProbeIDPasses when
MCP server cardmcp_server_card/.well-known/mcp/server-card.json is served as JSON.
API catalogapi_catalog/.well-known/api-catalog (RFC 9727) is served as JSON.
OAuth server metadataoauth_discovery/.well-known/oauth-authorization-server (RFC 8414) is served as JSON.
Agent skills indexagent_skills/.well-known/agent-skills/index.json is served as JSON.

The discovery files come from specifications that are still emerging. Most sites do not serve them yet, and "not served" here is not a defect. They are included so you can see when one appears or disappears.

Three things worth knowing about how it judges

  • A 200 is not enough. A single-page app answers every unknown path with its HTML shell and a 200. A file that comes back as an HTML page is treated as wrong (robots.txt, llms.txt) or not served (everything else), never as present.
  • robots.txt is read the way a crawler reads it. An agent's own User-agent group applies if there is one, otherwise the * group. The longest matching rule wins, and Allow wins a tie.
  • Markdown parity compares figures, not prose. It looks for prices ($, £, €) and percentages in the Markdown and checks that each appears in the page's visible text. A Markdown version that says $29 when the page says $39 fails. So does one that lists prices the page does not show at all, for example because the page shows a visitor's local currency. The details list every figure it could not find so you can judge which it is.

The headline finding: robots.txt says yes, the site says no

With agent personas on the same monitor, the robots_consistency probe compares what robots.txt promises with what each persona actually received:

robots.txt matches reality   Fail
1 agent allowed by robots.txt was turned away
ChatGPT-User: allowed by robots.txt, but got HTTP 403

This is the case where a WAF or bot rule disagrees with your published policy. The reverse (an agent robots.txt disallows that was served anyway) is listed as a note and does not fail the probe: robots.txt is a request, and serving a client you asked not to come is not an outage.

For the two to line up on the same run, give personas and readiness the same How often setting.

The score

The score is the percentage of evaluated probes that passed. Probes that were not checked are left out.

A site with a valid robots.txt and sitemap and nothing else scores around 20. That is a description, not a grade: whether you want an llms.txt or a Markdown version is your call. The score is most useful as a line over time, where a drop means something that used to work stopped.

Switch it on (Web UI)

  1. Open an HTTP monitor and click Edit.
  2. Switch on Agent readiness.
  3. Choose How often: every run, or 1 in N runs. The default is 1 in 12.
  4. Under Must pass, tick any probes that should fail the run when they do not pass.
  5. Save. The next run includes the probes, unless they already ran from that location in the last 5 minutes.

Switch it on (Monitoring as Code)

monitors:
  - name: Pricing page
    type: http
    url: https://www.example.com/pricing
    frequency: 5m
    agentReadiness:
      every: 12
      require:
        - robots_txt
        - llms_txt
        - markdown_parity

Writing the block switches readiness on. To pause it and keep the settings, add enabled: false. The fields are in the configuration reference.

As with agent personas, deploy with an up-to-date CLI everywhere, including CI. A CLI release from before this feature ignores the block and removes it from the monitor on its next deploy.

Which page to point it at

The site-level probes (robots.txt, llms.txt, sitemap, discovery files) give the same answer for any page on a site, so one monitor per site covers them.

The Markdown probes are about the monitored page itself. Put readiness on the monitor for the page whose Markdown version matters most, typically pricing or the docs landing page.

Alerting

By default every probe is reported and the run stays successful.

A probe listed under Must pass (require) fails the run when it does not pass, with an error naming the probe and the reason, for example Agent readiness, llms.txt: 2 broken links in llms.txt; 8 of 14 links checked. Your existing alert rules then apply. See Set Up Alerts.

A required probe that could not be evaluated also fails the run. For robots_consistency that means personas must run on the same run.

Only require what you intend to keep. Requiring a discovery file you do not serve makes the monitor fail permanently.

Two requirements are rejected when you save, because the monitor could never meet them:

  • markdown_negotiation on a monitor whose method is not GET.
  • robots_consistency without agent personas enabled on the same monitor, or with a persona How often that does not line up. Personas must run on every run that readiness does, so readiness's every has to be a multiple of the personas' every (12 and 6 work; 12 and 5 do not).

Private locations need a current runner

The probes run in the runner. A private location running an image from before this feature ignores the setting: it reports ordinary results with no readiness section, and a required probe never fails there. Update the runner image at each private location before relying on Must pass for alerting. If runs from a location never show an Agent Readiness section, that location's runner is the first thing to check.

Read the results

Open a run that included readiness and scroll to Agent Readiness: the score, then each probe grouped as above with its outcome, a one-line summary, and details such as the broken links or the figures that differ. Probes you have marked as required are labelled Must pass.

Runs that were not sampled for readiness do not show the section.

OpenTelemetry

Each readiness run produces one synthetics.agent.readiness.checked log event in your backend, carrying the score and the lists of passed, failed and not-served probes. It also names any probe that regressed or improved since the previous readiness run from the same location, so "what changed" is one query:

event.name = "synthetics.agent.readiness.checked"
AND synthetics.agent.readiness.regressed_probes IS NOT EMPTY

Severity is WARN when any probe fails or regressed, otherwise INFO. The attribute list is in OpenTelemetry > Agent readiness signals.

Cost and limits

  • A readiness run counts as one additional HTTP run per location, however many probes it makes. A run that had no time to carry out any probe is not billed.
  • It makes up to 19 probe requests to the monitor's site, one after another: the files above plus up to 10 links from llms.txt. Each may follow one redirect on the same site.
  • It only requests the monitor's own site. Links in llms.txt and sitemaps that point at other hosts are counted and reported, not followed, and a redirect to another host is refused.
  • If the monitored URL itself redirects to another host, the probes are not run and say so. Point the monitor at the final URL.
  • Only the Accept: text/markdown request, which goes to the monitored URL itself, carries the monitor's headers and auth. Every other probe request is sent without them.
  • The probes share a 30 second budget per run, with 8 seconds per request. Anything that does not fit is reported as not checked.
  • At most one readiness run every 5 minutes per monitor per location, whatever the monitor's frequency and How often setting. A 10-second monitor set to "every run" still probes every 5 minutes.
  • These requests count toward the per-hostname rate limit.