---
title: 'Check Agent Readiness'
description: 'Monitor the things AI agents look for on your site: robots.txt rules, llms.txt, a Markdown version of the page, and discovery files. What each probe checks, how the score works, alerting, and the OpenTelemetry event.'
section: 'Guides'
canonical_url: 'https://yorkermonitoring.com/docs/guides/agent-readiness'
---

# Check Agent Readiness

Before an AI agent renders your page, and often instead of rendering it, it looks for a handful of machine-readable things: what `robots.txt` lets it do, whether there is an `llms.txt`, whether the page is available as Markdown, and whether the site publishes any discovery files.

These are easy to get right once and easy to break later. A deploy replaces `llms.txt` with the app's HTML shell. A docs restructure leaves its links pointing at nothing. The Markdown version of the pricing page keeps last quarter's prices. None of it shows up in an uptime check.

**Agent readiness** is an option on HTTP monitors. On the runs you choose, the monitor probes its own site for these things, reports each one, and gives the run a score. You decide which probes, if any, should fail the run.

One-off scanners will give you a score today. This gives you the same checks on a schedule, from every location the monitor runs in, with history, alerting and events in your own OpenTelemetry backend.

## What gets checked

Every probe reports one of four outcomes: **pass**, **fail** (present but wrong), **not served** (absent, which is often a choice rather than a fault), or **not checked** (could not be evaluated on this run).

### Access

| Probe | ID | Passes when |
|---|---|---|
| robots.txt | `robots_txt` | `/robots.txt` is served as a text type, is not an HTML page, and can be parsed. The result lists which AI agents it disallows for the monitored page, and any `Content-Signal` directives. |
| robots.txt matches reality | `robots_consistency` | No agent that `robots.txt` allows was blocked or challenged on this run. Needs [agent personas](/docs/guides/monitor-agent-access) on the same monitor. |
| Sitemap | `sitemap` | The sitemap named in `robots.txt`, or `/sitemap.xml`, is served as XML. A sitemap on another host is not fetched and is reported as not checked. |

### Content for agents

| Probe | ID | Passes when |
|---|---|---|
| llms.txt | `llms_txt` | `/llms.txt` is served as a text type, has a title line, and the links on this site that were checked on this run resolve. Up to 10 links are checked per run; with more than that, each run checks 10 starting from a random point in the list, so the whole file is covered over successive runs. |
| Markdown on request | `markdown_negotiation` | The monitored page returns Markdown when requested with `Accept: text/markdown`. Checked for `GET` monitors only. |
| Markdown file alongside the page | `markdown_sibling` | A `.md` version sits next to the page: `/pricing.md` for `/pricing`, `/index.md` for `/`. |
| Markdown agrees with the page | `markdown_parity` | Every price and percentage in the Markdown version also appears on the page. |

### Discovery files

| Probe | ID | Passes when |
|---|---|---|
| MCP server card | `mcp_server_card` | `/.well-known/mcp/server-card.json` is served as JSON. |
| API catalog | `api_catalog` | `/.well-known/api-catalog` (RFC 9727) is served as JSON. |
| OAuth server metadata | `oauth_discovery` | `/.well-known/oauth-authorization-server` (RFC 8414) is served as JSON. |
| Agent skills index | `agent_skills` | `/.well-known/agent-skills/index.json` is served as JSON. |

The discovery files come from specifications that are still emerging. Most sites do not serve them yet, and "not served" here is not a defect. They are included so you can see when one appears or disappears.

### Three things worth knowing about how it judges

- **A `200` is not enough.** A single-page app answers every unknown path with its HTML shell and a `200`. A file that comes back as an HTML page is treated as wrong (`robots.txt`, `llms.txt`) or not served (everything else), never as present.
- **robots.txt is read the way a crawler reads it.** An agent's own `User-agent` group applies if there is one, otherwise the `*` group. The longest matching rule wins, and `Allow` wins a tie.
- **Markdown parity compares figures, not prose.** It looks for prices (`$`, `£`, `€`) and percentages in the Markdown and checks that each appears in the page's visible text. A Markdown version that says `$29` when the page says `$39` fails. So does one that lists prices the page does not show at all, for example because the page shows a visitor's local currency. The details list every figure it could not find so you can judge which it is.

## The headline finding: robots.txt says yes, the site says no

With [agent personas](/docs/guides/monitor-agent-access) on the same monitor, the `robots_consistency` probe compares what `robots.txt` promises with what each persona actually received:

```
robots.txt matches reality   Fail
1 agent allowed by robots.txt was turned away
ChatGPT-User: allowed by robots.txt, but got HTTP 403
```

This is the case where a WAF or bot rule disagrees with your published policy. The reverse (an agent `robots.txt` disallows that was served anyway) is listed as a note and does not fail the probe: `robots.txt` is a request, and serving a client you asked not to come is not an outage.

For the two to line up on the same run, give personas and readiness the same **How often** setting.

## The score

The score is the percentage of evaluated probes that passed. Probes that were not checked are left out.

A site with a valid `robots.txt` and sitemap and nothing else scores around 20. That is a description, not a grade: whether you want an `llms.txt` or a Markdown version is your call. The score is most useful as a line over time, where a drop means something that used to work stopped.

## Switch it on (Web UI)

1. Open an HTTP monitor and click **Edit**.
2. Switch on **Agent readiness**.
3. Choose **How often**: every run, or 1 in N runs. The default is 1 in 12.
4. Under **Must pass**, tick any probes that should fail the run when they do not pass.
5. Save. The next run includes the probes, unless they already ran from that location in the last 5 minutes.

## Switch it on (Monitoring as Code)

```yaml
monitors:
  - name: Pricing page
    type: http
    url: https://www.example.com/pricing
    frequency: 5m
    agentReadiness:
      every: 12
      require:
        - robots_txt
        - llms_txt
        - markdown_parity
```

Writing the block switches readiness on. To pause it and keep the settings, add `enabled: false`. The fields are in the [configuration reference](/docs/reference/configuration#agent-readiness).

As with agent personas, deploy with an up-to-date CLI everywhere, including CI. A CLI release from before this feature ignores the block and removes it from the monitor on its next deploy.

## Which page to point it at

The site-level probes (`robots.txt`, `llms.txt`, sitemap, discovery files) give the same answer for any page on a site, so one monitor per site covers them.

The Markdown probes are about the monitored page itself. Put readiness on the monitor for the page whose Markdown version matters most, typically pricing or the docs landing page.

## Alerting

By default every probe is reported and the run stays successful.

A probe listed under **Must pass** (`require`) fails the run when it does not pass, with an error naming the probe and the reason, for example `Agent readiness, llms.txt: 2 broken links in llms.txt; 8 of 14 links checked`. Your existing alert rules then apply. See [Set Up Alerts](/docs/guides/set-up-alerts).

A required probe that could not be evaluated also fails the run. For `robots_consistency` that means personas must run on the same run.

Only require what you intend to keep. Requiring a discovery file you do not serve makes the monitor fail permanently.

Two requirements are rejected when you save, because the monitor could never meet them:

- `markdown_negotiation` on a monitor whose method is not `GET`.
- `robots_consistency` without agent personas enabled on the same monitor, or with a persona **How often** that does not line up. Personas must run on every run that readiness does, so readiness's `every` has to be a multiple of the personas' `every` (12 and 6 work; 12 and 5 do not).

### Private locations need a current runner

The probes run in the runner. A private location running an image from before this feature ignores the setting: it reports ordinary results with no readiness section, and a required probe never fails there. Update the runner image at each private location before relying on **Must pass** for alerting. If runs from a location never show an Agent Readiness section, that location's runner is the first thing to check.

## Read the results

Open a run that included readiness and scroll to **Agent Readiness**: the score, then each probe grouped as above with its outcome, a one-line summary, and details such as the broken links or the figures that differ. Probes you have marked as required are labelled **Must pass**.

Runs that were not sampled for readiness do not show the section.

## OpenTelemetry

Each readiness run produces one `synthetics.agent.readiness.checked` log event in your backend, carrying the score and the lists of passed, failed and not-served probes. It also names any probe that **regressed** or **improved** since the previous readiness run from the same location, so "what changed" is one query:

```
event.name = "synthetics.agent.readiness.checked"
AND synthetics.agent.readiness.regressed_probes IS NOT EMPTY
```

Severity is `WARN` when any probe fails or regressed, otherwise `INFO`. The attribute list is in [OpenTelemetry > Agent readiness signals](/docs/concepts/opentelemetry#agent-readiness-signals).

## Cost and limits

- A readiness run counts as **one additional HTTP run** per location, however many probes it makes. A run that had no time to carry out any probe is not billed.
- It makes up to 19 probe requests to the monitor's site, one after another: the files above plus up to 10 links from `llms.txt`. Each may follow one redirect on the same site.
- **It only requests the monitor's own site.** Links in `llms.txt` and sitemaps that point at other hosts are counted and reported, not followed, and a redirect to another host is refused.
- If the monitored URL itself redirects to another host, the probes are not run and say so. Point the monitor at the final URL.
- Only the `Accept: text/markdown` request, which goes to the monitored URL itself, carries the monitor's headers and auth. Every other probe request is sent without them.
- The probes share a 30 second budget per run, with 8 seconds per request. Anything that does not fit is reported as not checked.
- **At most one readiness run every 5 minutes** per monitor per location, whatever the monitor's frequency and **How often** setting. A 10-second monitor set to "every run" still probes every 5 minutes.
- These requests count toward the per-hostname rate limit.

## Related

- [Monitor AI Agent Access](/docs/guides/monitor-agent-access) for how the site responds to specific agents
- [Monitor MCP Servers](/docs/guides/monitor-mcp-servers) for agents that call your tools
- [Configuration reference](/docs/reference/configuration#agent-readiness)
