Navigation
Getting Started
Guides
Integrations
AI & Agents
Agentic Workflows
A worked example of setting up Yorker with a coding agent: prompts to use, the config each one produces, and how to keep the agent's changes reviewable.
This guide takes a service from no monitoring to HTTP, browser, and MCP monitors with alerting, using nothing but prompts to a coding agent. Each step shows the prompt, what the agent does, and the config you should expect to see.
The examples use Claude Code. Any agent that can edit files and run shell commands works the same way.
Agents are not deterministic. Your agent's output will differ in detail from the examples here. The safeguards in this guide (validate, test, diff, review) exist so that does not matter.
Before you start
You need a Yorker account and a terminal in the repo you want to monitor.
npm install -g @yorker/cli
yorker login
yorker skills installyorker login opens your browser and asks you to approve the CLI. This is the one step the agent cannot do for you. yorker skills install adds the Yorker agent skill to .claude/skills/yorker/; restart your agent so it picks the skill up.
It also helps to tell the agent a couple of facts once, in CLAUDE.md or AGENTS.md, so you do not repeat them in every prompt:
## Monitoring
- Production base URL: https://shop.example.com
- Monitoring is managed with Yorker in `yorker.config.yaml`. Never deploy without showing me `yorker diff` first.
- Alerts go to oncall@example.com.1. Create the first monitor
I want to monitor this service with Yorker. Set up the config and add a monitor that checks the homepage returns 200 every five minutes.
The agent scaffolds a config non-interactively:
yorker init --name shop --url https://shop.example.com --type http --frequency 5mwhich writes a yorker.config.yaml to the repo root. The scaffold names the monitor after the hostname, so the agent gives it a readable name:
project: "shop"
defaults:
frequency: "5m"
locations:
- loc_us_east
monitors:
- name: "Homepage"
type: http
url: "https://shop.example.com"
assertions:
- type: status_code
value: 200Before anything is deployed, the agent proves the config is sound and the monitor passes:
yorker validate # schema and cross-references, no network call
yorker test # sends each HTTP monitor's request from this machine
yorker diff # the plan: what would be created, updated, deletedYorker deploy plan for "shop"
Checks:
+ CREATE http "Homepage" (300s, 1 locations)
Summary: 1 to create, 0 to update, 0 to delete, 0 unchanged
Read the plan. If it is what you asked for, say so:
Looks right. Deploy it.
yorker deploy --wait--wait holds until the new monitor reports its first result from a hosted location. That first run is where assertions are evaluated, so the agent can tell you the monitor is live and passing rather than just submitted.
2. Add an API monitor with real assertions
A status code check tells you the server is up. It does not tell you the response is right. Ask the agent to read your code:
Add a monitor for
GET /api/products. Read the route handler and the response type, and assert on the fields a client actually depends on.
Because the agent has the repo, it can open the handler, see the shape of the response, and write assertions that mean something:
monitors:
- name: "Products API"
type: http
url: "https://shop.example.com/api/products"
frequency: "1m"
assertions:
- type: status_code
value: 200
- type: response_time
max: 2000
- type: header_value
header: "Content-Type"
operator: contains
value: "application/json"
- type: body_json_path
path: "$.products[0].id"
operator: exists
- type: body_json_path
path: "$.products[0].price.currency"
value: "USD"Ask the agent to check its assertions against a real response (a curl of the endpoint is enough) before it deploys. yorker test confirms the request reaches the endpoint and shows the status code, but assertions are only evaluated on hosted runs. yorker deploy --wait then exits non-zero if the new monitor fails its first run, so a wrong field name is caught within a minute of deploying, not by an alert later.
For an authenticated endpoint, the agent references a secret rather than writing the value:
secrets:
- SHOP_API_TOKEN
monitors:
- name: "Orders API"
type: http
url: "https://shop.example.com/api/orders"
auth:
type: bearer
token: "{{secrets.SHOP_API_TOKEN}}"Every secret a config uses must be named in the top-level secrets: list, and its value comes from the YORKER_SECRET_SHOP_API_TOKEN environment variable at deploy time. The token never has to appear in the repo or the conversation. See Secret Interpolation.
3. Turn a Playwright test into a browser monitor
If you already have end-to-end tests, the most valuable monitor is one you have mostly written.
Turn
tests/checkout.spec.tsinto a Yorker browser monitor. Run it every 10 minutes from the US and Europe. Add step markers so the filmstrip is readable.
Yorker runs vanilla Playwright and accepts a standard @playwright/test file as it is. The agent adds // @step: comments, which split the run into named steps with a screenshot and a timing each:
import { test, expect } from "@playwright/test";
test("checkout", async ({ page }) => {
// @step: Open shop
await page.goto("https://shop.example.com");
// @step: Add to cart
await page.getByRole("button", { name: "Add to cart" }).click();
// @step: Checkout
await page.getByRole("link", { name: "Checkout" }).click();
await expect(page.getByText("Order summary")).toBeVisible();
});and points a monitor at the file:
monitors:
- name: "Checkout Flow"
type: browser
script: "./tests/checkout.spec.ts"
frequency: "10m"
timeoutMs: 60000
locations:
- loc_us_east
- loc_eu_centralThe same file is now your end-to-end test in CI and your synthetic monitor in production. When the page changes, you fix one script.
Browser monitors run on hosted Chromium, not on your machine, so yorker test lists the steps but does not execute them. The agent runs the spec with your own Playwright first (npx playwright test tests/checkout.spec.ts), then uses yorker deploy --wait to confirm the first hosted run.
If you do not have a test yet, pair your agent with the Playwright MCP server so it can drive a real browser, look at the page, and write the script from what it sees:
Use the Playwright MCP tools to open https://shop.example.com, search for "socks", and add the first result to the cart. Then write that journey as a Yorker browser monitor with step markers.
4. Monitor an MCP server
If you ship an MCP server, your users are agents, and a 200 OK on the endpoint says very little about whether they can use it. An MCP monitor runs the real protocol session.
We expose an MCP server at
https://shop.example.com/mcp. Add a Yorker monitor that checks the handshake, confirms thesearch_productsandget_ordertools are advertised, and callssearch_productswith a real query.
secrets:
- SHOP_MCP_TOKEN
monitors:
- name: "Shop MCP"
type: mcp
endpoint: "https://shop.example.com/mcp"
frequency: "5m"
auth:
type: bearer
token: "{{secrets.SHOP_MCP_TOKEN}}"
expectedTools:
- search_products
- get_order
testCalls:
- toolName: search_products
arguments:
query: "socks"
expectedOutputContains: "socks"
detectSchemaDrift: true
rejectUnauthenticated: truedetectSchemaDrift records each tool's input schema, description, and annotations, and reports when any of them change between runs. rejectUnauthenticated also sends a request with no credentials on every run and fails unless the server refuses it. See Monitor MCP Servers.
5. Add alerting
Email oncall@example.com when any monitor fails twice in a row. For the Products API, also alert if response time goes over two seconds.
The agent defines the channel once, sets a default rule that every monitor inherits, and adds the extra rule where you asked for it:
alertChannels:
oncall-email:
type: email
addresses:
- oncall@example.com
defaults:
alerts:
- name: "down"
conditions:
- type: consecutive_failures
count: 2
channels:
- "@oncall-email"monitors:
- name: "Products API"
type: http
url: "https://shop.example.com/api/products"
alerts:
- name: "down"
conditions:
- type: consecutive_failures
count: 2
channels:
- "@oncall-email"
- name: "slow"
conditions:
- type: response_time_threshold
maxMs: 2000
channels:
- "@oncall-email"A monitor's own alerts replace the defaults rather than adding to them, which is why the agent repeats the down rule. This is the kind of detail the skill carries so you do not have to.
Slack, PagerDuty, ServiceNow, and webhooks work the same way. See Notification Channels.
6. Ship the monitor with the pull request
The strongest version of this workflow is one where the agent never deploys at all.
Open a PR with the new endpoint and its monitor. Add a GitHub Actions workflow that validates the Yorker config on every push, comments the diff on PRs, and deploys on merge to main.
The agent writes the workflow from the CI/CD guide. From then on:
- The monitor is reviewed in the same diff as the code it watches.
yorker diffin the PR shows reviewers exactly what will change in production monitoring.- The deploy runs in CI, with an API key held in your CI secrets. The agent's job ends at the pull request, so it never needs to run
yorker deployfrom a laptop.
7. Investigate a failure
The same CLI that creates monitors reads their results, so the agent that knows your code can also debug your production.
Are any monitors failing? If so, work out why.
A typical investigation looks like this:
yorker status --json # which monitors are unhealthy
yorker incidents list --json # correlated alerts, grouped
yorker results list "Checkout Flow" --status failure --since 24h --json
yorker results get "Checkout Flow" res_abc123 --json # failing step, error, console errors
yorker log --json # did someone change the monitor?
git log --since="2 days ago" --oneline # did someone change the code?The agent lines up when the failures started against what shipped, checks whether every location fails or only one, and reads the failing step and console errors from the browser run. A good answer looks like:
Checkout Flow has failed from all three locations since 14:12 UTC. The failing step is "Checkout": the "Checkout" link is no longer found. Commit
a41f9c2at 14:05 renamed that link to "Proceed to payment". The site works; the monitor's selector is stale. I can update the script and show you the diff.
For harder cases, Yorker can run its own analysis across recent results, baselines, and correlated signals:
yorker incidents analyze inc_abc123 --json
yorker monitors analyze "Checkout Flow" --jsonPrompts worth keeping
| Goal | Prompt |
|---|---|
| Cover a new endpoint | "Add a Yorker monitor for the endpoint I just wrote. Read the handler and assert on the response shape." |
| Review coverage | "Compare our route files with yorker.config.yaml and list the public endpoints that have no monitor." |
| Tighten a threshold | "Look at the last 7 days of results for Products API and suggest a response_time assertion that would not have fired on normal traffic." |
| Add an SLO | "Add a 99.9% availability SLO over 30 days for Checkout Flow, with burn-rate alerts to the on-call channel." |
| Adopt existing monitors | "Run yorker pull and commit the config so our monitors are managed as code from now on." |
| Triage | "Summarise active incidents and tell me which one to look at first and why." |
| Clean up | "Run yorker deploy --prune --dry-run and show me what it would delete. Do not deploy." |
Key takeaways
- Install the skill once and commit it. Every agent on the team then follows the same rules.
- Ask for outcomes, and ask the agent to read your code. Its advantage over a web form is that it knows what your endpoint returns.
- Validate, test, and diff are free. Insist on all three before a deploy, and use
--waitso the first hosted run is checked too. - Review the plan from
yorker diff, and let CI ownyorker deploywhen you can.