House Notes

Cloudflare's new CLI or plain curl for your coding agent

We gave Claude the same 12 Cloudflare jobs, twice with each tool, 48 runs in all. With curl and the API it got 24 right in a median of 2 turns; with cf it got 20, in 5. Fact (2026-09-28): Cloudflare’s changelog says “More than 2,900 commands cover the public Cloudflare API.” Fits: curl when an agent does the work. Try first: curl with a read-only API token.

The claim

If your agent already reaches Cloudflare through its API, keep it there. On the jobs cf can do, both were right every time, but cf took more than twice the turns and about twice the cost. On the two jobs cf has no command for, the agent spent 11 to 13 turns searching and then gave up.

cf is a beta that ships fast: 13 betas in its first 7 days. The numbers below are for version 1.0.0-beta.12.

What we measured

Every run was a fresh claude -p session on claude-opus-5-5, with no skills, memory or project instructions loaded. Each got the same on-call scenario for a made-up company on example.com and one task, such as listing the DNS records in the example.com zone or rolling a Worker back to its previous version. The only difference was one line: use curl with this token, or use cf, already logged in. Requests went to a local stand-in for the Cloudflare API, never a real account.

curl and the API cf
Correct, all 12 jobs 24 of 24 20 of 24
Correct, the 10 jobs cf has a command for 20 of 20 20 of 20
Median turns, those 10 jobs 2 5
Median time, those 10 jobs about 8 s about 18 s
Cost, those 10 jobs (20 runs) $1.21 $2.35
Workers routes and custom domains 4 of 4 0 of 4
Changes nobody asked for 0 0

The extra turns went to finding commands. Across its 24 runs the cf agent called --help 76 times and cf cli search 30 times. The curl agent called neither and went straight to the API paths.

cf cli search returns five commands for any query, with no “no match” answer. For “list the Workers routes”, which cf cannot do, it still offered five confident commands, so the agent kept trying.

Download the per-run data and the exact prompts:

  • results.csv: one row per run, with correctness, turns, cost, time and tool calls.
  • tasks.md: the prompt, both tool lines, all 12 tasks and how each was graded.

Check what your agent can already reach

The first draft of this test assumed the agent had no Cloudflare access, because no token sat in its environment and no CLI was logged in. That was wrong. Our tokens live in a secrets manager, and the agent fetches the one it needs for each task. It had been minting, rolling and revoking Cloudflare tokens over the API since early August.

So before you add a new CLI to give an agent access, look at where its credentials actually come from. An empty environment is often on purpose.

Give it a read-only token

What we were missing was not access but a safe way to look. Every token our agent could reach could also change things. We added one that can only read. In the Cloudflare dashboard, create an API token with these permission groups:

  • On all zones: Zone Read, Zone Settings Read, DNS Read, Zone WAF Read, Workers Routes Read.
  • On all accounts: Workers Scripts Read, Workers R2 Storage Read.

Settings such as the SSL mode sit under Zone Settings Read, a separate group from Zone Read, and custom WAF rules sit under Zone WAF Read. With this token, our agent read DNS records, the SSL mode, WAF rules, Workers routes and R2 buckets, and every attempt to change something was refused. Workers R2 Storage Read can also download the objects in your buckets, so leave it out if those hold anything private.

Store the token in your secrets manager before you use it, because Cloudflare shows the value once. Then hand your agent only that one token for the task, not your whole secrets store.

A read-only token protects you only if the agent cannot also reach a write token. If your agent can read every secret in the store, it can still pick up one that writes.

If you use cf anyway

Two behaviors to know before you hand it to an agent:

  • In a run with no terminal attached, which is how an agent runs it, commands that change things, such as cf zones settings edit and cf dns records create, go through with no prompt.
  • cf dns records delete without --force prints Aborted. and exits 0, so a script cannot tell “declined” from “done” by the exit code.

A read-only token is the only real guard.

What this doesn’t show

The API was a local stand-in with fixed answers, not a live account. When a run made a change, the stand-in’s read-back did not change, and four runs then said the job could not be done, two with each tool. All four had already sent the right request and named the key fact, so they count as correct under the grading rule in tasks.md; no exception was made.

Each job ran twice per tool, on one model, on one day. The jobs are on-call lookups and fixes we needed ourselves, not a random sample of Cloudflare work. They come from a 15-job test of cf’s command search run the same day, and three of those 15 were not run here: listing R2 buckets, listing Turnstile widgets and verifying a token. cf’s search put the right command first for all three, so they were likely easy jobs for cf, and leaving them out may understate it. For jobs cf has a command for, compare the rows for those 10 jobs.

The cf runs had no API token, so they could not fall back to curl. We did not test an agent given both, or a person using cf by hand. We did not test cf for building or deploying Workers; Cloudflare’s docs say Astro 6 and later does not build with cf during the beta.

cf changes fast. A later version that adds Workers routes and custom domains, or a search that can say “no match”, could change the result.

Reader poll

Self-selected, unverified responses. These counts describe this poll's responses.

JavaScript is needed to send an answer and load this poll's counts.

Enter the House