OPENLAWN FIELD NOTES SEPTEMBER 22, 2026

Can Opus 5.5
measure up to Astra?

We put vision models to work on the same residential aerials. Here is what finished, how long it took, and how closely the lawn outlines matched our reviewed example.

10 models60 model trials3 properties1 reviewed reference
PRIMARY TEST / MEDIUM REASONING
Claude Opus 5.5 · medium reasoning70.1%

mean outline overlap · 3 reference runs

Completed
5/5
Median time
23.4s
Mean reported cost
$0.121
GPT-6 Astra · medium reasoning77.3%

mean outline overlap · 3 reference runs

Completed
5/5
Median time
61.5s
Mean reported cost
$0.220

Outline overlap is intersection-over-union with one user-corrected map. It is not a surveyed accuracy percentage. Time and cost above summarize completed runs across the test set.

At medium reasoning, Opus 5.5 finished every trial with about 2.6× faster median time and 45% lower mean reported cost than Astra, but lower outline agreement on the reviewed home. Older Opus 5 and Astra scored similarly in this small sample. These results support a speed–quality tradeoff, not a universal model ranking.

01 / PRIMARY OUTLINE AGREEMENT

The same lawn.
Different interpretations.

Ranked by mean overlap across the completed repetitions on the reviewed property. All other properties contribute to speed, cost and completion metrics only.

Claude Opus 5medium reasoning · n=3
77.9%
GPT-6 Astramedium reasoning · n=3
77.3%
Claude Opus 5.5medium reasoning · n=3
70.1%
GPT-6 Astralow reasoning · n=3
67.6%
GPT-5.6 Terramedium reasoning · n=3
56.9%

Bars start at zero. The table below includes the observed range; these ranges are not confidence intervals.

02 / HIGH-REASONING FOLLOW-UP

Give both models
more room to think.

After the primary test, we requested high reasoning from Opus 5.5 and Astra on the same reviewed home. These runs use an 8,192-token response allowance and a 180-second call limit, compared with 2,400 tokens and 90 seconds in the primary comparison.

Claude Opus 5.5 · high reasoning73.4%

mean outline overlap · 3 reference runs

Completed
3/3
Median time
36.0s
Mean reported cost
$0.173
GPT-6 Astra · high reasoning78.9%

mean outline overlap · 1 reference run

Completed
1/2
Median time
174.8s
Mean reported cost
$0.410

Opus 5.5 averaged 73.4% overlap at high reasoning, compared with 70.1% at medium. Its observed ranges overlap. Astra’s only completed high run scored 78.9%, compared with a 77.3% medium mean. This small follow-up does not establish a reliable improvement from high reasoning.

One property only. Astra’s second high trial stopped at the aggregate budget guard after starting; its third was blocked before a model call. These are budget limits, not model completion failures. The Astra high score therefore represents just one completed run. Completion denominators exclude attempts blocked before any model call.

Compare settings on the same reviewed home

Model / reasoningScored completionsMean overlapMedian timeMean cost
GPT-6 Astramedium reasoning377.3%67.1s$0.201
Claude Opus 5.5medium reasoning370.1%23.3s$0.125
Claude Opus 5.5high reasoning373.4%36.0s$0.173
GPT-6 Astrahigh reasoning178.9%174.8s$0.410

High-reasoning speed and cost cover this home only; the primary summary covers three homes. Effort, output allowance and timeout changed together, so this does not isolate the effect of reasoning effort. All attempted trials remain in the export.

03 / THE FULL SCORECARD

Keep the failures
in the picture.

A fast successful run is only part of the story. Every attempted run remains in the downloadable data, including output limits, timeouts and unfinished tool calls.

Model / reasoningCompleteOverlap / rangeMedian timeFirst polygonMean cost
Claude Opus 5medium · 3 properties · anthropic/claude-opus-55/577.9%76.5–78.8 · n=333.6s4.8s$0.2865 known-cost completions
GPT-6 Astramedium · 3 properties · openai/gpt-6-astra5/577.3%68.8–82.9 · n=361.5s21.1s$0.2205 known-cost completions
Claude Opus 5.5medium · 3 properties · anthropic/claude-opus-5.55/570.1%63.0–78.7 · n=323.4s9.5s$0.1215 known-cost completions
GPT-6 Astralow · 3 properties · openai/gpt-6-astra5/567.6%55.8–82.9 · n=329.2s7.2s$0.1665 known-cost completions
GPT-5.6 Terramedium · 3 properties · openai/gpt-5.6-terra5/556.9%55.0–58.8 · n=339.5s15.0s$0.0485 known-cost completions
Claude Fable 5.1medium · 3 properties · anthropic/claude-fable-5.11/5—No scored completions41.6s29.1s$0.3161 known-cost completions
Gemini 3.1 Pro Previewmedium · 3 properties · google/gemini-3.1-pro-preview0/5—No scored completions——unknown0 known-cost completions
Gemini 3.8 Flashmedium · 3 properties · google/gemini-3.8-flash0/5—No scored completions——unknown0 known-cost completions
Grok 4.7medium · 3 properties · spacexai/grok-4.70/5—No scored completions——unknown0 known-cost completions
Fugu Ultra v2medium · 3 properties · sakana/fugu-ultra-v20/5—No scored completions——unknown0 known-cost completions
Kimi K3medium · 3 properties · moonshotai/kimi-k30/5—No scored completions——unknown0 known-cost completions
Claude Opus 5.5high · 1 property · anthropic/claude-opus-5.53/373.4%68.7–80.2 · n=336.0s15.1s$0.1733 known-cost completions
GPT-6 Astrahigh · 1 property · openai/gpt-6-astra1/21 stopped by budget after starting1 additional attempt blocked by budget78.9%78.9–78.9 · n=1174.8s98.4s$0.4101 known-cost completions

30/60 dispatched trials completed; 1 additional attempt was blocked before a model call. Costs are Gateway-reported model usage, not OpenLawn’s full measurement cost. 11 requests have unavailable billed usage. Our conservative accounting reserves those unknown charges: $21.76 against the $25 limit; $7.23 is confirmed reported usage, not a final billing total.

04 / HOW WE TESTED

Small sample.
Open methodology.

Production playground route at pinned source hashes; identical saved pinned aerial and whole-yard parcel for each residence; maintained extensions enabled; no account guidance; 2400 output tokens except existing Fugu 16384/32768; medium reasoning except explicit Astra-low profile; model-specific production tool choice. Three homes, with 3 repetitions on the sole user-corrected home and 1 on each other home. Reference withheld. At most 3 concurrent runs, order rotated. 180-second overall timeout. No automatic reruns of failed trials. Direct pipeline timing excludes geocoding, browser drawing, upload and saving. Saved model-input images are 2560 × 2560 pixels for every trial (the interactive playground currently downsamples to 1024 × 1024); these results are not end-to-end timings for the default playground image setting.

Separate follow-up on the single reviewed residence: planned 3 repetitions each for Astra and Opus 5.5, alternating sequentially. Requested high reasoning, 8192 output-token allowance, 180-second per-call timeout and 300-second overall timeout. Local route adapter permits high reasoning and extends the timeout; outgoing request adapter raises max_tokens. Prompt, tools and geometry reducer are unchanged. The same $25 aggregate guard includes original requests. This changes reasoning, output allowance and timeout together; it does not isolate the causal effect of reasoning alone.

  • One user-corrected outline, not surveyed ground truth. Repeats assess variation on that home, not generalization to new homes.
  • This is an OpenLawn workflow comparison. The prompt was developed around Astra. Shared reasoning labels do not guarantee equal compute.
  • Production model-specific settings differ: automatic tools for Opus 5.5 and Fable 5.1; larger output allowance for Fugu. Other models retain a 2400-token allowance. Incomplete output is a failure of this configuration, not proof that the model cannot measure lawns.
  • Models receive identical saved images, not newly captured imagery. Account-specific feedback is disabled; references are withheld.
  • Speed and cost summarize successful attempts only; completion denominators include failures. Missing billed usage is unknown, not zero.
  • Gateway provider routing, cache warmth and load were not controlled. Up to three runs were concurrent.
  • Addresses, coordinates, raw images, account identifiers and model responses remain private. The export supports metric auditing, not full independent reproduction.
  • The September 6 study used an older pipeline. Its 41 trials remain available separately and are not pooled into this ranking.
  • High-reasoning follow-up uses only the reviewed home, a larger output allowance and longer timeout. Its speed/cost averages are not directly comparable to the primary three-home averages. The effort field records what we requested; equal labels are not proof of equal internal compute.

The model draws.
You make the final call.

Astra was OpenLawn’s production default when this benchmark was published. Agent View now defaults to GPT-6.1 Sol and offers an Astra remeasurement after a completed Sol run. Neither model replaces a human review of a customer’s property.

Try OpenLawn free ↗