# OpenLawn Benchmark Report

Canonical page: https://openlawn.ai/#benchmark
Publisher: OpenLawn
Published: 2026-09-06

Internal workflow comparison, not independently audited or a universal model leaderboard. Three residential properties; only one has a user-corrected reference, not surveyed ground truth.

41 attempts, 25 completed; nine models and three residential properties. Only one property has a user-corrected reference. No surveyed accuracy claim is made.

## Overall quality index (0–100)

The landing-page score is mean outline overlap with the one user-corrected property, not a combined speed/reliability score or real-world accuracy percentage. Higher means closer agreement with that outline. Astra and Gemini Pro scored similarly; Astra completed maps sooner under its tested settings. Both comparisons are limited by the small sample and different configurations.

| Model identifier | Quality index | Reference runs | Configuration |
| --- | --- | --- | --- |
| openai/gpt-6-astra | 82.9 | 2 | Initial configuration; includes available repeats |
| google/gemini-3.1-pro-preview | 82.7 | 1 | Adjusted 8,192-token budget |
| anthropic/claude-opus-5 | 75.7 | 2 | Initial configuration; includes available repeats |
| bytedance/seed-1.8 | 73.4 | 1 | Initial configuration; includes available repeats |
| openai/gpt-5.6-terra | 61.7 | 1 | Initial configuration; includes available repeats |
| anthropic/claude-sonnet-5 | 31.0 | 1 | Initial configuration; includes available repeats |

Fable, Gemini Flash, and Grok have no reference score. Missing evidence is not a zero score. All failures and alternate configurations remain below and in the trial download.

## Every configuration, including failures and repeats

| Model identifier | Phase | Completed | Mean model-pipeline seconds | Reference overlap (IoU, %) | Reference trials |
| --- | --- | --- | --- | --- | --- |
| anthropic/claude-sonnet-5 | screen | 3/3 | 16.1 | 31.0 | 1 |
| openai/gpt-5.6-terra | screen | 3/3 | 30.5 | 61.7 | 1 |
| anthropic/claude-fable-5.1 | screen | 0/3 | Not available | Not available | 0 |
| openai/gpt-6-astra | screen | 3/3 | 46.4 | 90.7 | 1 |
| google/gemini-3.8-flash | screen | 0/3 | Not available | Not available | 0 |
| anthropic/claude-opus-5 | screen | 3/3 | 28.7 | 76.2 | 1 |
| google/gemini-3.1-pro-preview | screen | 0/3 | Not available | Not available | 0 |
| bytedance/seed-1.8 | screen | 3/3 | 34.4 | 73.4 | 1 |
| spacexai/grok-4.6 | screen | 0/3 | Not available | Not available | 0 |
| anthropic/claude-fable-5.1 | compatible | 1/3 | 33.1 | Not available | 0 |
| google/gemini-3.8-flash | compatible | 1/3 | 88.3 | Not available | 0 |
| google/gemini-3.1-pro-preview | compatible | 3/3 | 123.9 | 82.7 | 1 |
| openai/gpt-6-astra | low | 3/3 | 25.4 | 65.5 | 1 |
| openai/gpt-6-astra | repeat | 1/1 | 45.5 | 75.2 | 1 |
| anthropic/claude-opus-5 | repeat | 1/1 | 32.3 | 75.1 | 1 |

Mean time is computed only from successful attempts. A missing value is not zero. Astra medium's two corrected-reference runs ranged from 75.2% to 90.7% overlap; Opus's ranged from 75.1% to 76.2%. The landing-page overlap bars use the mean of both runs for those models. Others have one or no reference trial. These observed ranges are not confidence intervals.

## Protocol and limitations

- initial: Same saved pinned aerial, parcel, prompt, tools and parser per residence. Medium reasoning, 2400 output tokens per turn, up to 7 turns. No SegFormer. Corrected reference withheld.
- compatible: Gemini: 8192 tokens per turn. Claude Fable: auto tool choice instead of forced tool choice.
- low: Astra: low reasoning.
- repeat: One extra corrected-residence run each for Astra medium and Opus.
- limits: 90 seconds per turn; 180 seconds per run; at most 3 concurrent runs. Fixed image frame: reframe requests remain incomplete.
- timing: Direct model-pipeline wall time, not end-to-end app time. Excludes geocoding, parcel lookup, queue, browser rendering and saving. First polygon means tool callback arrival.
- limitations: Prompt was developed for the existing Astra workflow. Equal effort labels are not equal compute across providers. Gateway routing, load and cache warmth were not controlled. Too few references or repeats for general accuracy or reliability claims.
- privacy: Addresses, coordinates, imagery, record identifiers and raw responses are withheld. This export supports metric auditing, not full independent reproduction.

Initial Fable runs rejected forced tool choice; initial Gemini runs hit the output cap; Grok timed out. Compatibility retries increased Gemini's token budget and used auto tool choice for Fable. Integration failures do not establish poor vision capability.

[All 41 anonymized trials as JSON](https://openlawn.ai/benchmarks/2026-09-06.json)
