How to Retest GPT-6 Vision Agents After the Sept 25 Fix (2026 Guide)

Matt Payne··Updated ·7 min read
Key Takeaway

OpenAI fixed a GPT-6 image-encoding bug on Sept 25, 2026. Luna's RefCOCOg score jumped from 28.8% to 60.0%. Retest visual agents with 20 real files, a before/after harness, and routing rules before turning workflows back on.

GPT-6 Vision Broke. Re-Test It in 60 Minutes

The GPT-6 image-encoding bug fix changed how Sol and Luna processed image inputs. Reports say it affected API, Codex, and computer-use tasks.

The model received a damaged or degraded visual representation.

That distinction matters.

Step 1: Build a 20-File Test Set in 15 Minutes

Don't test with random stock images. Use files from your real sales and marketing workflows.

Twenty files are enough for a fast signal:

WorkflowFilesWhat to test
CRM screenshots4Contact name, status, owner, next action
LinkedIn screenshots2Name, title, company, visible activity
Lead PDFs4Company facts, tables, footnotes, dates
Invoices or order forms2IDs, totals, line items, addresses
Ad dashboards4Spend, clicks, CPL, campaign status
Web interfaces4Button location, disabled states, navigation

Include one clean file and one ugly file for each task.

The ugly files matter more. Use small text, dense tables, cropped pages, and mobile screenshots.

Write the correct answer before running anything. Don't score the model from memory.

For invoice numbers, require exact characters. A missing leading zero is a failure.

For ad dashboards, define tolerances. A spend value of `$10,421.17` shouldn't pass as `$10,412.17`.

Hash every input with SHA-256. That proves the before and after runs used identical bytes.

Python's built-in `hashlib` costs $0. ImageMagick also costs $0 and can record width, height, format, and color profile.

Record these fields:

  • Test ID
  • File hash
  • File format
  • Pixel dimensions
  • Model ID
  • Provider route
  • Image detail setting
  • Prompt version
  • Expected answer
  • Actual answer
  • Pass or fail
  • Latency
  • Provider-reported cost
  • Tool action result

Don't claim a before-and-after improvement without saved pre-fix outputs. Call it a "post-fix evaluation" instead.

Expected outcome: A fixed, scored set of 20 real files with known answers.

Step 2: Run a Controlled Before/After Harness in 20 Minutes

Keep your prompt unchanged.

Keep the image, parser, tool settings, and output schema unchanged too.

Run each file three times. Twenty files produce 60 runs.

That shows output variance without burning an afternoon.

Compare four modes where your current API supports them:

1. Existing production route and settings 2. Same route with higher image detail 3. Vision-only extraction without tool use 4. Full agent flow with tool use enabled

This split tells you where the failure lives.

If extraction fails, inspect the image path. If extraction passes but clicking fails, inspect the action tool.

Ofox recommends separating "understand" from "act." That's right.

Ask the model to identify the target first. Then let the agent click it.

A correct description followed by a bad click points to coordinates, the browser, or the action tool.

Use a JSON schema like this:

{"id":"crm-01","passed":true,"latency_ms":1840,"cost_usd":0.0042} {"id":"pdf-03","passed":false,"latency_ms":2910,"cost_usd":0.0081}

This Python script compares saved result files:

import json import statistics import sys

def load(path): with open(path) as f: return {x["id"]: x for x in map(json.loads, f)}

before = load(sys.argv[1]) after = load(sys.argv[2]) shared = sorted(set(before) & set(after))

def rate(data): return sum(data[i]["passed"] for i in shared) / len(shared)

print("Cases:", len(shared)) print("Before pass rate:", round(rate(before) * 100, 1)) print("After pass rate:", round(rate(after) * 100, 1)) print("Delta:", round((rate(after) - rate(before)) * 100, 1), "points") print("After median latency:", statistics.median(after[i]["latency_ms"] for i in shared), "ms") print("After total cost: $", round(sum(after[i]["cost_usd"] for i in shared), 4))

Run it with:

python compare.py before.jsonl after.jsonl

Python costs $0. Your model cost depends on current image and token rates.

Don't use an old pricing screenshot. Pull usage from the API response or provider bill.

Lookonchain reported that Luna's RefCOCOg score rose from 28.8% to 60.0%. That benchmark measures whether a model locates the requested image region.

That 31.2-point jump matters. It doesn't prove your invoice parser improved by 31.2 points.

Expected outcome: A measured accuracy delta tied to your files, prompts, and route.

Step 3: Find the Real Failure Layer in 15 Minutes

Visual agents have at least five failure layers.

Don't blame the model before you check the rest of the stack.

1. Capture failure

Open the exact file sent to the API.

Check for blank screenshots, missing overlays, bad cropping, or login-protected URLs. A local filename in a prompt isn't an image upload.

2. Encoding failure

Compare PNG and JPEG versions of the same screenshot.

Image preprocessing can change what a model sees. A 2026 study found that resizing methods changed whether hidden visual instructions appeared.

That paper reported payload rendering rates from 55.7% to 92.5% across three models. Resizing can change the result.

3. Perception failure

Ask for extraction only.

Require exact text, visible state, and target location. Don't let the model click yet.

4. Action failure

Compare the predicted target with the browser's actual DOM geometry.

If the DOM places the button elsewhere, check more than the screenshot reader. Fixed screenshots can hide clipping and off-screen content.

5. Parser failure

Inspect the raw model answer before your workflow parses it.

The model may return the correct invoice ID. Your regex may drop the first zero.

For PDFs, test both native file input and rendered pages. The GDP.pdf benchmark found repeated failures on merged headers, footnotes, legends, and cross-page references.

Those problems show up in sales contracts, media reports, and pricing sheets.

Use ImageMagick and Python for file checks. Use Playwright for browser tests. n8n Community Edition can run the workflow. Each has a $0 software license.

Expected outcome: Assign every failure to capture, encoding, perception, action, or parsing.

Step 4: Add Routing Rules Before Turning It Back On

A visual agent shouldn't get one attempt and unlimited authority.

Give it routes based on risk.

InputFirst routeValidationFallback
Text-heavy PDFNative PDF inputRequired fields presentRender problem pages
Scanned PDFHigh-detail page renderOCR pattern checksAlternate vision model
Web interfaceDOM locator firstLabel and state matchScreenshot grounding
CRM screenshotVision extractionName and status requiredHuman review
InvoiceOCR plus visionTotals reconcileStop processing
Ad dashboardVision extractionMetric ranges checkedSecond model review

For UI actions, require three facts before clicking:

  • Exact visible label
  • Current enabled or disabled state
  • Target location or DOM selector

If any fact is missing, don't click.

For financial documents, reconcile the math. Subtotal plus tax should equal the total.

For lead research, require identity agreement. The name, company, and title must point to the same person.

For ad QA, validate against account rules. A campaign marked "Paused" shouldn't trigger a launch action.

Track failures by route and file type:

SELECT model_id, provider_route, file_type, COUNT(*) AS runs, AVG(CASE WHEN passed THEN 1.0 ELSE 0.0 END) AS pass_rate, AVG(latency_ms) AS avg_latency_ms, SUM(cost_usd) AS total_cost FROM visual_agent_runs WHERE created_at >= CURRENT_TIMESTAMP - INTERVAL '24 hours' GROUP BY model_id, provider_route, file_type ORDER BY pass_rate ASC;

Alert when pass rate drops five percentage points from your seven-day baseline. Also alert on three consecutive failures for one task.

The second rule catches narrow regressions faster.

Selenium began at ThoughtWorks in 2004 because web interfaces kept changing. Browser automation has always needed repeatable tests.

Visual agents need the same discipline.

Expected outcome: Bad reads stop before they become bad clicks, emails, or CRM updates.

FAQ

Has OpenAI fixed the GPT-6 image encoding bug?

Secondary reports say OpenAI fixed the issue on September 25, 2026. They cover GPT-6 Sol and Luna across visual API and Codex tasks.

Astra's scope wasn't clearly disclosed. Don't assume every third-party route updated at the same time.

How do I know whether my visual agents were affected?

Check for more OCR errors, wrong UI targets, missed labels, or PDF extraction failures. Compare saved pre-fix runs with identical post-fix inputs.

If only action steps fail, the issue may be browser tooling. If extraction fails too, check capture and encoding first.

How should I rerun image evals?

Use 10 to 30 real files with written answers. Keep image bytes, prompts, settings, routes, and parsers fixed.

Run each case three times. Score accuracy, action success, latency, and actual provider cost separately.

Did the GPT-6 image-encoding fix change pricing or endpoints?

No supplied report ties the September 25 fix to new endpoints or pricing. GPT-6 Sol and Luna launched separately on September 22 with usage-based billing.

One launch report cited prices 50% below GPT-5.6 promotional rates. Don't confuse launch pricing with the bug fix.

What routing rules should visual agents use now?

Use native PDF parsing before page images when possible. Use DOM selectors before screenshot coordinates for web actions.

Require validation before any click, send, update, or payment step. Send failed checks to another model or a person.

Related Reading

AI Answer

Did OpenAI fix the GPT-6 image encoding bug?

Secondary reports say OpenAI fixed the GPT-6 image-encoding bug on September 25, 2026. Luna's RefCOCOg score rose from 28.8% to 60.0% after the fix. The fix covered GPT-6 Sol and Luna across visual API and Codex tasks.

AI Answer

How long does it take to retest visual agents after the GPT-6 vision fix?

A full retest takes 60 minutes using 20 real files with written answers. Build the test set in 15 minutes, run the before/after harness in 20 minutes, and diagnose failure layers in 15 minutes. The Python comparison script and SHA-256 hashing tools cost $0.

AI Answer

What pass rate drop should trigger an alert for visual agent workflows?

Alert when pass rate drops five percentage points below your seven-day baseline. Also alert on three consecutive failures for a single task. The second rule catches narrow regressions faster than the baseline threshold alone.