THE SIGNAL IN ONE SENTENCE
OpenAI slipped a small sentence into its API changelog on September 25 that should make every team running image evaluations sit up straight. The company says it fixed a bug in image encoding that degraded image understanding in GPT-6 Sol and GPT-6 Luna. It says the fix improves visual tasks in the API and Codex, including computer use. Then comes the important instruction: if a use case includes image inputs, rerun the evaluations and retry workflows affected by the issue. That is not a routine polish note. It means a score collected before the fix may describe a faulty input path rather than the model teams thought they were testing. The model may have been perfectly capable of reading a button, chart, receipt or screenshot that arrived through the wrong visual pipeline. Or the model may still fail after the fix. Until the same cases are rerun, nobody knows which explanation owns each miss. The notice is precise about the existence of a bug and broad about everything else. OpenAI does not publish when the faulty path began, how many requests were affected, which image formats or dimensions were most vulnerable, whether every request using the two models passed through the same encoder, or how large the quality change is by task. It does not publish old and new benchmark scores, affected model snapshots, request identifiers, regional differences or a severity classification. There is no public recipe for identifying one historical request as affected. OpenAI released GPT-6 Sol and GPT-6 Luna to the API on September 22. Three days later it published the image-encoding fix. The changelog says both models accept text and image inputs through the Responses and Chat Completions APIs. The short interval creates an obvious review window for early adopters, but the public notice does not say that September 22 through 25 is the complete affected period. A team should not quietly assume that it is. Ask OpenAI for a bounded incident window or inspect every visual result generated before the confirmed fixed version. An image model is not one box. A production request passes through a stack. The application loads or captures an image. It may resize, crop, compress, rotate or convert it. The client serializes it as a URL, file reference or encoded payload. The service validates and decodes the data. An image encoder converts pixels into a representation the reasoning model can use. The model interprets that representation alongside the prompt. A tool or application then turns the response into an action. A fault anywhere in that chain can look like model stupidity at the end. That distinction matters because teams often save the answer and forget the path. A benchmark spreadsheet might record model name, prompt and score while omitting the exact image bytes, request date, API endpoint, SDK version, image-detail setting, response identifier and preprocessing code. When the vendor fixes an encoding bug, those missing fields become the whole investigation. Without them, a team cannot say which results are stale, reproduce the old behavior or prove that the rerun used the same input. The cheap response is to rerun one glossy benchmark and replace the number. The useful response is to build an evaluation ledger. Start with the original image hash. Save the bytes used in the request, not merely the source file before resizing. Record width, height, format, color profile, orientation and file size. Record every crop, conversion, compression and scaling step. Save the complete request payload or a privacy-safe equivalent, including the model identifier, endpoint, prompt, image order, detail setting, tool configuration, reasoning setting, time, region and SDK version. Keep the response identifier, raw output, grader result, human judgment and downstream action. Add a pipeline version that names the code responsible for preparing the image. Then rerun the identical case after the fix. A clean comparison holds everything still except the thing being investigated. Use the same source bytes, prompt, option order, tools, grader and acceptance threshold. If the API does not expose the old model snapshot or buggy encoder, label the comparison honestly. It is a before-and-after service comparison, not an isolated encoder experiment. Run each case more than once when the output can vary. Separate a deterministic perception task from an open-ended reasoning task. Reading a seven-digit invoice total, locating a small button and describing a crowded street scene are different tests. A single average can hide the task that still breaks. Computer use deserves its own lane because one visual error can cascade. An agent may misread a button, click the wrong control, land on a different screen and then fail several later steps correctly given the wrong state. Scoring only the final task makes the entire chain look like one model miss. A better replay records screenshots and actions at every step. Recheck element recognition, target selection, action coordinates, state transition and recovery separately. The encoding fix may improve the first step without fixing planning, tool timing or action precision. A better screenshot interpretation does not automatically make an autonomous workflow safe. Teams should also review decisions already made from visual inputs. That does not mean every image request caused harm. OpenAI has not said that. It means that a known quality defect existed and the vendor cannot tell the public how to identify the affected subset. Triage by consequence. Start with workflows that moved money, changed access, submitted forms, approved claims, reviewed safety conditions, extracted medical or legal information, or acted without a person checking the image. Then review high-volume customer workflows and benchmark results used for model selection. A failed meme caption can wait. A misread medication label cannot. Rerunning an evaluation is not the same as repairing a production outcome. If an old visual result triggered a human decision or automated action, compare the corrected output with the action that actually happened. Create a remediation path for material differences. That could mean reopening a case, alerting an operator, correcting a record or contacting an affected customer. Keep the review proportional and privacy-aware. Do not copy sensitive images into a new test bucket without the same access, retention and deletion controls that governed the original workload. The notice also exposes a familiar procurement problem. Model names are not complete version identifiers. GPT-6 Sol on September 24 and GPT-6 Sol after the September 25 fix share a public name while the service behavior may differ. Hosted models change behind stable aliases because infrastructure, safety systems and preprocessing improve. That is normal service operation. It becomes a measurement problem when an evaluation report says only which model was used. Every externally meaningful score should carry a date, response identifier and provider change-state note. If the provider offers immutable snapshots, record them. If it does not, treat the time of execution as part of the version. OpenAI tells developers in its own evaluation guidance to evaluate early and often, design task-specific tests, log everything and continuously evaluate on every change. The encoding incident is a concrete example of why those habits matter. A continuous evaluation should not run only when a team edits its prompt. It should also run when a provider changes a model, encoder, tool, safety system or serving layer. Changelogs need to be machine-readable inputs to the evaluation schedule, not bedtime reading for whichever engineer still checks the docs. Build a change watcher with a small blast-radius map. Which products use GPT-6 Sol or Luna? Which accept images? Which use computer control? Which have stored test cases? Which can rerun automatically, and which require a human because the source images are sensitive or the outcome is subjective? When a relevant vendor change arrives, freeze the old report, create a new evaluation run and link the two. Never overwrite the historical score. The old number still documents what users experienced at that time. It simply needs a status such as superseded, affected window unknown or rerun required. Public leaderboards have the same obligation. If a visual benchmark used the affected service path, the result should carry a footnote and rerun date. A corrected score should not be presented as though it came from the original model launch. Researchers comparing systems need to know that the preprocessing changed. Procurement teams need to know whether a vendor decision was based on the old path. Product teams need to know whether a regression alert was actually a provider incident. An honest chart can contain both numbers. There are limits to what outsiders can conclude. OpenAI says the bug degraded image understanding and that the fix improves visual tasks. It does not say all image requests were wrong, that text-only performance changed, that a security boundary was crossed, or that the underlying model weights changed. It does not publish evidence that every affected workflow is now correct. It recommends retrying workflows, not blindly accepting new outputs. The repair addresses a named input-path problem. Teams still own their prompts, tools, graders, application code and decisions. The plain signal is that multimodal evaluation begins before the model sees a pixel. The image bytes, transformations, encoder, service date and tool chain are part of the system under test. OpenAI fixed one piece of that chain and did the responsible minimum by telling teams to rerun. Now the burden moves downstream. If a company cannot identify which visual scores came from which pipeline, it never had a durable benchmark. It had a spreadsheet with amnesia.
01
WHAT ACTUALLY CHANGED
OpenAI posted an API changelog fix on September 25, 2026.
The notice names GPT-6 Sol and GPT-6 Luna.
OpenAI says an image-encoding bug degraded image understanding in both models.
The company says the update improves visual tasks in the API and Codex.
OpenAI specifically includes computer-use workflows in the affected class of visual tasks.
Developers using image inputs are advised to rerun their evaluations.
OpenAI also advises retrying workflows affected by the issue.
GPT-6 Sol and GPT-6 Luna were released to the API on September 22.
The models accept text and image inputs through the Responses and Chat Completions APIs.
The public notice does not specify the start of the affected period.
The public notice does not specify how many requests or customers were affected.
OpenAI does not publish task-level before-and-after measurements for the fix.
OpenAI does not publish the exact encoder failure mechanism.
The notice does not provide a request-level test for identifying affected historical calls.
The notice does not say that text-only requests were affected.
The notice does not say that all image requests failed.
The notice does not state that the underlying model weights changed.
The recommendation makes earlier visual evaluation results candidates for review rather than automatic deletion.
OpenAI evaluation guidance separately recommends task-specific tests, comprehensive logging and continuous evaluation.
The fix turns request timing and image-pipeline provenance into necessary evaluation metadata.
02
WHY THIS MATTERS
A model can receive a degraded visual representation even when the original image is correct.
An input-pipeline fault can be mistaken for weak model reasoning.
A benchmark result is not reproducible if it omits the exact image bytes and preprocessing path.
A stable public model name can cover service behavior that changes after a provider fix.
The execution date therefore becomes part of the effective model version.
Image-input scores collected before the fix may no longer describe current service behavior.
An average benchmark score can hide task-specific failures in reading, localization or scene understanding.
Computer-use errors can cascade from one misread screen into several downstream actions.
Final-task scoring alone cannot show where a visual agent left the correct path.
Rerunning a benchmark does not repair decisions already made from an earlier output.
High-consequence visual workflows need outcome review in addition to technical retesting.
Without an affected-window disclosure, customers must choose a conservative review boundary.
Providers may be unable to expose internal infrastructure details, so customers need their own request ledger.
Historical scores should remain visible with a superseded or rerun-required status.
Overwriting an old score destroys evidence about what users experienced at that time.
A corrected visual pipeline does not prove that prompts, tools, graders or policies are correct.
Task-specific reruns are more informative than one broad visual benchmark.
Sensitive source images require the same privacy controls during reruns as during production.
Vendor changelogs should trigger evaluation workflows just as code changes do.
The incident shows that multimodal quality belongs to the complete system, not only the model weights.
03
WHERE IT COULD HELP
- Inventory every product using GPT-6 Sol or GPT-6 Luna with image inputs.
- Separate text-only products from image and computer-use workflows before triage.
- Ask OpenAI for the affected start time, end time, request scope and identification method.
- Preserve the original report and mark it rerun required instead of overwriting it.
- Hash the exact image bytes sent in every evaluation request.
- Record image dimensions, format, color profile, orientation and file size.
- Version every resize, crop, rotation, conversion and compression step.
- Store the model identifier, endpoint, SDK version, image detail and request time.
- Keep response identifiers, raw outputs, grader results and human judgments together.
- Rerun identical cases with the same bytes, prompt, tools and acceptance threshold.
- Repeat nondeterministic cases and report the distribution rather than one lucky output.
- Break visual evaluation into text reading, object recognition, localization and reasoning slices.
- Replay computer-use tasks one screen and one action at a time.
- Score element recognition, action choice, coordinate accuracy and state transition separately.
- Review past outcomes first where visual outputs moved money, access, claims or records.
- Create a remediation path when a corrected output materially conflicts with an earlier action.
- Apply the original retention, access and deletion policy to every rerun image.
- Trigger targeted evaluations from provider changelog events.
- Publish the rerun date, provider state and known limitations beside every updated score.
- Keep a regression set that includes typical, edge and adversarial visual cases.
KEEP A HAND ON THE WHEEL
OpenAI verifies a September 25 image-encoding fix for GPT-6 Sol and GPT-6 Luna, says the bug degraded image understanding, and recommends rerunning image-input evaluations and retrying affected workflows. The company identifies visual API tasks, Codex and computer use as areas that may improve. Its public changelog does not disclose the affected start time, request count, customer scope, image formats, regional distribution, exact encoder mechanism, task-level severity, before-and-after benchmark results, immutable model snapshots or a request-level method for identifying affected calls. It also does not claim that all image requests failed, that text-only performance changed or that the fix makes every visual workflow correct. Watch for a bounded incident window, customer notices, reproducible before-and-after measurements, request-level diagnostics, snapshot identifiers and clearer guidance for reviewing past automated actions.
04
TERMS WORTH KEEPING
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on September 28, 2026.
PUBLICATION RECEIPT: Revision 1. Published September 28, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US