Community report2026-08-20

DeepSeek V4 Pro Gray-Test Reports: Is a Stronger Build Returning?

Some DeepSeek users say they are seeing an earlier-looking reasoning style and better long-task results in V4 Pro. The reports may point to a traffic experiment, but DeepSeek has not confirmed a new model, checkpoint, or routing change.

Chinese DeepSeek status-page screenshot showing Chat Service labelled with five components
Community screenshot supplied to DeepSeek V4 Hub, originally circulated on Xiaohongshu and bearing the watermark AI6667789. It does not establish what the five components represent or prove a model-routing change. Select the image to open the full-size version.

Some DeepSeek V4 Pro users believe they may be seeing a new gray test. The evidence is a change in the visible reasoning text: a few users say replies that once opened with phrases such as We need or Let’s now begin with lines like I’m weighing, I’m settling, or I’m validating.

That is interesting, but it is not proof of a new model. DeepSeek has not announced a new V4 Pro checkpoint, a model-route change, or a fresh public test. For now, this is a community report about uneven behavior across sessions and configurations, not a confirmed release.

Why users are watching the “I’m doing” pattern

A post on the Linux.do technical community said the pattern appeared while using DeepSeek Harness in Standard mode with V4 Pro and the Max reasoning setting. Other users reportedly reproduced something similar in the web product and through the official API, while others continued to receive the older style. The contrast is why the phrase has become a rough community fingerprint rather than a reliable version label.

The claim has a backstory. During a reported V4 Pro gray test in late July, some developers said the model felt much stronger than the build that later became the August 13 public release. A few compared the earlier experience with Anthropic’s flagship Fable 5. Those comparisons were personal impressions, not a shared benchmark run, and they do not establish that the same checkpoint is now returning.

Visible reasoning text is also a fragile signal. It can change with the prompt, tool definitions, system instructions, model settings, or the surrounding agent harness. A different opening phrase may reflect a configuration change without reflecting a material change in model quality.

The August release is part of the speculation

The August 13 V4 Pro release prompted reports that its visible reasoning prose and task behavior differed from the earlier gray-test experience. Some community testing claimed that scores shifted with different agent-tool exposure, and one later analysis suggested that tool configuration, rather than the model itself, constrained results. Those reports point to a plausible testing problem: model quality and harness quality are hard to separate when an agent has a changing tool set.

They do not demonstrate that DeepSeek shipped a reduced model and is now restoring a hidden full version. Nor do they support a more dramatic rumor that DeepSeek is silently routing requests to another company’s model. There is no public evidence for either claim in the material reviewed for this article.

What the reported timing could, and could not, mean

One account said the “I’m doing” pattern became difficult to trigger after 5 p.m. Beijing time. The same report described more than 160 direct official-API calls, using a standard DeepSeek Harness configuration, in which the wording appeared only once. That is an observation from one test, not a controlled study. It may be consistent with a time-based traffic allocation, but prompt variance and harness sensitivity are alternative explanations.

The timing has become a joke in the community because DeepSeek’s official V4 API pricing now distinguishes peak and off-peak periods. The official pricing page lists peak hours of 09:00–12:00 and 14:00–18:00 Beijing time, with off-peak rates at half the peak rate. The schedule is a billing rule; it does not say that different model versions run at different hours.

A screenshot of five Chat Service components is not a release note

A separate Chinese social post pointed to a status-page screenshot where “Chat Service” is labelled with five components and claimed the service had previously shown two. That could be worth watching if DeepSeek publishes an explanation. By itself, though, the screenshot cannot tell us whether the count reflects an operational status grouping, capacity work, an internal routing layer, or a new dialogue model.

The post also links the timing to the recent API price change and suggests that lighter compute pressure may have made room for broader testing. That is editorial speculation. DeepSeek has not publicly connected its pricing policy, capacity, and any model experiment.

What would count as confirmation

A meaningful confirmation would be a DeepSeek release note, an updated model identifier, a documented routing policy, or a reproducible evaluation with the same prompts, tools, settings, and account conditions. Until then, reports of better long-running tasks are useful leads for developers to test on their own workloads, not a basis for a capability claim.

For now, the narrow conclusion is straightforward: some users may be encountering a different V4 Pro behavior, and the available clues are too weak to identify why. If a stronger gray-test build is in circulation, DeepSeek has not said so. If it is not, the episode still shows how much tool configuration and visible reasoning style can shape perceptions of an agent model.

Sources