Skip to main content
This site is an independent third-party technical service provider. Claude™ and Anthropic® are trademarks of Anthropic, PBC. This site has no affiliation, endorsement, or partnership with Anthropic.

Claude Marshmallow and Melon: what the leaks suggest

What the Marshmallow and Melon codenames reveal about Anthropic's model testing cadence, and which claims remain unverified.

Newsclaudeanthropicclaude apiai modelsindustry newsEst. read10 min
2026.08.25 published
Claude Marshmallow and Melon: what the leaks suggest

Two unannounced Claude model names, claude-marshmallow-eap and claude-melon-eap, surfaced in developer discussions over the past two days. The names have triggered claims about performance and product positioning, but Anthropic has not listed either model in its documentation or newsroom as of August 25, 2026.

Some developers said the names appeared in API access records or model lists shared within developer communities. The story quickly expanded: Marshmallow was supposedly stronger than Melon, its conversational style may have felt better than Opus 5, and neither model appeared to reach the Fable tier.

That is also where the evidence becomes less clear.

Anthropic’s model catalog, newsroom, and developer documentation currently list Fable 5, Opus 5, Sonnet 5, and Haiku 4.5 as the main public models. Marshmallow and Melon do not appear in those sources.

This is therefore not a Claude model launch report. A more accurate description is that the developer community found infrastructure traces that may belong to prerelease models. Those traces offer clues about Anthropic’s testing process and release cadence, but not confirmed product specifications.

Claude Marshmallow and Melon: facts, signals, and speculation

Discussion around the two codenames combines several kinds of evidence. They should not be treated as equally reliable.

Claim Current status How to interpret it
Screenshots show claude-marshmallow-eap and claude-melon-eap Supported by multiple reposts and screenshots The strings circulated in developer communities, but screenshots alone do not prove general API availability
Both names use the eap suffix Directly observable The community commonly reads this as Early Access Program, but Anthropic has not defined it for these models
Marshmallow performs better than Melon Subjective early-user feedback No public test set, parameter configuration, or reproducible result is available
Marshmallow feels better to chat with than Opus 5 A single class of experience reports A more pleasant conversation style does not prove better reasoning, coding, or tool use
The models map to Opus 5.1, Sonnet 5.1, or a new Haiku Community speculation Anthropic has not confirmed any product mapping
A release is close Timing prediction There is no announced release date, and test models may be renamed, merged, or canceled

Later reporting made the situation more nuanced.

Developer Chetaslua subsequently said that the plain claude-marshmallow-eap and claude-melon-eap names did not produce verifiable first-party output. The model that may have left API traffic on August 21 was instead named claude-marshmallow-ht-eap. According to the developer’s traffic analysis, that model returned its own name and fell back to Opus 4.8 for some safety-sensitive tasks.

That correction matters. A name appearing in a list, a request reaching a live model, and a model returning verifiable output are three separate events.

Early model reporting often collapses those three evidence levels into one.

Why the reported Claude model tests do not prove performance

Two screenshots from the original discussion illustrate the problem. One user said Marshmallow already felt better to chat with than Opus 5. Another said Melon was better than Opus 5.

The claims conflict, and neither includes the prompt, full output, sampling settings, number of trials, or comparison conditions. They show what the community is discussing, but they do not belong in a model leaderboard.

“Better to chat with” is also highly subjective. A response can feel more natural because it is shorter, follows the user’s tone, or avoids excessive explanation. The result could be very different on long-running agents, repository edits, retrieval accuracy, or complex reasoning tasks.

A useful model comparison should answer at least five questions:

  1. Did both requests use the same prompt and system instructions?
  2. Were temperature, effort level, and tool permissions identical?
  3. Was there any routing, fallback, or safety-classifier switch?
  4. Was the task a subjective conversation or an objectively verifiable workload?
  5. Were there enough trials to separate a pattern from a one-off result?

The public evidence does not yet meet that standard.

Food codenames matter less than an unpredictable naming system

Claude’s public product names have traditionally followed literary forms: Haiku, Sonnet, Opus, Fable, and Mythos. Reported internal names have followed several other themes, including animals and food.

Marshmallow and Melon fit the same food theme as other community-reported names such as Honeycomb and Fruitcake. That has led to a theory that Anthropic uses a “food name plus EAP” pattern for some prerelease checkpoints.

Internal codenames, however, do not reliably reveal the eventual product line.

One internal model may go through several training, post-training, and safety configurations before being folded into another release. A name seen outside the company might also identify a router, experiment group, or deployment environment rather than a standalone foundation model. The community previously connected Honeycomb to Opus 5, but Anthropic never confirmed a one-to-one relationship.

The codenames do not prove that “Opus 5.1 is here.” They suggest that Anthropic is testing models more frequently and across a wider set of systems. Developers may see the first traces in a model list, error message, client configuration, or response field long before the company settles on a public name and product position.

Why internal Claude model names are easier to spot

Model launches once looked like coordinated events: a newsroom post, model card, API documentation, and product entry appeared together. Agent-oriented deployment is more complicated, which creates more opportunities for prerelease traces to surface.

Before a checkpoint reaches general availability, it may pass through routing, access control, red-team testing, safety fallback, client compatibility, context-window testing, tool-use evaluation, and billing integration. If any of those stages use production-adjacent infrastructure, the model ID can briefly appear in logs, lists, or errors.

That makes the EAP suffix more informative than the dessert theme. It may point to a prerelease testing process rather than a marketing name:

  • Limit access to selected accounts or testers.
  • Observe behavior through real clients and API infrastructure.
  • Apply classifiers or fallback models to higher-risk tasks.
  • Expand access only after stability, cost, and safety reviews.
  • Finalize the public name, pricing, and regional availability later.

This workflow is a reasonable inference from public traces, not an Anthropic description of the Marshmallow or Melon programs.

Anthropic’s 2026 Claude model release cadence

The codenames make more sense when placed alongside Anthropic’s recent product releases.

Anthropic released Fable 5 on June 9, Sonnet 5 on June 30, and Opus 5 on July 24. Its current model catalog gives each one a distinct position: Fable 5 targets the highest-capability, long-horizon agent work; Opus 5 targets complex agentic coding and enterprise tasks; and Sonnet 5 balances speed, capability, and cost.

Model Anthropic’s stated positioning Public release date
Claude Fable 5 Highest capability and long-running agents June 9, 2026
Claude Sonnet 5 A balance of speed and intelligence for scaled workloads June 30, 2026
Claude Opus 5 Complex agentic coding and enterprise work July 24, 2026

Three primary product lines changed within two months. The appearance of more EAP codenames at least suggests that Anthropic no longer treats a major annual release as its only unit of iteration.

Model competition is shifting from “how often can a company announce a generation?” to “how continuously can it train, evaluate, route, and deploy multiple checkpoints?” A public launch is only the most visible moment in a much longer operating cycle.

What Marshmallow and Melon may be testing

Without model cards or reproducible evaluations, it would be inaccurate to rename these models Opus 5.1 or Sonnet 5.1. Anthropic’s current lineup still points to several useful areas to watch.

1. Experience improvements may matter more than a higher ceiling

Community reports repeatedly mention conversational quality rather than a large benchmark gain. The experiments may therefore include tone, instruction following, response length, and multi-turn consistency.

Those qualities matter in production. A model with a higher benchmark score may still be a poor fit if it overexplains routine tasks or loses direction during a long workflow.

2. New checkpoints may target finer cost tiers

Anthropic already maintains a capability and pricing ladder across Fable, Opus, Sonnet, and Haiku. Two EAP models with different reported strengths could be testing new cost-performance points rather than a single new flagship.

Most organizations do not need a frontier model for every request. Summarization, classification, code review, and complex reasoning can run on different tiers. A model family that is easier to route can become more useful in long-term workflows.

3. Safety fallback may remain part of model deployment

If the reported marshmallow-ht-eap fallback to Opus 4.8 is accurate, the test may cover safety routing as well as model capability.

Anthropic has already documented a related mechanism for Fable 5: queries in some topic areas are handled by the next-most-capable model. Evaluating a future Claude release may therefore require looking beyond its primary model name to understand the router, classifiers, and fallback behavior around it.

How developers should prepare for new Claude models

The practical answer is deliberately conservative: do not migrate early, and do not put internal codenames into production configuration.

Anthropic’s documentation says that public Claude model IDs identify pinned model versions. The currently documented IDs include claude-fable-5, claude-opus-5, claude-sonnet-5, and claude-haiku-4-5-20251001. Marshmallow and Melon are not on that list.

If an application needs to adopt new models quickly, keep the model ID in configuration or a routing layer instead of hard-coding it into business logic:

CLAUDE_FAST_MODEL=claude-sonnet-5
CLAUDE_STRONG_MODEL=claude-opus-5
CLAUDE_FRONTIER_MODEL=claude-fable-5
CLAUDE_FAST_MODEL=claude-sonnet-5
CLAUDE_STRONG_MODEL=claude-opus-5
CLAUDE_FRONTIER_MODEL=claude-fable-5

When a new model reaches the documentation and your account’s model list, run a staged evaluation before changing the configuration. That keeps naming, pricing, or capability changes from forcing a broader application rewrite.

The same principle applies when using Apito, an AI API gateway. Call only the models that are currently visible in your account console. The live model catalog, pricing, and availability should come from the console, not from an internal codename circulating in developer communities.

Model release mechanics matter more than guessing the name

No one outside Anthropic can yet say what Marshmallow and Melon will be called, or whether either model will reach a public release.

The codename story still reveals a more durable change. Models are no longer developed only through a release event every few months. They are tested continuously across allowlists, clients, routers, and safety systems.

What looks like a leak may be nothing more than a label that fell off that assembly line.

The label may disappear, the product may be renamed, and the checkpoint may never reach the public catalog. Three signals remain more reliable than community speculation: the model appears in Anthropic’s documentation, the API is available to the intended account, and the model card and pricing are published.

Until those conditions are met, Marshmallow and Melon are worth watching, but they are not a reason to change production code.

FAQ

Have Claude Marshmallow and Claude Melon been released?

No. As of August 25, 2026, neither name appears in Anthropic’s model catalog or newsroom. The available information comes primarily from developer-community screenshots and reports.

Does eap mean Early Access Program?

That is the community’s common interpretation, but Anthropic has not defined the suffix for these two model names. Treat it as an early-access signal, not evidence of a release date.

Is Marshmallow Opus 5.1?

There is no evidence confirming that mapping. Opus 5.1, Sonnet 5.1, and a new Haiku are all community theories. An internal codename does not have to map directly to a public product.

Is Marshmallow stronger than Opus 5?

Only scattered subjective reports are available, and some community claims conflict with each other. No public, reproducible evaluation shows that Marshmallow is broadly more capable than Opus 5.

Can developers call these models through the Claude API now?

Do not treat community screenshots as evidence of availability. Use Anthropic’s documented model list or the model catalog actually visible to your account.

How should a production system prepare for new models?

Keep model IDs in environment variables or a routing layer, and maintain a fixed acceptance-test suite. When a model is formally available, run a staged evaluation before replacing an existing production model.

Sources

Verification note: Marshmallow, Melon, and marshmallow-ht-eap remain unconfirmed community reports. Use Anthropic’s published announcements and live model catalog for product names, capabilities, pricing, and release dates.

Related Articles