IMMERSIVE COMMONS · THE SIGNALISSUE 18 · 9 — 15 AUG 2026
OPEN INTELLIGENCE · ISSUE 18

THE SIGNAL
9 — 15 AUG 2026
FRONTIER TOWER
18

Two Labs Declined To Ship

Anthropic's risk report disclosed an internal model called Model 2 that beats its shipped flagship and said plainly that it has no plans to release it, while raising its own misalignment estimate from "very low" to "low" and citing recent cybersecurity incidents. OpenAI, a week after slowing Astra for the same reason, went the opposite direction on the same premise: it opened Daybreak Red and shipped GPT-5.6-Cyber, a model trained to refuse LESS for vetted defenders. Capital did not read either as a warning. Anthropic is being priced off a 2028 revenue forecast of $190 to $200 billion against a $47 billion run rate, and General Catalyst put $1.1 billion into a company that was two months old. Underneath all of it, the contested layer stopped being the model and became the memory: Spotify open-sourced its institutional-context harness, a paper measured agentic prompt files growing 226% and named the disease, and someone proposed a wire format so agent memory can move between vendors at all.

BEATS 05
DISPATCHES 09
CHAIN MYTHOS × 03
PUBLISHED 2026-08-15
I.

THE LABS DECLINED TO SHIP

Anthropic disclosed a model better than its flagship and said it has no plans to release it. OpenAI answered the same cyber premise by shipping a model trained to refuse less.

220FIELD REPORTMYTHOS · CHAIN

Anthropic Built A Better Model And Shelved It.

The latest risk report discloses an internal model that beats the shipped flagship, raises the company's own misalignment estimate, and says there are no plans to release it.

Axios illustration accompanying its report on Anthropic's unreleased Model 2
IMAGEAxios

Anthropic's latest risk report, covered by Axios on August 14th, describes an internal model the company calls Model 2 that shows *"noticeable improvement"* on many internal tasks over the shipped Mythos tier. Both Mythos 5 and Model 2 are used *"heavily"* inside the company for coding, agentic work, and data generation. The report's disposition is one sentence long: *"We do not currently have plans to release this model externally."* In the same document Anthropic raised its broad estimate of misalignment risk in high-stakes situations from *"very low"* to *"low,"* citing recent cybersecurity incidents.

The mechanism worth reading twice is not the shelving, it is the reason given for the uncertainty. The report says the company is *"less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations... no longer capture increases in models' capabilities."* That is a lab stating on the record that its measuring instrument has fallen behind the thing it measures. Anthropic frames Model 2 as routine — *"we internally train and evaluate many different exploratory versions of models that we don't intend to release"* — and notes the jump is smaller than Opus 4.6 to Mythos earlier this year. Both things can be true, and the eval gap is the load-bearing one.

This is the second time in eight days a frontier lab has taken a capability off the table and said cyber. OpenAI slowed Astra on August 7th because it *"cannot rule out critical cyber capabilities."* Anthropic rolled back its own training-pause commitment in a February update to its Responsible Scaling Policy, on the argument that unilateral restraint makes the world less safe — and has now unilaterally restrained anyway, on a model nobody outside the company had asked about. Restraint that arrives as a disclosure rather than a policy is not a commitment. It is a decision that can be revisited next quarter by the same people who made it.

Axios (primary)Axios — OpenAI slows Astra
221FIELD REPORT

OpenAI Shipped A Model Trained To Refuse Less.

Daybreak Red hands vetted defenders GPT-5.6-Cyber, tuned down on refusals for exploit development — and it has already found a real Chrome bug.

OpenAI's announcement of Daybreak access tiers and the GPT-5.6-Cyber model
IMAGEOpenAI

On August 10th OpenAI expanded Daybreak into two access tiers and introduced GPT-5.6-Cyber. *Daybreak Blue* gives approved defenders GPT-5.6 Sol with the production request-screening guardrails removed, for vulnerability discovery, secure code review, malware analysis, and incident response. *Daybreak Red* goes further: it carries a purpose-trained cybersecurity model built on Sol and explicitly *"trained to further reduce refusals"* on exploit-chain development, authentication bypass, and privilege escalation. The company's framing is a clock — defenders have *"a narrowing window to prepare"* before offensive AI arrives at scale.

The proof point is not a benchmark, it is a CVE. OpenAI pointed GPT-5.6-Cyber at **V8**, the JavaScript engine in Chrome, and uncovered two previously unknown vulnerabilities that could be chained to corrupt memory and escape the V8 heap sandbox. Researchers validated them, disclosed to Google, and the fix shipped as CVE-2026-15903 — the optimizing compiler skipping a safety check on integer conversion. On ExploitBench, which leaves the V8 sandbox enabled and withholds detail about the target bug, Sol under Daybreak Blue is the more token-efficient solver at the standard 300-turn limit; the gap narrows at 600 turns.

Set this beside Anthropic shelving Model 2 four days later and the week resolves into one question asked twice with opposite answers. Both labs accept that frontier cyber capability is the binding risk. Anthropic's response is to withhold the capability from everyone; OpenAI's is to withhold it from most people and deliberately un-refuse it for a vetted list including SpecterOps, SentinelOne, and Palo Alto Networks. The second answer is the harder one to audit, because the safety property now lives in an access-control decision rather than in the model. A refusal you can measure has been replaced by an allowlist you cannot.

OpenAI (primary)
II.

THE MONEY DID NOT BLINK

In the same seven days, Wall Street priced Anthropic off a 2028 forecast four times its current run rate, and a two-month-old company raised $1.1 billion.

222FIELD REPORT

Wall Street Is Pricing Anthropic Off 2028.

A $190 to $200 billion revenue forecast two years out, against a $47 billion run rate today, because nobody can value the company on what it earns now.

Reuters illustration of the Anthropic logo with a keyboard and robotic hand
IMAGEReuters

Reuters reported on August 14th that Anthropic is projecting 2028 revenue of roughly $190 billion to $200 billion, a figure two people familiar with the company's financials say has not been reported before. The comparison that matters is internal: Anthropic publicized a $47 billion run rate as recently as May. Bankers are applying enterprise value-to-revenue multiples to the 2028 number rather than to anything the company earns today, four sources told Reuters, ahead of an analyst day.

Revenue multiples are ordinary for high-growth software without a mature profit profile. Reaching *two years forward* is not. It is what you do when current EBITDA does not describe the business — and Anthropic's does not, because GPUs, model training, inference, and hiring are compressing margins that investors are betting expand later. The comp set is the tell: Palantir at 53 times this year's expected revenue, SpaceX and Cloudflare both at 41.6 times expected 2026 revenue, per LSEG. Cerebras cited 2028 expectations before its IPO this year; SpaceX's projections ran to 2029 before it went public in June at a record valuation.

Read this against the same week's other Anthropic story and the dissonance is the point. The risk report says the company's own evaluations *"no longer capture increases in models' capabilities"* and raised its misalignment estimate. The IPO book says underwrite four times growth through 2028. Neither document is lying; they are answering to different committees. But a valuation built on a 2028 forecast is a bet that nothing between here and there forces a pause — and the lab itself just demonstrated, by shelving Model 2, that it is willing to be the thing that pauses.

Reuters (primary)
223FIELD REPORT

A Two-Month-Old Company Raised $1.1 Billion.

Igor Babuschkin left xAI to rebuild the stack end to end for agents that belong to you, and General Catalyst wrote the round before there was a product to price.

River AI, the startup founded by xAI co-founder Igor Babuschkin
IMAGETechCrunch

River AI, founded by xAI co-founder Igor Babuschkin, closed $1.1 billion in a combined seed and Series A led by General Catalyst and AMP PBC, with Nvidia, AMD Ventures, Y Combinator, and Temasek participating. AMP PBC is itself new — an AI investment firm started in 2026 by former Andreessen Horowitz general partner Anjney Midha. River came out of stealth in June. It was two months old.

The thesis is a deliberate fork from every other lab's. Babuschkin, previously at DeepMind and OpenAI, wants agents that are *personally trainable* rather than human-worker replacements, and argues the whole stack has to be rebuilt for it: *"training, models, the product layer, and new hardware that lets personal AI live close to you."* The shipped surface today is an API billed per million tokens across open models, exposing both reinforcement learning and LoRA fine-tuning to developers — a product whose stated purpose is to be an antidote to prompt engineering. Not a chat box. A training loop you point at yourself.

Nvidia and AMD both being on the cap table of a company that intends to build new hardware is the detail to file. But the number is the story. A $1.1 billion round into a two-month-old company, in the same seven days two frontier labs documented why they are slowing down, is capital saying it does not believe the deceleration is real — or does not believe it will last long enough to matter to a 2032 exit. Both readings price the same thing: that the constraint the labs are describing is temporary.

TechCrunch (primary)
III.

MEMORY BECOMES THE CONTESTED LAYER

Spotify shipped the context harness, a paper named the disease, and a protocol proposed that agent memory should be able to leave the vendor that stored it.

224FIELD REPORT

Spotify Shipped Its Institutional Memory.

Xirp is a model-agnostic coding harness whose product is not the agent but the context — who owns the service, what depends on it, why it was built that way.

Xirp, Spotify's agentic development environment built on Backstage Portal
IMAGEXirp

Spotify opened the beta for **Xirp**, an agentic development environment built on top of Portal, its commercial Backstage distribution. The pitch is narrow and unusually honest about what went wrong: AI coding tools *"solved the generation problem"* — output went up, more services got built — and the second-order effect was that new engineers took months to ramp and agents made decisions that were *"technically correct and operationally wrong, because they didn't know what they were actually working inside."*

The mechanism is a deliberate refusal to treat this as a documentation problem. Spotify's framing is that the knowledge was never missing — it was *"in Slack threads nobody could find, in the heads of the three people who built that service in 2021"* — which makes it a retrieval problem. So Xirp connects the agent to service ownership, dependency graphs, and architectural decision records at session start, and closes the loop the other way: every session writes back work items, sessions, and generated documentation into a shared Workspace, so the next engineer or agent inherits it. It is a harness, not a model — it runs Claude, Gemini, or Codex, and the explicit selling point is *"knows your org, not theirs."*

That last clause is the strategic move. If the durable asset is organizational context rather than the model, then the harness is the layer with lock-in and the model is the commodity underneath it — which is exactly the bet a company with a decade of service-catalog data should make. For anyone running agents against a codebase of real size, the transferable lesson costs nothing to adopt: the ceiling on your agent is not the model you picked, it is whether the session starts knowing what it is inside.

Xirp (primary)Backstage
225FIELD REPORT

Somebody Measured Why Your Agent File Only Grows.

Across 247,694 instruction lifetimes in 1,867 repositories, agentic prompt files more than tripled — and deleting a line safely costs exponentially more than adding one.

arXiv paper naming catastrophic remembering in agentic coding prompt files
IMAGEarXiv

A paper posted this week, **Why Does Claude.md Keep Growing? Catastrophic Remembering in Agentic Coding**, does the thing nobody had bothered to do: it measured the file. Across 247,694 instruction lifetimes in 1,867 repositories, agentic prompt files grow without bound, stopping only *"when the repository retires or someone rewrites the file wholesale."* They more than triple over their lifetime — +226% — gaining 4.9 net instructions every commit. And the older an instruction gets, the *less* likely anyone is to remove it, at a log-hazard of -0.032 per commit.

The mechanism is an asymmetry in cost, not a failure of discipline. Appending an instruction is always cheap. Deleting one is not, because once the *rationale* for a line is gone, establishing that removing it will not cause a regression costs O(2^|D|) in a prompt of |D| instructions — you cannot know which of the other instructions it was load-bearing against. The authors name the result catastrophic remembering, deliberately inverting the catastrophic forgetting that all of continual learning is organized around. Their fix is almost insultingly simple: *comments*. Encoding the latent reasoning next to the instruction removed 99.3% of excess instructions in verifiable worlds built by inverting IFEval, collapsing growth from +211.3% to +1.4%, and improved real-world instruction-following on WildIFEval by up to 23.1%.

Anyone maintaining a `CLAUDE.md`, an `AGENTS.md`, or a system prompt of any age is running the experiment this paper describes, and the ratchet is invisible because every individual append was correct at the time. The actionable finding is not "prune your prompt" — pruning is the expensive operation, which is the whole point. It is that a one-line *why* written beside a rule at the moment you add it converts an undeletable instruction into a deletable one. The paper closes on the sentence the field will end up quoting: *"If English is the new code, why don't we have comments yet?"*

arXiv 2608.11095 (primary)
226FIELD REPORT

Agent Memory Got A Wire Format.

Six memory frameworks, six SDKs, no way to migrate. memorywire proposes five operations and a governance channel so a human can review a write before it becomes permanent.

memorywire, a vendor-neutral wire format for agent memory operations
IMAGEarXiv

**memorywire** starts from an inventory rather than an idea. mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, and MemTensor each ship their own SDK, storage layout, and operational vocabulary. There is no shared wire format, so *"every integration is bespoke, every migration rebuilds memory from scratch,"* and — the part that should bother anyone running agents in production — no framework ships a governance surface that lets a human review writes before they enter long-term storage. The proposal is a JSON Schema 2020-12 format covering five operations (remember, recall, forget, merge, expire) across four memory types, with a store interface, a fan-out router, and an optional human-in-the-loop channel.

The result worth reading is adversarial, not architectural. In a fusion sweep injecting a single poisoned result at rank 0, **Reciprocal Rank Fusion** held recall@5 at 1.000 across every K tested, while max-score fusion collapsed to 0.500 with an 80% leak rate at K of 5 or more. That is a concrete statement that how you merge memory retrievals is a security property, not a tuning knob. The reference implementation ships five backend adapters — sqlite-vec, mem0, Letta, Cognee, pgvector — and a 16-scenario conformance suite passing 68 of 80 cells with zero failures.

The authors are direct that this is not a new algorithm; it is a packaging of RRF, state machines, short-to-long-term consolidation, and diff-and-approve workflows into a venue-neutral protocol positioned to compose with MCP rather than compete with it. That is the correct read of where the fight is. MCP standardized how an agent reaches a tool; nothing has standardized what an agent is allowed to remember, who reviews it, or whether it can leave. Pick your memory vendor today and you have picked a data format you cannot export.

arXiv 2606.01138 (primary)
IV.

THE SKILL FILE IS THE ATTACK SURFACE

Two enterprise coding agents, 5,629 runs, and a compromise rate that should stop a procurement process.

227FIELD REPORT

Gemini CLI Fell For 96% Of Poisoned Skills.

2,826 adversarial skill files, 5,629 runs, two enterprise coding agents — and explicit safety recognition in under two percent of them.

arXiv paper assessing malicious skill files in enterprise coding agents
IMAGEarXiv

**Towards a Risk Assessment of Malicious Skill Files in Coding Agents** attacks the interface every agent platform shipped in the last year: folders of instructions and scripts an agent loads dynamically to specialize itself. The authors used six LLMs across four model families to transform 471 real-world shell commands into benign-appearing skills, releasing 2,826 of them mapped to 11 [MITRE ATT&CK](https://attack.mitre.org/) tactics as a public benchmark. Then they ran two enterprise-grade agents against it — 5,629 completed runs.

Gemini CLI is exploited in 95.5 to 96.1% of runs. Qwen Code in 71.6 to 74.0%. The range spans raw majority vote to a declared-intent-corrected estimate, and both land inside a blind human gold standard the pipeline was validated against at Cohen's kappa 0.85 — this is not a judge artifact. The number that should end the argument is the third: explicit safety recognition occurs in only 1.99% of runs. The agents are not weighing the risk and choosing wrong. They almost never notice there is a risk. And the exploitation rate is *nearly invariant to which model generated the malicious skill*, which puts the vulnerability in the interface rather than in any one adversary's cleverness.

A skill file is a natural-language document carrying shell commands, which the agent reads with the same trust it extends to its own instructions. No signature, no provenance, and per this paper effectively no recognition. The industry spent this week arguing about whether frontier models should be released at all, while the deployed enterprise attack surface turned out to be a folder you can drop a file into. The recommendation is stated without hedging: *"Enterprises must assess and mitigate skill-interface risk before adopting coding agents."*

arXiv 2608.05223 (primary)
V.

MATTER: A RUNTIME FOR THE BODY

The robot stack has spent two years shipping models and no runtime. Someone wrote the llama.cpp for embodiment.

228FIELD REPORTMATTER

The Robots Finally Got A Runtime.

Two years of shipping embodied models on model-specific Python stacks, and someone wrote the portable C++ inference runtime a control loop actually needs.

Embodied.cpp, a portable C++ inference runtime for embodied AI models on heterogeneous robots
IMAGEarXiv

**Embodied.cpp**, revised to v3 on August 9th, argues that embodied AI has a deployment problem no better model will fix. Vision-language-action models and world-action models both ship as *"model-specific Python stacks, backend assumptions, and robot-side glue code,"* and the inference runtimes that exist were built for request-response serving. That is the wrong contract for a body. Embodiment needs multi-rate execution inside a closed control loop, latency-first batch-1 inference on heterogeneous edge hardware, and I/O that is not fixed-shape tokens.

So the paper does for embodiment what llama.cpp did for text. It analyzes representative VLAs and world-action models, extracts the shared execution path, and organizes it into five layers — input adapters, sequence builders, backbone execution, head plugins, and deployment adapters — behind one backend abstraction spanning devices, robots, and simulators. Evaluated on three VLA and two world-action models using normalized Python-versus-C++ quantization comparisons, it reports 1.05x to 2.70x inference speedups and 7% to 77% lower VRAM.

The bottom of that VRAM range is the number that matters, because it decides whether a policy runs on the robot or on a workstation the robot is tethered to. Nearly every embodied demo that impressed anyone this year ran its inference somewhere with a wall socket. A runtime that is portable, quantization-aware, and organized around the control loop rather than the API endpoint is the unglamorous layer that has to exist before any of it walks out of the lab — and it is worth noticing that it arrived as a systems paper rather than a model release.

arXiv 2607.02501 (primary)