Newsgather
ZurückMeta enters coding agent race with Muse Code and Muse Spark 1.2
Meta enters coding agent race with Muse Code and Muse Spark 1.2
Technik
Decryptvor 8 StundenTechnik2 Min. Lesezeit

Meta enters coding agent race with Muse Code and Muse Spark 1.2

Meta launches Muse Code, a terminal coding agent powered by Muse Spark 1.2, featuring a crash-safe runtime and multimodal capabilities.

Auf einen Blick

Meta released Muse Code, a terminal coding agent powered by the new Muse Spark 1.2 model, competing with Anthropic and OpenAI in software engineering tasks.

KI-generierte Zusammenfassung

Warum es wichtig ist

Meta is expanding its AI model lineup to compete with established coding tools from Anthropic and OpenAI.

Schriftgröße

Meta is the latest tech giant to ship a coding agent, racing to compete with leading AI behemoths Anthropic and OpenAI.

"We're excited to release Muse Code (beta), a terminal coding agent powered by Muse Spark 1.2, our newest model," the company wrote in an official announcement. "This marks our next step toward the frontier, with larger and much more capable models on the way."

As an agentic coding tool, Muse Code is built for software engineering across large repositories. Per Meta, it "takes on complex software engineering tasks across large repositories: planning changes, writing code, and validating the results. It can coordinate multiple persistent subagents for each task, solving difficult problems faster, more accurately, and with less intervention."

The detail that stands out is the runtime. Muse Code logs every model call, tool run, approval, and edit to a local event log that acts as a single source of truth. "This single source of truth makes the runtime replay-exact and restart-safe: after a crash, the agent can resume precisely where it stopped," Meta said. For long-running jobs, that's the feature that matters more than raw speed—and it's the part competitors haven't made a selling point.

It also ships with default skills. The "/plan" command turns a task into an approval-gated plan, while "/grill" stress-tests that plan until it holds up and "/goal" works toward successful completion of the objective similar to what Hermes does. Meta said it co-trained Muse Spark 1.2 with Muse Code so the core LLM and the agent work together in synergy.

The benchmarks, and the catch

Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1. Meta said it "significantly scaled up training compute on coding tasks while expanding training environment diversity, delivering improvements in code generation, complex debugging, and end-to-end developer workflows." The charts tell a clear story.

On Terminal-Bench 2.1, Muse Spark 1.2 with Muse Code scored 82.9%, behind Claude Code on Opus 5 at 86.7% but ahead of GPT-5.6 Terra on Codex (81.8%) and Grok Build (81.6%).

DeepSWE 1.1, which measures agentic coding capabilities, was closer: 59.3% for Muse versus 65.0% for Opus 5 and 64.8% for Codex. On Meta's internal coding bench, Muse hit 70.6% to Opus 5's 79.4%.

The speedup charts flip the order. Over 1,000-plus tool calls, Opus 5 posted the biggest gain versus baseline (about 74–75%), with Muse Spark 1.2 mid-pack at roughly 61–69% depending on the run. Meta's point is that the agent keeps improving as tool calls accumulate, the behavior you want from a long-horizon coder.

The most interesting demos are long-horizon and multimodal. In stress testing, Meta said Muse Code "iteratively optimized GPU kernels over 1,000+ tool calls (up to 24 hours) on Nvidia Hopper GPUs." That means it was able to improve over time.

There's also a visual-coding angle. In one demo, a user drops a fly-through video of a house into the terminal as an mp4, and Muse Code "interprets the video and produces a visually rich website with booking capabilities." Reading raw video into a working web app is the multimodal pitch Meta has been making across the Muse line.

The field is already crowded

That said, Meta is late to the fight. OpenAI's Codex already runs parallel cloud agents; DeepSeek has built its own rival to Claude Code and agentic tools like Hermes or OpenClaw are already good substitutes with more capabilities. Muse Code's edge is the crash-safe runtime and the subagent design, not benchmark supremacy.

The risk is the usual one for agentic coding: an agent that resumes after a crash and keeps calling tools for 24 hours is powerful and unpredictable. Meta is betting developers want that autonomy, and it's shipping now.

Worauf zu achten ist

KI-Ausblick — Möglichkeiten, keine Fakten

  • Meta will release larger and more capable models following Muse Spark 1.2.

    Wahrscheinlich · Innerhalb von Monaten

Offene Fragen

  • When will Muse Code be generally available beyond beta?
  • How will enterprise adoption compare to existing tools?

Verwandte Themen

This article was originally published by Decrypt.

Ähnliche Meldungen

Mehr zu diesem Themameta